A training approach where human preferences over pairs of model outputs are used to train a reward model, which then guides RL fine-tuning to align model behaviour.
training
A training approach where human preferences over pairs of model outputs are used to train a reward model, which then guides RL fine-tuning to align model behaviour.