A simpler alternative to RLHF that trains directly on preference pairs without a separate reward model. Widely adopted post-2024.
training
A simpler alternative to RLHF that trains directly on preference pairs without a separate reward model. Widely adopted post-2024.