📖 The AI Tool Bible
training

Direct Preference Optimization (DPO)

A simpler alternative to RLHF that trains directly on preference pairs without a separate reward model. Widely adopted post-2024.

Related terms