Preference alignment
10 methods in the atlas attack this one problem. They are rivals: each wins something the others do not.
Phrasings that mean this problem
Preference alignmentReward modeling
llm-training-alignment
- RLHFProximal policy optimization on a reward modelcanonllm-training-alignment
- Direct preference optimizationImplicit reward reparameterizationcanonllm-training-alignment
- IPOIdentity preference objectivespecialistllm-training-alignment
- KTOProspect-theoretic utility lossspecialistllm-training-alignment
- ORPOOdds-ratio preference lossspecialistllm-training-alignment
- GRPOGroup-relative advantage baselinestandardllm-training-alignment
- Constitutional AISelf-critique and revisionstandardllm-training-alignment
- Rejection sampling fine-tuningBest-of-n distillationspecialistllm-training-alignment
- Reward model trainingBradley-Terry pairwise lossstandardllm-training-alignment
- Process reward modelingStep-level supervisionspecialistllm-training-alignment