algo
now
.net
new ·
18
pairs
atlas
problems
fields
listen
quant
AI
philosophy
algonow
/
algorithms
/
Rejection sampling fine-tuning
Rejection sampling fine-tuning
Pairings in the atlas
Best-of-n distillation
specialist
Preference alignment
llm-training-alignment
Rivals: other methods for the same problems
Constitutional AI
Direct preference optimization
GRPO
IPO
KTO
ORPO
Process reward modeling
RLHF
Reward model training
Where it sits
llm-training-alignment
·
Machine Learning & AI