RLHF
Also known as Reinforcement learning from human feedback. This is the canonical page; those names redirect here.
Pairings in the atlas
- Proximal policy optimization on a reward modelcanonPreference alignmentllm-training-alignment
Also known as Reinforcement learning from human feedback. This is the canonical page; those names redirect here.