now.net

Process reward modeling

Also known as Step-level reward model. This is the canonical page; those names redirect here.

The usual abbreviation PRM is not claimed here: robotics already redirects it to Probabilistic roadmap.

Pairings in the atlas

Rivals: other methods for the same problems

Where it sits

llm-training-alignment · Machine Learning & AI