In reinforcement learning, the simplest way to train a model is an Outcome Reward Model (ORM): if the model reaches the correct final answer ($+1$), reward it; if the final answer is wrong ($-1$), penalize it. But in complex multi-step reasoning, outcome rewards suffer from a fatal flaw: the credit assignment problem.
The Illusion of the Lucky Guess
Consider a model solving a ten-step physics problem. If the model makes a catastrophic mathematical error in Step 3, but makes an equal and opposite blunder in Step 7 that coincidentally cancels out, it arrives at the correct final number. An outcome reward model blindly rewards this broken trajectory with a score of $+1$, reinforcing flawed reasoning.
[Outcome Reward Model (ORM): Evaluates Only Final Output]
Step 1 ✓ ──► Step 2 ✓ ──► Step 3 (Fatal Math Blunder!) ──► Step 4 ──► Final: 42 (Lucky Guess!) ──► Score: +1.0 (Flawed!)
[Process Reward Model (PRM): Dense Step-by-Step Credit Assignment]
Step 1 [Score: 0.99 ✓] ──► Step 2 [Score: 0.98 ✓] ──► Step 3 (Math Blunder!) [Score: 0.05 ✗ - PRUNED!]
│
└──► Backtrack & Explore Alternate Branch ✓
The Process Reward Model (PRM) Revolution
Process Reward Models evaluate reasoning densely, assigning a verified probability score to every individual deductive leap:
- Instant Error Localization: The system pinpoints the exact token position where reasoning diverged from truth.
- Active Search Guidance: During test-time search (MCTS or Best-of-N), the search engine prunes branches the instant a low-scoring step is detected, avoiding wasting compute on doomed trajectories.
- Elimination of Hallucination Loops: Models are rewarded strictly for sound deductive steps, eliminating sycophantic rationalizations.
Dense step-level verification is the cornerstone of reliable mathematical and systems reasoning.