Unlocking the Critic
Research note · RL post-training for LLMs Hongyang (Kevin) Li et al. · 2026

Reward-Free Policy Optimization

Unlocking the Critic

RL post-training for LLMs increasingly removes the critic, and even where a critic is trained it is discarded once training ends, although it has learned to predict outcomes. RFPO repurposes a single calibrated, frozen, well-pretrained critic as the reward. With half of every batch cut off at a 4,096-token cap, it validates above supervised PPO under the same cap, without verifier labels during RL and on 19% fewer GPU-hours.

University of Luxembourg · Seafill Open-Source Community · Université Paris-Saclay

The frozen critic's value at the last token

Share of rollouts per value bin, correct above the axis, incorrect below.

correctincorrect
47.6%mean validation over steps 10–300, against 46.6% for supervised PPO at the same cap.
~50%of RFPO's training rollouts are unfinished at every step and still receive a reward.
264 hGPU-hours for 300 steps, against 327 for PPO (−19%). No critic update, no verifier calls.
(i)Critic instability is an optimization artifact (ii)One frozen critic, three roles (iii)Calibration guards against length bias (iv)Parity at lower cost, built for long horizons

Contribution (i)

Critic instability is an optimization artifact

Critics were pushed out of long chain-of-thought RL because critic-based training looked unstable. We find the instability comes from the update recipe, not the value network: it recedes once each policy update is kept small and low in variance. Under this single-update rule, none of the three zero-label runs degenerates over 310 to 378 steps. Response length stays within 13% of its starting value, and policy entropy declines gradually from 0.32 to 0.21–0.22, against 0.25 for supervised PPO.

Contribution (ii)

One frozen critic, three roles

A well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. RFPO repurposes a single calibrated, frozen critic as the rollout-level reward, as the value baseline for GAE, and as a success forecaster for unfinished prefixes, so the RL loop contains no verifier and trains no value network. Any procedure that yields a critic able to rank trajectories well will do.

ROLLOUTS (tokens →) generation cap frozen critic V(x, y≤t) calibrated once reward1[v − b(ℓ) > τ], last token baselineV at every token, for GAE forecastscores rollouts cut off illustrative scores: 0.93 → reward 1 · 0.08 → reward 0 · truncated rollout scored at its last visible token
RFPO in one picture. No value network is trained and no verifier runs once the critic is frozen. The raw score has a length bias, so we subtract a length baseline b(ℓ) and binarize at a threshold τ; a continuous score lets the policy exploit the bias.

The value histogram at the top of this page is why this works: the frozen critic pushes incorrect rollouts toward 0 and correct ones toward 1 (AUC 0.96 on rollouts from the converged policy). It ranks attempts at the same problem just as well when the attempt never finished, so the reward is not simply detecting truncation. From a prefix alone, within-problem AUC reaches 0.84 at 4,096 tokens, and 0.84–0.92 on prefixes that have not produced an answer yet.

Ranking attempts at the same problem

Within-problem AUC of the frozen critic, averaged over ten policy checkpoints (16,384 rollouts each, 19–54% truncated).

Truncated attempts are ranked as well as the full set. Finished-only is lower because a problem that mixes finished and truncated attempts offers easy contrasts, and removing the truncated attempts removes them.

Ranking from a prefix

Within-problem AUC when the critic sees only the first n tokens.

all prefixesno answer yetshare finished
Hover to read a prefix length.

Contribution (iii)

Calibration guards against length bias

The critic's raw score carries a length bias, and a continuous reward lets the policy exploit it in either direction. We subtract a length baseline b(ℓ) and binarize at a threshold τ, which closes that channel. Used as a continuous reward, the same critic first climbs above fully supervised PPO, to 29.2% on AIME 2026 pass@1 against 25.0%, and then gives the gain back while response length drifts.

Response length drifts under a continuous reward

Change in mean training response length, last step against first, by debiasing strength. The shaded band is the range the binarized runs stay within.

Binarizing trades the ceiling for stability. With no debiasing the policy shortens its answers by 35% in 87 steps; with full debiasing it lengthens them by 34% and ends up truncating almost every rollout.

Contribution (iv)

Parity at lower cost, built for long horizons

Binarized, RFPO matches supervised PPO with no verifier labels during RL while cutting compute and memory: at the standard 5,120-token cap a step takes 421 s instead of 582 s, and peak memory falls by 9.4 GB per GPU. Because the critic scores truncated rollouts as accurately as complete ones, training does not have to wait for every trajectory to finish. Under a 4,096-token cap, with about half of every batch unfinished for all 310 steps, RFPO validates above supervised PPO trained under the same cap, on 19% fewer GPU-hours.

Validation under a 4,096-token training cap

Macro-average of AIME 2026 and AMC 2023 at 12,288 tokens. Dots: single evaluations; lines: centered means of five.

RFPO, 0% labelsPPO, 100% labels
Hover the chart to read a step.

Share of training rollouts truncated

PPO shortens its answers to fit the cap; RFPO keeps about half unfinished.

RFPOPPO
Hover to read a step.

Median seconds per RL step

Same node, same cap. The gap is PPO's critic update.

At the standard 5,120-token cap

Separate 16-sample evaluation after 300 RL steps, pass@1 (%)

ModelAIME25AIME26AMC23GPQAAvg
Initial policy22.721.071.444.940.0
PPO, 100% labels25.221.773.646.641.8
RFPO, 50% labels22.721.272.848.541.3
RFPO, 0% labels25.421.770.047.341.1

Level or ahead on AIME and GPQA, behind on AMC 2023 (40 problems, 2.5 points each). An RL step takes 421 s, against 582 s for PPO.

As reasoning traces lengthen and agentic episodes stretch, waiting for the outcome buys less and less for what it costs, and the value network is the one component of the standard recipe that speaks before the outcome arrives. So far our evidence comes from mathematical reasoning with one 4B model family and traces up to 12,288 tokens; larger models and other tasks are next.

The question is not whether a critic is affordable, but how much of what it already knows the current recipe throws away.

RFPO is our answer: it keeps that knowledge and turns it into a reward-free, compute-efficient method for LLM post-training.

Cite

@article{li2026unlocking,
  title   = {Unlocking the Critic: Reward-Free Policy
             Optimization for LLM Post-Training},
  author  = {Li, Hongyang and Li, Xiao and Wu, Caesar and
             Mammar, Said and Danoy, Gr{\'e}goire and Bouvry, Pascal},
  journal = {arXiv preprint},
  year    = {2026}
}