Reward-Free Policy Optimization
RL post-training for LLMs increasingly removes the critic, and even where a critic is trained it is discarded once training ends, although it has learned to predict outcomes. RFPO repurposes a single calibrated, frozen, well-pretrained critic as the reward. With half of every batch cut off at a 4,096-token cap, it validates above supervised PPO under the same cap, without verifier labels during RL and on 19% fewer GPU-hours.
University of Luxembourg · Seafill Open-Source Community · Université Paris-SaclayThe frozen critic's value at the last token
Share of rollouts per value bin, correct above the axis, incorrect below.
Contribution (i)
Critics were pushed out of long chain-of-thought RL because critic-based training looked unstable. We find the instability comes from the update recipe, not the value network: it recedes once each policy update is kept small and low in variance. Under this single-update rule, none of the three zero-label runs degenerates over 310 to 378 steps. Response length stays within 13% of its starting value, and policy entropy declines gradually from 0.32 to 0.21–0.22, against 0.25 for supervised PPO.
Contribution (ii)
A well-pretrained critic estimates the posterior probability of eventual success from later trajectory states and unfinished prefixes. RFPO repurposes a single calibrated, frozen critic as the rollout-level reward, as the value baseline for GAE, and as a success forecaster for unfinished prefixes, so the RL loop contains no verifier and trains no value network. Any procedure that yields a critic able to rank trajectories well will do.
The value histogram at the top of this page is why this works: the frozen critic pushes incorrect rollouts toward 0 and correct ones toward 1 (AUC 0.96 on rollouts from the converged policy). It ranks attempts at the same problem just as well when the attempt never finished, so the reward is not simply detecting truncation. From a prefix alone, within-problem AUC reaches 0.84 at 4,096 tokens, and 0.84–0.92 on prefixes that have not produced an answer yet.
Ranking attempts at the same problem
Within-problem AUC of the frozen critic, averaged over ten policy checkpoints (16,384 rollouts each, 19–54% truncated).
Truncated attempts are ranked as well as the full set. Finished-only is lower because a problem that mixes finished and truncated attempts offers easy contrasts, and removing the truncated attempts removes them.
Ranking from a prefix
Within-problem AUC when the critic sees only the first n tokens.
Contribution (iii)
The critic's raw score carries a length bias, and a continuous reward lets the policy exploit it in either direction. We subtract a length baseline b(ℓ) and binarize at a threshold τ, which closes that channel. Used as a continuous reward, the same critic first climbs above fully supervised PPO, to 29.2% on AIME 2026 pass@1 against 25.0%, and then gives the gain back while response length drifts.
Response length drifts under a continuous reward
Change in mean training response length, last step against first, by debiasing strength. The shaded band is the range the binarized runs stay within.
Contribution (iv)
Binarized, RFPO matches supervised PPO with no verifier labels during RL while cutting compute and memory: at the standard 5,120-token cap a step takes 421 s instead of 582 s, and peak memory falls by 9.4 GB per GPU. Because the critic scores truncated rollouts as accurately as complete ones, training does not have to wait for every trajectory to finish. Under a 4,096-token cap, with about half of every batch unfinished for all 310 steps, RFPO validates above supervised PPO trained under the same cap, on 19% fewer GPU-hours.
Validation under a 4,096-token training cap
Macro-average of AIME 2026 and AMC 2023 at 12,288 tokens. Dots: single evaluations; lines: centered means of five.
Share of training rollouts truncated
PPO shortens its answers to fit the cap; RFPO keeps about half unfinished.
Median seconds per RL step
Same node, same cap. The gap is PPO's critic update.
At the standard 5,120-token cap
Separate 16-sample evaluation after 300 RL steps, pass@1 (%)
| Model | AIME25 | AIME26 | AMC23 | GPQA | Avg |
|---|---|---|---|---|---|
| Initial policy | 22.7 | 21.0 | 71.4 | 44.9 | 40.0 |
| PPO, 100% labels | 25.2 | 21.7 | 73.6 | 46.6 | 41.8 |
| RFPO, 50% labels | 22.7 | 21.2 | 72.8 | 48.5 | 41.3 |
| RFPO, 0% labels | 25.4 | 21.7 | 70.0 | 47.3 | 41.1 |
Level or ahead on AIME and GPQA, behind on AMC 2023 (40 problems, 2.5 points each). An RL step takes 421 s, against 582 s for PPO.
As reasoning traces lengthen and agentic episodes stretch, waiting for the outcome buys less and less for what it costs, and the value network is the one component of the standard recipe that speaks before the outcome arrives. So far our evidence comes from mathematical reasoning with one 4B model family and traces up to 12,288 tokens; larger models and other tasks are next.
The question is not whether a critic is affordable, but how much of what it already knows the current recipe throws away.
RFPO is our answer: it keeps that knowledge and turns it into a reward-free, compute-efficient method for LLM post-training.
Cite
@article{li2026unlocking,
title = {Unlocking the Critic: Reward-Free Policy
Optimization for LLM Post-Training},
author = {Li, Hongyang and Li, Xiao and Wu, Caesar and
Mammar, Said and Danoy, Gr{\'e}goire and Bouvry, Pascal},
journal = {arXiv preprint},
year = {2026}
}