Appearance
Four Ways to Fix LLM Post-Training Instability
The instability problem
Post-training an LLM is a battle against drift. You start with a model that answers questions reasonably well, run policy optimization on it, and within a few hundred steps the outputs get weird: repetitive, collapsed, or silently worse at the original task. The standard fix was a KL penalty clamping the policy to a reference model, plus group sampling so you could compute relative advantages without training a critic.
Group sampling, the GRPO approach, works. It's also expensive. Estimating an advantage for one prompt means generating 8 to 16 responses, scoring them all, and using each response's rank within the group as its advantage. You pay in sampling cost during training, and you lose token-level credit assignment. A group-relative advantage tells you response 3 beat response 7. It doesn't tell you which token in response 3 made the difference.
Four papers from the August 2026 arXiv batch attack this problem at four different stages of the pipeline. BPCO makes the critic stable enough to replace group sampling. SRPO removes the critic and the reward model entirely and has the model critique its own trajectories. ERPO moves regularization from the response side to the query side. CRPO fixes the preference data itself so alignment transfers across languages. None of them need a new architecture. All of them plug into existing GRPO, PPO, or REINFORCE pipelines.
All four are preprints, so treat the numbers as pre-review. The recipes are concrete enough to test on your own pipeline this week.
Key Numbers
- 8 to 16 responses per prompt: what GRPO samples to estimate advantages without a critic.
- 1 response per prompt: what BPCO needs to match or beat group-based results.
- 0.08x FLOPs: SRPO's training cost relative to scaled SFT for its AIME'24 result.
- 6 math benchmarks: where ERPO's Query-KL beats the standard Policy-KL regularizer.
BPCO: keep the critic, make it stable
The reason GRPO dropped the critic isn't that critics are useless. It's that critic training is unstable. Value heads drift, bootstrap errors compound, and the advantages they produce quietly corrupt the policy gradient. BPCO (Best-Practice Critic Optimization) is a recipe that fixes the specific failure modes.
The recipe combines five pieces: DPPO as the base algorithm, value predictions bounded to the reward range, Monte Carlo value targets, unnormalized policy advantages, and length-adaptive GAE. Two of these do most of the work. Bounding value predictions to the reward range stops the value head from drifting into nonsense territory. Monte Carlo targets, instead of bootstrapped ones, stop error from compounding across training steps.
BPCO also exploits a trick group-based methods can't. Because the critic is used only during training, you can condition it on reward-defining information the policy never sees: reference answers, grading rubrics, whatever the reward depends on. The critic gets to peek at the answer key. The policy doesn't.
Across math reasoning tasks with models from 1.5B parameters to a 30B-A3B mixture of experts, BPCO matches or beats a group-based baseline while sampling one response per prompt. If you're running GRPO with 8 responses, that's an 8x cut in sampling cost during training, plus token-level advantages, in exchange for training a value head. The same recipe also improves learning with rubric-based rewards, which is where the answer-key conditioning pays off.
Quick Take: The field's default answer to RLHF instability was to throw away the critic and sample groups, but these four papers show a better path: fix the specific broken piece, whether that's the critic, the regularizer, the reward signal, or the data.
SRPO: teach the model to critique itself
SRPO takes the opposite bet. Don't fix the critic, remove it. Self-Reflective Policy Optimization has the model analyze its own completed trajectories, compress the errors into short "reflection patches," and then use reflection-conditioned teacher scores on fresh on-policy rollouts as dense token-level training signals.
The key move is converting sparse terminal feedback into dense signals without external machinery. No reward model, no critic, no larger teacher. The model generates the reflection itself, and the reflection conditions the scoring of the next rollout. Instead of "this trajectory failed," the model gets "this trajectory failed because step 3 picked the wrong tool, and here's what correct behavior looks like."
The numbers tell the story. A Qwen3-8B base model trained with SRPO hits 73.3% on AIME'24 using 8% of the training FLOPs that scaled SFT needed. That's an 8B model at near-frontier math performance on a budget that fits on a single node. On agentic benchmarks it reaches 64.7% on WebShop, 76.8% on ALFWorld, and 31.2% on SWE-Bench-Lite.
The practical payoff: if your bottleneck is compute or reward signal quality, SRPO is the most direct answer. It's also the most debuggable, because reflection patches are inspectable. You can read what the model thinks it did wrong, and that's a debugging surface neither GRPO nor critic-based methods give you.
ERPO: regularize the input, not the output
ERPO is the cheapest fix in this batch, and the one I'd try first. Environment-Regularized Policy Optimization targets the stability-exploration dilemma directly.
Here's the dilemma. Policy optimization needs a KL regularizer to stop the policy from drifting off the reference distribution. But the standard regularizer sits on the action side: it penalizes responses that diverge from the reference policy. That penalty constrains response behavior and consumes the exploration budget. Crank it up and training is stable but the model won't explore. Turn it down and the model explores but drifts.
ERPO's argument is that you're regularizing the wrong side. The thing that actually drifts during training is the query distribution, not the response distribution. As the policy changes, the distribution over training queries induced by the current policy shifts away from its pre-RL reference. So ERPO adds a Query-KL term that bounds that shift, plus a per-query weight derived from the static reference that biases updates toward typical queries.
The important detail is what the QKL gradient touches. It flows through the query likelihood only. The response score function from the policy-gradient estimator doesn't appear in the QKL term, so QKL exerts no direct pressure on the response distribution. Exploration is preserved. And because it's just a loss term, it plugs into GRPO, PPO, or REINFORCE without extra forward passes.
Across six math reasoning benchmarks, ERPO replaces the standard Policy-KL regularizer, controls query drift, and delivers better accuracy with substantially more stable behavior under high-temperature decoding and long-horizon training. If your runs are drifting after a few hundred steps, this is a one-line change to the loss. Try it before you rebuild anything.
CRPO: fix the data, not the algorithm
The last paper in the batch is about data, not optimization dynamics. CRPO (Cross-Lingual Ranking Preference Optimization) starts from a fact that's easy to forget: most preference data is English, and alignment trained on it doesn't transfer well to other languages.
The fix is a hierarchical structure over parallel preference pairs. For a given prompt, you have candidate responses in the target language and in English. CRPO optimizes intra-lingual preferences, the ranking within each language, and inter-lingual preferences, the ranking across languages, jointly. The English preference knowledge acts as a bridge: the model learns what good looks like in English and carries that signal into the target language.
CRPO also goes beyond binary comparison. It's built on the LambdaLoss framework, so it uses a relative ranking signal across multiple candidate responses instead of pairwise win/loss. That difference matters. Binary preference data throws away the information in the ordering of candidates 2 through 5. A ranking signal keeps it.
Across five languages with varying resource levels, CRPO beats standard approaches on instruction-following and knowledge utilization. The paper also reports larger reward margins and higher log-probability of desirable responses, which is a proxy for a more stable preference manifold. If you're aligning a multilingual model and your preference data is mostly English, this is the paper to read.
How the four approaches compare
All four methods respond to the same failure mode, but they're not interchangeable. They replace different components.
| Method | Core idea | Replaces | Cost profile | Best fit |
|---|---|---|---|---|
| BPCO | Stable critic: bounded values, MC targets, length-adaptive GAE | Group sampling in GRPO | One response per prompt, plus value-head training | Math reasoning, rubric-based rewards |
| SRPO | Self-reflection patches produce dense token-level teacher scores | Critics, reward models, group sampling | No RM or critic; 0.08x SFT FLOPs | Long-horizon reasoning, agentic tasks |
| ERPO | Query-KL bounds input distribution drift | Policy-KL regularizer | No extra forward passes | Long-horizon training, high-temperature decoding |
| CRPO | Hierarchical intra- and inter-lingual ranked preferences | English-only binary preference data | Preference data construction | Multilingual alignment |
The choice comes down to which part of your pipeline is the weakest link. If sampling cost dominates your training budget, BPCO is the change with the biggest payoff. If your runs are unstable and you don't know why, ERPO is a five-minute experiment. If you're aligning a multilingual model, CRPO addresses a problem the other three don't touch. If you have a long-horizon agentic task with sparse rewards, SRPO is the one that turns terminal feedback into something learnable.
Common Pitfalls
Don't let the critic predict values outside the reward range. BPCO found that bounding value predictions to the reward range is one of the largest stability wins. An unconstrained value head drifts, and the advantages it produces will corrupt your gradient estimates before the loss curve shows anything wrong.
Don't crank up the Policy-KL coefficient to fix instability. You'll consume the exploration budget and the policy will collapse to safe, repetitive outputs. ERPO's whole argument is that the drift you're fighting lives on the query side. Regularize there instead.
Don't collapse multi-candidate preference data into binary pairs. If you have four ranked responses, a LambdaLoss-style ranking signal beats four pairwise comparisons. CRPO shows the ranking signal produces a more stable preference manifold, and you get it for free from data you already collected.
Don't train a multilingual model on English-only preference pairs and expect alignment to transfer. CRPO's hierarchical intra- and inter-lingual ranking is what actually moves target-language behavior. Without the cross-lingual term, the English preference knowledge stays locked in English.
If you're using self-reflection, don't feed the model raw outcome rewards. SRPO's reflection patches are what convert sparse terminal feedback into dense token-level signals. Skip the reflection step and you've just recreated sparse reward with extra steps.
One thing to remember
Every method in this batch is a drop-in change to an existing pipeline. No new architecture, no new training framework, no new data collection setup. The fastest win for most teams is ERPO's Query-KL, because it costs nothing to try and it addresses the drift that shows up in every long training run.
The Bottom Line
If you're running GRPO with 8 or more responses per prompt, adopt BPCO: it cuts sampling cost to one response per prompt, gives you token-level advantages, and the critic can condition on the rubric without leaking it to the policy.
If your training runs are unstable or your compute budget is tight, start with ERPO's Query-KL regularizer. It's a one-line change to the loss that fixes drift without touching exploration, and you'll know within a few hundred steps whether it works.
One thing to watch: self-reflection methods like SRPO are moving fast, and the reflection patch is an inspectable intermediate artifact. Expect to see reflection-conditioned signals used for reward modeling and data filtering within the next year, not just for policy updates.