Skip to content

Reasoning Models Are Won in Post-Training

#rlvr #post-training #reasoning-models #grpo #self-distillation #test-time-training #granite-4-2

The distinguishing capability of a reasoning model is no longer decided by the architecture. It's decided by the post-training pipeline. A recent cluster of papers, plus IBM's Granite 4.2 release, makes the same point from different angles: RL with verifiable rewards (RLVR) is table stakes, and the real levers are exploration control, label-free adaptation, expert consolidation, and the systems engineering that keeps the whole loop affordable.

None of these methods touch the transformer block. They change what happens after pre-training. The cluster includes a critique-based inference method that reuses weak-model failures, a systematic comparison of three ways to consolidate RLVR experts, a test-time optimization that works without ground-truth labels, a video-LLM adaptation of on-policy self-distillation, and a systems taxonomy for parallel RL training. Granite 4.2 shows what the stack looks like when a vendor ships it: dense models from 3B to 30B, a five-phase pre-training run on 15T tokens, and a multi-stage RL ladder that ends in agentic tool use.

The Capability Lever Moved to Post-Training ​

The reason post-training dominates is simple. Base models have converged on the same architectural recipe: dense transformers, GQA, RoPE, SwiGLU. What separates a good reasoner from a mediocre one is how the model is pushed past imitation into sustained chain-of-thought behavior. That happens in the RL loop.

The papers in this cluster are best read as attacks on three specific bottlenecks. First, RLVR collapses the policy's diversity, hurting pass@k exactly where you want coverage. Second, RL and on-policy self-distillation (OPSD) both lean on ground-truth labels, which you don't have at test time. Third, the compute bill for all of this runs to millions of GPU-hours. Each result picks one bottleneck and takes it apart.

RLVR's Silent Killer: Entropy Collapse ​

Start with the problem that degrades most RLVR runs before you notice: entropy collapse. GRPO on verifiable rewards pushes a policy toward a few high-scoring solution templates. Average accuracy climbs. But sample K answers and you'll find they're variations of the same few approaches, so pass@k flattens exactly where it should reward coverage. When I've run GRPO loops on competition math, this shows up as a beautiful reward curve and a policy that writes the same three proof styles over and over.

The weak-model guidance paper in this cluster treats entropy collapse as a search problem, not a regularization problem. Instead of KL penalties or entropy bonuses, it forces the target model to continue partial reasoning trajectories sampled from a smaller, weaker model. Those unfamiliar prefixes disrupt the overconfident policy and push it into neighboring regions of the solution space. No extra SFT, no reward shaping, no prompting tricks.

The method beats vanilla RLVR across math benchmarks, and the gain grows as k in pass@k scales up. That shape is the tell: the improvement is coverage expansion, new solution paths rather than reweighted old ones. The compute trade is favorable too. A weak model's rollouts are cheap, and you spend them during training to perturb the teacher, not at inference time.

Key Numbers

  • 8.6 points: the widest single-benchmark gap between Merge, Mix RL, and MOPD, even though average performance differs by at most 1.4 points.
  • 38.0% to 45.2%: Qwen3-1.7B's test-time accuracy gain from TTPO on competition math, with zero ground-truth labels.
  • 15T tokens: Granite 4.2 pre-training volume, ending with a 512K-token context window.
  • 4,096: the GRPO batch size in Granite's RLVR stage, from 256 prompts with 16 responses each.

Inference-Time Scaling Has a Cheaper Sibling ​

Inference-time scaling buys reasoning with tokens. The standard recipe is simple: sample many candidates at test time and pick the best one. It works, but it multiplies inference cost by the number of samples. CritICL makes a cheaper bet: failure modes are structured across model scales within a family, and a weak model's mistakes are reusable guidance for a strong one.

CritICL turns failures into critique-based in-context examples. The static variant builds a global failure-mode profile from a small model and prepends it to the prompt. The dynamic variant predicts the failure mode for the specific input, retrieves matching critiques, and uses those instead. Either way, the weak model's failure modes become a critique layer over the strong model's output.

It's weak-to-strong generalization at inference time, and the cost is a few hundred tokens per prompt rather than a dozen parallel rollouts. The paper reports results competitive with, or better than, test-time scaling methods on reasoning benchmarks. The open question is cross-family transfer, since failure modes are structured within a family, not across all models. But if you already serve a small model alongside a large one, this is nearly free signal.

Quick Take: the papers in this cluster don't compete on architecture. They compete on what surrounds the policy: how exploration is controlled, how labels are replaced or routed, how experts are consolidated, and how many tokens the whole thing burns.

Three Ways to Fuse RLVR Experts ​

If you've trained separate RLVR experts for different domains, you'll eventually face the consolidation problem: how do you turn three specialists into one model? The fusion-paradigms paper organizes the options by what they reuse.

ApproachReusesNeeds experts?Main constraintBest when
MergeExpert task vectorsYesCompresses all updates into one vector; sensitive to task-vector geometryExperts exist and cheap fusion is the priority
Mix RLDomain datasetsNoSensitive to domain mixture proportionsTraining a unified model from scratch without experts
MOPDDatasets and teacher policiesYesBounded by teacher qualityDomain-specific gains matter more than beating teachers or minimizing cost

Headline averages are deceptively calm. The three approaches differ by at most 1.4 points in average performance across the benchmark suite. The distribution tells the real story: on a single benchmark the gap reaches 8.6 points, and the variation tracks cross-domain relations visible in task-vector geometry. If your experts learned conflicting behaviors, Merge smears them together. I've hit that failure mode myself: two expert vectors pointing in different directions, and the merged model solved neither task well.

Mix RL is sensitive to mixture proportions, so the paper recommends adjusting them for cross-domain transfer. MOPD stays bounded by its teachers, which is fine when preserving domain gains matters more than beating them. All three improve single-sample accuracy without measurable gains in solution coverage, and none hurt held-out capabilities.

Test-Time Training Without Labels: TTPO ​

Test-time training keeps adapting the model after deployment, on the current problem, with no labels. The natural approach is majority voting: sample K rollouts, take the vote as the pseudo-label, train against it. It's fragile in a specific way. One incorrect vote corrupts the teacher, and every token trained against it carries the error. In my own experiments, the failure shows up as a sudden divergence a few steps after a bad vote lands.

TTPO starts from an asymmetry that makes this tractable: rollouts that disagree with the pseudo-label are usually wrong, whether the pseudo-label itself is right or not. So TTPO routes the two branches differently. Agreeing rollouts are distilled with OPSD. Disagreeing rollouts are penalized with grouped RL. Token-level selection sharpens both: distillation down-weights positions the model already converged on, and RL penalizes only confident errors.

The results are strong enough that the method feels underrated. Without any labels, TTPO matches label-supervised OPSD on five competition-level benchmarks. It lifts Qwen3-1.7B from 38.0% to 45.2% at test time, and the gains survive without thinking mode, +25.2% to +36.4%. A 1.7B model is small enough to run on-device, which makes label-free test-time adaptation the difference between a fixed model and one that improves on the user's actual problems. The mechanism compounds too: majority-vote routing produces tighter self-supervision as the model improves.

OPSD Escapes the Text Box: Video-OPSD ​

OPSD doesn't get the attention it deserves. The setup is elegant: the policy itself, augmented with privileged information, acts as the teacher, and the student trains against the teacher's token-level targets on its own rollouts. Video-OPSD extends this to video LLMs, where the privileged information hides inside the input.

Long videos are mostly redundant. A question usually depends on a small subset of frames, and the paper's Evidence-Grounded Self-Teacher conditions the teacher only on those evidence frames while the student reasons over the full video. A companion mechanism, Evidence-Guided Token Optimization, weights the distillation loss per token by how much that token relies on the privileged evidence.

Video-OPSD beats standard OPSD across backbones and matches GRPO-level performance with substantially less training time. The broader lesson extends beyond video: any domain where you can cheaply isolate the decisive subset of the input has this structure, and the evidence-weighting trick is a template for it.

The Systems Bill Is the Real Constraint ​

None of this is cheap. The systems paper in the cluster puts a number on it: state-of-the-art reasoning model training requires millions of GPU-hours and multi-model pipelines that stress hardware far beyond supervised training. It frames RL-for-LLM as a distributed systems problem as much as an algorithmic one, and delivers a taxonomy of parallelism: data, tensor, pipeline, sequence, context, and expert parallelism, plus newer options like disaggregated placement, stage fusion, and asynchronous execution.

The async pattern is the one that shows up in production. Granite 4.2 runs asynchronous GRPO end to end. A pool of generation workers samples responses into a shared buffer, the trainer consumes a full step from the buffer, takes an optimizer step, and streams updated parameters back without pausing the generators. A refresh can land mid-rollout, leaving one trajectory stitched from two adjacent policy versions. Rather than paying to prevent that, they reuse the KV cache and cap the drift: workers stay at most one update behind the trainer, and truncated importance sampling clamps the train-versus-generation log-probability ratio so stale tokens can't dominate an update. That's the difference between a research loop and a product.

Granite 4.2 Runs the Whole Playbook ​

Granite 4.2 is the most complete public example of this playbook, and it's Apache 2.0 licensed, so you can actually run it. The family is dense and decoder-only, 3B, 8B, and 30B, with GQA attention, RoPE at theta 10M, SwiGLU, and RMSNorm. Pre-training runs from scratch on roughly 15T tokens in five phases, with the last phase extending the context window to 512K. That's enough to drop in a large document set or a whole repository without chunking.

The SFT stage mixes 68.4% non-agentic data with 31.6% agentic data, roughly 7.2 million samples and 100B tokens, of which 65B are trainable. The agentic mix is lopsided on purpose.

Post-training is a graduated curriculum of separate GRPO runs, each warm-started from the previous checkpoint. The KL schedule tracks the reward type: zero for verifiable stages (RLVR, SWE 2), nonzero for preference, safety, and skill-graft stages. The stage shapes are just as deliberate.

StagePrompts/stepGens/promptMax seq lenRollout turnsKL
RLVR (x3)2561664K10
IF booster2561664K10
Code booster641664K10.05
SWE 16416128K10.01
SWE 23216128K1280
Terminal83264K640.01
Search3216128K640.01
RLHF1281648K10.05

Look at where the compute goes. RLVR spends 4,096 rollouts per step on broad skill coverage. Terminal runs only 8 prompts per step but 32 generations each, because agentic reward is a sparse bit at the end of a long tool-use trajectory. SWE 2's 128 rollout turns are the same idea: you need many attempts before the model earns a single confirmation that it solved the issue. Every model gets the foundational ladder plus RLHF. Only the 8B and 30B get the agentic block. The models also ship a thinking/non-thinking switch and a low-effort thinking mode that spends a short reasoning budget on easy questions, a sensible admission that not every prompt needs a chain of thought.

Common Pitfalls ​

  1. Grind RLVR to convergence and ignore the entropy curve. Reward goes up while the policy collapses onto a few solution templates. Watch pass@k at large k; if it flattens, inject cheap diversity like weak-model prefixes before you touch reward shaping.

  2. Apply majority-vote pseudo-labels symmetrically at test time. One wrong vote corrupts the teacher and every token trained against it. The asymmetric split in TTPO, distill agreements, penalize disagreements, and only confident errors at that, is what keeps the update grounded.

  3. Merge expert task vectors without checking their geometry. Conflicting experts produce an 8.6-point regression on a single benchmark, which averages hide. Check task-vector similarity first, and expect Merge to compress away exactly what the experts disagree on.

  4. Run every GRPO stage with the same KL and rollout shape. Granite runs KL 0 for verifiable rewards and 0.05 for preference stages, and moves rollout turns from 1 to 128 depending on whether the reward is dense or a sparse terminal bit. Copying one config across all stages leaves capability on the table.

  5. Treat the systems engineering as an afterthought. Asynchronous generation, KV-cache reuse, and truncated importance sampling are what make a multi-stage RL loop affordable. Without them, the algorithm doesn't matter, because you won't be able to run it long enough to see results.

Sources ​

  • CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes (arXiv:2608.27455)
  • Consolidating RLVR Capabilities Across Domains (arXiv:2608.27409)
  • Boosting LLM Exploration via Weak-Model Guidance in RLVR (arXiv:2608.27420)
  • TTPO: Test-Time Policy Optimization (arXiv:2608.27448)
  • Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video Large Language Models (arXiv:2608.27065)
  • Performance Foundations of Parallel & Distributed Reasoning Language Models (arXiv:2608.27046)
  • Granite 4.2 LLMs: How They're Built (Hugging Face Blog)

One thing to remember: the reward signal is half the story. Every strong result in this cluster pairs a reward with a mechanism that keeps the policy diverse and honest, weak-model prefixes against entropy collapse, asymmetric routing on pseudo-labels, evidence-weighted token supervision in video, or a KL schedule that matches the reward type. The reward says what good looks like. The surrounding pipeline decides whether the model can still find it.

The Bottom Line ​

  1. If you