Appearance
Rewards, Distillation, and 100 Steps: What RL Post-Training Actually Looks Like Right Now
The field moved while you weren't looking
RL post-training has split into two very different camps. One camp grinds out multi-million-dollar runs on frontier models, producing systems that outscore gold-medal human programmers. The other camp fine-tunes a 350M model for 100 GRPO steps on a free Colab GPU and picks up 7 points of structured-output accuracy.
Both are valid. Both tell you something important about where the field is heading.
The papers and tutorials landing this week share a common thread: everyone is trying to squeeze more signal out of less compute. That means better reward design, smarter distillation, and a willingness to ask what the model actually needs to learn rather than just throwing data at it.
Outcome rewards hit a wall, so people are shaping them
The big shift is from coarse outcome rewards to something with finer granularity. The Cliff paper makes the cleanest argument: once a reasoning process has made its first mistake, every subsequent token is conditioned on an invalid prefix. Evaluating that suffix tells you almost nothing. The signal you actually want is in the boundary where the reasoning first went wrong.
Cliff uses an off-the-shelf LLM as a teacher to locate that first mistake in each rollout. The rollout splits into a correct prefix and an incorrect suffix, and those get positive and negative token-level advantages respectively. That's a much more informative gradient signal than a single scalar at the end.
The setup is deliberately simple. Process reward models need a specialized RM and its training data. On-policy distillation assumes teacher and student reason identically. Cliff needs neither. The paper reports consistent gains across 12 scenarios, beating on-policy distillation by 15% and standard GRPO by 7%, even with modest teachers.
The practical implication: if you're running GRPO on reasoning tasks, a single binary correctness reward is leaving signal on the table. The same off-the-shelf judge that verifies final answers can tell you where the chain went off the rails.
LLM judges have a blind spot, and it's systematic
The user feedback paper from the same batch exposes something uncomfortable. When a model successfully fixes an issue because of feedback, LLM judges frequently fail to recognize the corrected response and prefer the inferior baseline instead.
This isn't a fluke of one judge setup. The authors built synthetic data with definitive ground truth, plus naturalistic data to check real-world validity. In both settings, feedback-informed revisions resolved targeted issues at significantly higher rates than baseline revisions. The problem was purely in the evaluation: judges systematically rated worse outputs higher.
I've seen this pattern myself. When I evaluate candidate responses with a strong LLM judge, it latches onto fluency and misses factual corrections, especially when the corrected version is more cautious or hedged. The implications for reward design are direct: if you're using an LLM judge as a reward model, the noisy signals you get aren't just noise. They're biased in a way that actively discourages the model from improving.
The fix isn't obvious. The paper doesn't offer a clean one, it just isolates the bias and proves it exists. But if you're building a reward stack, this is a strong argument for verifiable rewards when you can get them, and for careful calibration when you can't.
Distillation's dirty secret: domain labels lie
Another paper tackles the multi-domain distillation problem from a different angle. The standard approach routes each training sample to the teacher whose domain matches. The MT-SDPO paper points out the flaw: domain expertise holds on average, not per sample. The matched teacher often gets a sample wrong while a teacher from another domain nails it.
The reliable teacher has to be identified per sample, not per domain. MT-SDPO does this with three components: self-anchors, where a rollout is supervised by a correct rollout from the same group; answer-verified eligibility, where a teacher can only supervise samples its own answer passes a verifier; and privileged distillation, which merges anchor and verified feedback into one context for an EMA self-teacher that the student can't see.
The numbers make the case. Across five students from three model families, MT-SDPO lifts the weakest domain of Qwen3-8B by 14.79 points and narrows the domain gap by 74.7%. That's a better balance than serving one matched teacher per domain, which is the alternative most teams would reach for.
| Method | Weakest-domain lift | Domain gap reduction | Extra inference cost at deploy |
|---|---|---|---|
| Single matched teacher per domain | baseline | baseline | none, but requires per-sample routing |
| MT-SDPO | +14.79 points | 74.7% | none, single policy at serving time |
The deployment implication matters here. MT-SDPO keeps one policy at serving. You don't need to route, you don't need multiple models, and the student learns from whoever was right on any given sample.
500 samples, 100 steps, one free GPU
The IFStruct tutorial from Hugging Face is the counterweight to the multi-million-dollar runs. A 350M Liquid model scores 22.6% on the IFStruct structured-output benchmark. After 100 GRPO steps with about 500 training samples and a LoRA adapter, it scores 29.7%.
The gains are exactly where the training aimed. JSON pass rate jumps from 18.0% to 31.9%, bare-list outputs from 16.6% to 29.7%. YAML, which the training didn't target, barely moves. That's the signature of a well-specified reward.
Three reward functions drive this: JSON-format compliance, field-count accuracy, and schema validation, weighted 1.0, 0.5, and 2.0. The schema weight dominates because structural validity is the thing that actually breaks downstream integrations.
The takeaway isn't that 350M models are secretly great. It's that structured output is a narrow, well-defined skill, and a narrow reward signal can teach it fast. The notebook runs on free-tier Colab. The entire fine-tuning budget is smaller than most teams spend on a single prompt-engineering workshop.
When the reward is taste
The watercolour painting tutorial takes the opposite extreme. The model writes JavaScript that paints images through the p5.brush library, and the reward is aesthetic preference. There's no verifiable answer at all.
The reward stack is fascinating because it's entirely made of pieces that shouldn't work together. A gate checks that the sketch compiles and doesn't cheat. A pairwise judge, Qwen3-VL-30B-A3B-Instruct, compares the candidate painting against four references from a hand-curated pool of 178 images. HPSv3, an open 7B preference model, scores how much a person would prefer the render. Those two judges split 0.60/0.30 of the reward weight, with the gate and a length bonus making up the rest.
The pool is the reward function. The author hand-rated 178 model-generated paintings into love and okay tiers, and the pairwise judge draws half its references from each. The policy always faces rivals it can sometimes beat, and a win pays the same against either tier. The runner deck changes, the reward changes with it. No reward-function code needs to touch it.
I ran into the same wall during the debugging stretch: flat reward curves for a long time. The cause was boring, my learning rate was too low. The LoRA target-module list for the Qwen MoE model was the second trap, the usual assumption of dense-model naming meant the adapter was training ten out of forty layers. The fix was all-linear, which reaches every linear layer.
| Run | Steps | First third reward | Final third reward | Δ |
|---|---|---|---|---|
| hps-only | 60 | 0.58 | 0.71 | +0.13 |
| judge-led | 110 | 0.45 | 0.72 | +0.27 |
| hps-led | 110 | 0.57 | 0.82 | +0.24 |
The reward curves line up with how much of the author's taste the mix carries. Judge-led starts lower and noisier, and spends its first thirty steps nearly flat before moving. That's the cost of a harder objective.
The endgame: outscoring human gold at IOI
The competitive programming pipeline is where all these ideas converge. The Nemotron team built a specialization pipeline with 22,000 curated problems, synthetic reasoning traces, SFT, and RL, plus a test-time strategy called GenCorrect that iteratively generates, evaluates, and refines diverse solutions.
The numbers deserve attention. The 30B-A3B Nano-CC goes from 130 to 291 IOI 2025 points after post-training, and to 468 with GenCorrect. The gold threshold was 438.3. The 550B-A55B Ultra-CC reaches 502. Then, prospectively evaluated during IOI 2026 under the same time, internet-access, and submission constraints as human contestants, the Ultra-CC system scores 535.4 out of 600. The gold threshold was 361.12. The top human scored 498.27.
That last point is worth sitting with. This is the first AI system to outscore the highest-scoring human contestant on an IOI problem set, under identical constraints.
The system combines two things that the other papers treat as alternatives. RL post-training for the base capability, then test-time compute scaling through feedback-driven refinement for the final push. GenCorrect is the same shape as Cliff's feedback loop, applied at inference instead of training.
Common pitfalls
- Don't use
scale_rewards=groupwhen your gate rejects a meaningful fraction of rollouts. One gate rejection shrinks every other advantage in the group. The watercolour tutorial switched to no scaling and the reward finally moved. - MoE architectures silently break LoRA module lists. The standard
target_modulesassumes dense-model naming. On Qwen3.5-35B-A3B, the hand list trained ten of forty layers. Useall-linearor audit the actual module names. - LLM judges actively prefer worse corrections. The feedback paper shows judges systematically favoring inferior baseline outputs when a feedback-informed fix succeeds. Blindly trusting an LLM-as-reward-model stack will push your policy the wrong way.
- Learning rate too low is the most expensive bug you'll chase. Flat reward curves sent one author through dozens of runs before the fix was 2e-5 to 5e-5. Start near the ceiling that LoRA-without-regret papers report, then tune down.
- Outcome rewards ignore the reasoning that produced them. Cliff's core observation, that post-first-mistake tokens carry no signal, applies broadly. Binary correctness rewards on long reasoning chains train the model on noise.
One thing to remember
Every successful run in this batch, from competitive programming champions to watercolour painters, came from making the reward signal better, not the model bigger. The cheapest win in LLM post-training right now is looking at your reward function before you add compute.
What this means
If you're training a reasoning model, adopt token-level credit assignment in the style of Cliff. A binary correctness reward is leaving 7% or more on the table, and the off-the-shelf judge you already use can locate first mistakes.
If you're compute-constrained, the 100-step GRPO recipe is your starting point. A 350M model with 500 samples and three well-specified reward functions on a free GPU gave a 7.1-point structural-compliance gain, while the gap to a 2B model was cut more than half on JSON.
One thing to watch: LLM-judge evaluation bias is a real threat to the whole reward-stack approach. The feedback paper proves the bias exists, but nobody has published a clean fix yet. Expect a calibration or ensembling method targeted directly at this problem within the next couple of quarters.