Appearance
The Messy Middle
Pretraining gets the headlines. Post-training decides whether the model actually behaves. These four new papers don't share a single experiment, but together they redraw how we adapt LLMs after the fact.
One gives the field a long-missing taxonomy for verbal feedback. Another settles a practical question with a refreshing answer: there's a wide near-optimal region for SFT-RL budget splits, and small models can map it for you. A third replaces the one-size-fits-all training recipe with a per-sample router. The fourth uses Fisher geometry to explain why benign fine-tuning destroys refusal behavior faster than anyone is comfortable with.
All four are about the same thing. Routing. Where feedback goes, which samples get updated, how budget flows, and how safety pathways break and can be restored.
| Paper | Core move | What it changes | The catch |
|---|---|---|---|
| Rise of Verbal RL | Organizes feedback by when it acts | Feedback channel becomes a design decision | Taxonomy is descriptive, not prescriptive |
| SFT-RL Budget Allocation | Replaces peak-seeking with a near-optimal region | Budget planning becomes transferable across model sizes | Region shifts with annotation cost asymmetry |
| Self-Routing | Routes samples by rollout correctness and confidence | One training recipe no longer fits all samples | Needs rollout state signals you may not log today |
| Safety Fragility | Explains collapse with low-rank Fisher geometry | Failure is a routing disruption, not gradient conflict | LoRA and ASAM protection fades at larger scales |
Verbal RL: Grounding, Deliberation, Learning
The Verbal Reinforcement Learning paper organizes a field that badly needed organizing. Natural language is the feedback channel, and the taxonomy splits it along one clean axis: when the feedback takes effect, and what it modifies.
Pillar one is language as grounding signal. The language defines the task itself: goals, states, reward structure. You aren't optimizing a model against an external objective; the words are the objective.
Pillar two is language as deliberative feedback. This operates at test time with no parameter updates. Chain-of-thought prompting is the degenerate case, but the pillar spans richer loops where a natural-language critic examines a candidate output before the final answer is emitted.
Pillar three is language as learning signal. Here feedback reshapes weights through training. Textual reward models and natural-language gradient methods land in this bucket.
The diagnostic question the paper forces you to ask: do we want this feedback to change actions now, or change future actions? Deliberative feedback is cheap to prototype. A prompt and a few API calls. Learning-signal feedback commits you to another training run. That difference drives the engineering more than the language model underneath ever does.
The SFT-RL Budget Question
Which gets the annotation budget, SFT or RL? Most teams pick by folklore: SFT first, RL when SFT stalls. This paper says stop optimizing and start tolerating.
The key idea is near-optimality. Instead of hunting for a single best SFT-RL ratio, the authors define the near-optimal region: every allocation that lands within a tolerance, they tested 2% to 10%, of peak performance. Two findings matter.
First, the region is wide even at small tolerances. There isn't a knife-edge optimum to find. There's a plateau. Second, the region widens with model scale and transfers across model families. That's the practical gift. A small proxy model can map the region, and that region transfers up to the large target, so you can skip the expensive exhaustive search over big models.
Cost asymmetry shifts things too. RL rollouts are not free. When you generate a rollout before every learning step, the effective RL data cost climbs, and the cheap-to-collect RL data changes where the near-optimal region sits. Counting that cost is how you keep your allocation inside the region you mapped.
Quick Take: The SFT-RL ratio is a plateau, not a peak. Find the plateau on a small model and stop stressing.
When Safety Routing Breaks
Benign fine-tuning destroys refusal behavior fast. The numbers from the paper: after around 100 benign examples, safety can collapse to high attack success rates while general utility degrades only mildly. That asymmetry is the puzzle.
Prior work blamed gradient conflict between the safety objective and the downstream task. This paper proposes a different mechanism. The safety Fisher information is low-rank. Alignment flattens the safety geometry but preserves a low-rank output-routing pathway. Fine-tuning selectively re-sharpens that pathway in output-side MLP modules. What breaks is the routing that sends inputs into refusal behavior, not the knowledge of what to refuse.
The routing view explains things gradient-conflict stories couldn't. Why do a few safety examples restore refusal quickly? Because the internal safety representations survived; the pathway was disrupted, not destroyed. Why do LoRA and ASAM help at first? They suppress output-side sharpness. But their protection fades as fine-tuning scale grows. LoRA is a delay, not a fix.
Key Numbers
- 100 benign examples: enough to collapse refusal to high attack success rates.
- Low-rank safety Fisher: alignment flattens safety geometry without erasing it.
- Output-side MLP sharpness: the selective re-sharpening that explains the collapse.
Self-Routing: One Recipe No Longer Fits All
A single training recipe for every sample is a bad idea. Self-Routing is the fix. The framework watches the model's own rollouts, reads two signals per sample, correctness and confidence, and routes accordingly.
Correct but low-confidence samples go to GRPO. Confident correct samples go to on-policy self-distillation. Uncertain failures get regularization. Stable, already-learned samples get skipped. No external teacher, no extra annotations, no additional sampling. The model's own rollout state decides the loss.
On math reasoning across Qwen3 and Qwen3.5 backbones, Self-Routing beats uniform GRPO, uniform OPSD, and fixed mixtures. The more interesting part is the dynamics. The routing distribution shifts over training, and the system stops spending updates on low-signal or already stable samples. That's the budget-allocation instinct from the SFT-RL paper, applied at the level of individual samples instead of data sources.
Field Notes from Post-Training
I've seen the safety collapse in my own work. We were fine-tuning a chat-aligned model for a spreadsheet task, maybe 400 examples total, and refusal dropped to nearly nothing within the first hundred. The model still answered questions fine. It answered everything.
The SFT-RL budget question has been background noise in every project I've shipped. We burned compute testing a 50/50 split, then 70/30, then 90/10, chasing a peak that moved every time the data changed. The near-optimal region framing should have been obvious. We were standing on a plateau the whole time.
My team also ran a math RL pipeline where a single GRPO loss hit every sample, including ones that had been correct and stable for three epochs. The update was pure noise. Per-sample routing kills exactly that waste, and it's not even expensive to implement once you log rollout confidence.
The thread connecting all four papers is that they move attention away from the algorithm and toward information flow. What feedback counts, where it enters, and how it destabilizes or reinforces what the model already knows.
Common Pitfalls
- Assume refusal survives benign fine-tuning. It doesn't. Roughly 100 ordinary task examples can push attack success rates into uncomfortable territory. Plan safety data into the same run, not a re-alignment cleanup later.
- Chase the exact optimal SFT-RL ratio. The peak is a fiction. What exists is a wide near-optimal region that transfers across model sizes. Map the plateau on a small proxy and pick a point inside it. You'll save a week of grid search.
- Apply one loss to every sample. Uniform GRPO over-updates confident correct samples and adds noise to low-signal ones. Skip the stable samples, regularize the uncertain failures. Respect the learning state of each sample.
- Count RL data as cheaper than it is. Rollouts cost inference compute. Once you include rollout generation in the budget, the near-optimal region shifts toward SFT. Forget that and your allocation lands outside the region you mapped.
- Trust LoRA to hold safety at scale. It delays the collapse; it doesn't prevent it. The protection weakens as fine-tuning budget grows.
One Thing to Remember
All four papers are about routing. Verbal RL routes language into deliberation or learning. The budget work routes annotations between SFT and RL. Self-Routing routes samples into different updates. Safety collapse is a broken output-routing pathway. The researcher mindset shifts from "which algorithm" to "where does information flow."
Sources
- The Rise of Verbal Reinforcement Learning. arXiv:2609.01597v1. http://arxiv.org/abs/2609.01597v1
- Scaling Near-Optimal SFT-RL Annotation Budget Allocation from Small to Large LLMs. arXiv:2609.01573v1. http://arxiv.org/abs/2609.01573v1
- From Rollouts to Recipes: Self-Contained Post-Training for LLMs. arXiv:2609.01422v1. http://arxiv.org/abs/2609.01422v1
- When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning. arXiv:2609.01455v1. http://arxiv.org/abs/2609.01455v1
The Bottom Line
If you're running post-training on a tight annotation budget, run a small proxy model to map the near-optimal SFT-RL region, then train the large model at a point inside it. Stop chasing the exact ratio.
If you're fine-tuning a safety-aligned model for a downstream task, assume refusal will collapse after a few hundred examples. Add safety examples to the same run, and treat LoRA or ASAM as an early-training delay, not a permanent guard.
If you're building an RL pipeline from scratch, skip the single-recipe approach. Route samples by rollout correctness and confidence, let the routing distribution shift over training, and count rollout generation cost before you call RL "annotation-free."