Skip to content

The Post-Training Efficiency Stack: Fewer Tokens, Faster RL, and Agents That Remember

#post-training #reasoning-models #rl #lora #inference-optimization #coding-agents

The overthinking tax ​

Every reasoning request carries a hidden line item: the tokens the model burns thinking before it answers. On a hard GPQA question, Qwen 3.8 27B at maximum effort streams out roughly 6,600 tokens of internal monologue. The actual answer is maybe 200 tokens. You pay for all of it.

That's not a rounding error. Serving a 27B reasoning model, thinking tokens are most of your GPU bill, because the model produces them on every request, from graduate-level chemistry to "why is my build failing."

Over the past few months, three separate projects attacked this waste at different levels:

  • UkisAI post-trained Qwen 3.8 27B to emit up to 58% fewer thinking tokens while keeping top-effort accuracy within 1%.
  • TRL's AsyncGRPOTrainer now syncs a few-megabyte LoRA adapter instead of a three-gigabyte model, cutting a 500-step RL run from 3 hours 27 minutes to 53 minutes.
  • Randal Schwartz's Synthetic Scars project gives coding agents institutional memory, so the same production failure stops recurring across fresh conversation windows.

They look unrelated. They're not. Each one removes tokens, or repeat work, that doesn't contribute to the outcome. That's the whole game.

Teaching a model to think less without thinking worse ​

The UkisAI team started from an annoyance: their quantized Qwen 3.8 27B instances kept falling into reasoning loops, the "overthinking errors" that show up as long, anxious repetition inside the thinking trace. The loops persisted even at medium and low reasoning effort. A Meta paper proposed a fix for this in post-training-quantized models, but out of the box the results were mixed.

So they did the empirical thing. They generated large volumes of out-of-distribution traces across coding, language, vision, and agentic tasks on an 8xH100 box, isolated the traces with overthinking, and looked for common-denominator tokens between them. Then they built an inference-time penalizer for those tokens.

Two things surprised them. The penalizer worked better than the token set from the Meta paper, and it worked on bf16 models, not just quantized ones. They kept going from there. They built a loss function around those tokens, ran LoRA SFT over the traces, and reasoning fell off substantially. Good. But accuracy fell with it. Bad.

The fix took a lot of tinkering, literally from the day the base model shipped. They tried RL with GSPO, on-policy distillation, and adapter chunks from ThinkingCap 3.6, until accuracy came back to within 1% on almost all in-house tests. The training was deliberately indirect. They never targeted reasoning length. They targeted the tokens associated with overthinking and let the model find its own shorter path to the answer.

The results, benchmarked five runs per model on each task at xhigh effort:

BenchmarkQwen 3.8 27BSwift-27BMedian tokens saved
GPQA-Diamond88.4%88.3%58%
LiveCodeBench v676.8%81.6%46%
Terminal-Bench 2.166.7%65.8%39%
MMLU-Pro85.5%85.0%28%
C-Eval90.0%90.6%19%
IFBench73.5%71.8%51%
AIME 202698.7%94.0%50%
HMMT (Nov 2025)99.3%96.0%46%
ERQA (vision)67.5%66.3%55%

The token savings are the headline. A 58% reduction on GPQA means a question that took 6,642 thinking tokens now takes 2,771. For a token-bound workload, that's roughly 1.95x faster responses and close to half the compute bill. Since you pay per token on rented GPUs, this is a direct margin play.

Two caveats from the table. LiveCodeBench went up 4.8 points, which the team attributes to the harness's default truncation rather than a real gain. And AIME 2026 dropped 4.6 points, which they traced to a single token in their penalizer that math reasoning apparently depends on. They're fixing it in the next release. If you serve math-heavy workloads, that's the number to watch.

The team positions this as complementary to reasoning-effort settings, not a replacement. Base-medium on GPQA scores 84.1% at 1,753 median tokens. Swift at xhigh scores 88.3% at 2,771 tokens. You keep top-effort accuracy at a little over one and a half times the token budget of a worse answer. Their thesis is worth taking seriously: reasoning length is important, and shortening it by force is the wrong move. You want to remove only the unnecessary part of thinking.

Quick Take: Swift attacks cost per request. Async GRPO attacks cost per training run. Both work by removing work that doesn't change the outcome.

The RL loop was the next bottleneck ​

Once the model is cheaper to serve, the next lever is the training loop that improves the model. RL post-training needs the trainer and the rollout generator to talk constantly: generate rollouts, compute advantages, update the policy, ship the new weights back to inference.

In a conventional setup, that's NCCL inside a dense cluster. Every optimizer step moves gigabytes. TRL's AsyncGRPOTrainer already separated training and generation, but it assumed both sides shared a node or a filesystem. On Hugging Face Jobs, each Job is one container on one VM. There's no shared localhost, no cross-node network path, no NCCL group. Full-weight sync is a non-starter.

The trick is LoRA. A rank-1 adapter for a 1.5B model is a few megabytes. The full model is around 3 GB. There's also a statistical argument, from Thinking Machines' LoRA Without Regret, for why rank 1 is enough in RL: the advantage function gives roughly O(1) bits of information per episode. There isn't much to learn from any single step, so a rank-1 adapter has the capacity to absorb it.

A bucket, a proxy, and no NCCL ​

The resulting architecture has three moving parts: a trainer Job, two vLLM Jobs, and a Storage Bucket mounted at the same path inside every container. The trainer saves the adapter with an atomic rename, then POSTs its path to vLLM's runtime adapter endpoint. vLLM loads the files from disk. No tensors cross the network.

Key numbers:

  • Adapter size: a few MB, versus 3 GB for the full model, so sync travels through a mounted FUSE bucket instead of NCCL
  • Runtime impact: a 500-step run dropped from 3 h 27 min to 53 min, about 4x faster
  • Staleness math: with max_staleness=4, vLLM needs 6 adapter slots. Five silently evicts a policy that still has rollouts in flight
  • Preemption safety: checkpoints and the final adapter persist to the bucket, so a killed trainer resumes instead of restarting

The proxy earns its keep. Exposed Job ports require a Bearer token on every request, so the proxy adds the auth header. It broadcasts every state-changing call, adapter loads included, to all replicas. And it routes each rollout to the replica most likely to hold its KV prefix. That last piece is the subtle one.

Routing rollouts by KV prefix ​

Generation has two phases with very different profiles. Prefill processes the whole prompt at once and computes attention keys and values for every token. Decode then produces one token at a time, attending to everything before it. Because attention is causal, the KV cache of a token depends only on its prefix. Two requests that share a prefix can share KV blocks, and a replica that already has them skips the prefill entirely.

GRPO sends G=8 rollout requests with the same prompt. If all eight land on the same replica, the first request prefills and the other seven reuse the cache. Round-robin would split them, forcing half the requests to redo prefill work on a replica that hasn't seen the prompt. That's wasted GPU compute on every step.

The router tracks which replica has seen which block hash. vLLM stores prefix KV in 16-token blocks, and the hashes are chained: block 3's hash represents blocks 1, 2, and 3. The chain is seeded with the adapter name, because KV computed under policy v3 is useless for policy v4.

That seeding detail matters more than it looks. Publish new weights under the same adapter name, and vLLM's prefix cache will happily match KV blocks from the old policy. A rollout can then get its prefill from version 3 and its decode from version 4, and the trainer has no way to detect it. Versioned adapter names make the collision impossible: a name always means one set of weights.

Another trap hides in data-parallel mode. TRL refuses adapter-only sync when --data-parallel-size > 1, because /v1/load_lora_adapter only reaches the DP rank that answers it. The other ranks silently serve the base model under the new policy name. On Jobs, data parallelism lives at the replica level instead, and the proxy is what fans the adapter load out.

Scars, not rulebooks ​

The third project attacks a different kind of waste: repeated failure. It doesn't touch model weights at all. The "training" happens after training, and it persists in the repository rather than the parameters.

If you've run an AI coding agent against a real codebase, you know the pattern. The model writes clean, idiomatic code that passes the toy tests. Deploy it at 3 AM and it forgets what happens when a user double-clicks a button, leaks a stream subscription until the app dies, or assumes every network request returns in 50 milliseconds. You correct the mistake on Monday. It apologizes. Thursday, in a different file, it does the exact same thing.

Schwartz calls this the Straight-A Intern Paradox, and the explanation is statistical. Models train on the happy path, the sunny-day code from tutorials and homework repos that dominates the internet. Production engineering is mostly the sad path, and the sad path is rare in training data.

The obvious fix, a rulebook full of warnings, fails twice. First, the pink elephant trap: flood an LLM's context with "Don't do X, don't touch Y," and the tokens for X and Y dominate its attention, so the model fixates on the forbidden pattern. Second, statistical amnesia: every fresh conversation window starts with total amnesia. The three-hour debugging session from yesterday never happened. The agent lives in a permanent Groundhog Day.

The counter to statistical amnesia is to write the memory down. Synthetic Scars codify lessons as a three-part structure:

  1. The Wound. The exact production disaster: the crash trace, the memory leak, the corrupted image.
  2. The Trap. The tempting textbook pattern the model reaches for because it looks clean on the surface.
  3. The Permanent Reflex. The non-negotiable invariant that must hold before code is written or merged.

The workflow adds a mandatory post-mortem after every task. The agent replays its trajectory, finds where it stumbled, and if the failure mode is new, distills it into the three-part format and commits it to the project's memory repository. The next agent, even in a fresh conversation, loads the scar catalog during planning. The organization stops having Groundhog Day.

The field results, across 51 consecutive production tickets in two codebases (a Flutter state framework and an enterprise geotechnical telemetry monorepo):

MetricTraditional AI codingScar-equipped agent
Repeat failure rate~40-50%0.0%
First-pass success on complex tickets~24%52.9%
Institutional memory0 scars retained185 codified scars

The 0% repeat failure rate is the number that matters. Once a mistake was codified, across 51 tickets the same mistake never recurred. The internal traces back it up: across 14,600 thinking turns, the team found 78 instances where the agent started drafting the naive shortcut, collided with a scar in context, and stopped itself mid-thought to write the guarded version instead. That's the hot-stove recoil, mechanized.

I train a lot of developers on AI, and the pattern here is painfully familiar. The senior ones are the hardest to convince. "Prompt and pray" feels great for a month, then the codebase quietly turns to slop. The model follows your instructions, doesn't see the big picture, and doesn't care about tech debt. Worse, that debt becomes part of its predictions for what code should look like, so the problem compounds. Paying it down always loses to new features, because nobody is getting paged.

The community response to Swift showed the same hunger on the serving side. Within days there were GGUF quants from Q1 to Q8, Bartowski's builds, NVFP4 and W4A16 versions, even an uncensored variant. The license isn't Apache 2.0, but it only affects companies over $1M in revenue, which keeps it usable for most hobby and research work.

Common Pitfalls ​

Penalizing reasoning length directly. UkisAI's whole thesis is that this is the wrong lever. Trim tokens by force and accuracy follows them down. Their method targeted specific tokens correlated with overthinking, then restored accuracy with on-policy distillation. A naive token cap, or a length reward in RL, produces models that answer faster and worse.

Suppressing the wrong token. The AIME regression is the cautionary tale. A token that looked like overthinking on coding tasks turned out to be load-bearing for math reasoning, and the model lost 4.6 points. If you build a token penalizer, audit it per domain. The same token set won't transfer from code to math to agentic tasks.

Reusing one adapter name in vLLM. Publish new weights under the same name and vLLM keys its prefix cache by that name, so KV blocks from the old policy match after the swap. You get prefill from one policy and decode from the next, and the trainer sees ratio drift it can't explain. Versioned names fix it.

Setting max-loras too low. With max_staleness=4, you need the current policy plus four lagging versions plus one extra slot for the swap. That's 6, not 5. With 5, vLLM silently evicts a policy that still has rollouts in flight at every sync.

Writing "Don't do X" rules for agents. The pink elephant effect is real, and so is the amnesia. Every new conversation resets the lesson. If you want an agent to stop touching the hot stove, codify the burn as a reflex: the exact failure, the tempting pattern, and the guard. Rulebooks are where behavioral fixes go to die.

One thing to remember ​

The cheapest token is the one the model never generates. But the winning pattern across all three of these projects is identical: measure the waste, isolate it, cut it, and verify the outcome. Swift removed overthinking tokens, not reasoning. Async GRPO moved megabytes instead of gigabytes, and kept the numbers honest with versioned names and staleness accounting. Synthetic Scars encoded the exact failure, not a general platitude. None of them shipped a policy change and hoped for the best. That's the skill.

The Bottom Line ​

  • If you're serving a reasoning model in production and token spend is eating your margin, evaluate thinking-token pruning post-training like Swift