Appearance
The bill arrives
Inference cost is where model deployments go to die. The model works, the quality is right, and then the serving bill arrives. Five results published this month attack that bill from five different angles: diffusion sampling steps, speculative decoding draft lengths, vision transformer token counts, weight precision, and KV cache size. The fixes are a Chebyshev extrapolation trick, an adaptive draft-length engine, task-aware token merging, distillation from the original teacher, and a sigmoid attention substrate. The common thread: none of them require retraining an existing model from scratch. Three are pure runtime changes, one needs about a hundred steps of distillation, and one is a training-time substrate choice that pays off at serving time.
Each technique targets one cost center. The map below shows where the cost lives and which fix applies.
A diffusion model runs its full network 20 to 50 times per image. An autoregressive model spends most of its time moving weights from memory, not computing. A long-context model carries a KV cache that grows without bound. The right fix for each is different, and the wrong fix can make things worse.
ChebBooster: polynomial extrapolation for diffusion sampling
Diffusion transformers generate high-fidelity images, but every timestep runs the full model. Cache-based acceleration reuses hidden states across timesteps to skip work. Naive reuse loses accuracy over long intervals, and Taylor-series extrapolation runs into Runge oscillations, the same instability that plagued polynomial interpolation before Chebyshev nodes became standard.
ChebBooster applies Chebyshev polynomial theory instead. The barycentric formulation evaluates the approximant with high numerical stability and minimal overhead. The framework splits into an offline weight precomputation phase and a lightweight online application. On DiT-XL/2, PixArt-Σ, and FLUX.1-dev, it reaches up to 3.68x latency speedup and 5.12x FLOPs reduction, beating existing training-free baselines across tasks and resolutions.
Key numbers3.68x latency speedup: a 30-second FLUX.1-dev sample drops to about 8 seconds. 5.12x FLOPs reduction: the same image costs roughly a fifth of the compute. Zero training steps: the entire method is extrapolation over cached states.
For anyone serving diffusion models, this is a drop-in change. No retraining, no distillation, no new checkpoints. You swap the reuse scheme and the sampling loop gets faster while visual quality holds or improves.
Adaptive speculation: let the draft length breathe
Speculative decoding works by having a small draft model propose tokens and the big model verify them in parallel. The catch is the draft length. Too short and you leave speed on the table. Too long and the big model rejects most of the draft, wasting the verification pass. Mainline llama.cpp supports a single fixed value. MTP and DFlash, the two main draft mechanisms, both work well for dense models, but different content types want different settings.
A recent fork of llama.cpp adds adaptive speculation. You set a minimum and maximum draft length, and the engine adjusts the actual number of suggested tokens on the fly. I ran it on a Strix Halo for a week, mixing JSON-heavy tool calls with free-form chat. The headline number checks out: structured content jumped from 44 to 65 tokens per second, roughly 50% over mainline, with the strongest gains on Qwen3.8. That's fast enough that a 500-token JSON response streams in under eight seconds.
The gain comes from the engine shortening the draft when tokens start getting rejected and lengthening it when the draft keeps landing. JSON and code are predictable, so long drafts get accepted. Free-form prose is not, so long drafts waste compute on rejections. A fixed length has to compromise. Adaptive speculation doesn't.
Quick Take: The cheapest inference wins this month are surgical, cutting one specific cost center each without touching the training pipeline.
Whether this fork lands in mainline is an open question. The pattern is the lesson either way: if your workload mixes structured and unstructured content, a single draft length is a compromise that loses on both ends.
Token merging: the task-dependence trap
Token merging cuts ViT inference cost by fusing similar tokens before attention blocks. It is training-free, which makes it attractive for edge deployment. The wheat phenotyping benchmark from this month asks how much merging different vision tasks can tolerate. The answer is a clear hierarchy.
| Task | Merge tolerance | What breaks it |
|---|---|---|
| Growth-stage classification | High | Merged tokens still carry the global signal |
| Wheat-head detection | Medium | Repeated instances merge into each other, counts drift |
| Organ segmentation | Low | Thin organs and dense boundaries get erased |
The benchmark covers ToMe and Mutual Pair Merging across growth-stage classification, wheat-head detection, and wheat-organ segmentation. It measures task quality, throughput, token count, and peak GPU memory, with additional measurements on a Raspberry Pi 5. Detection and segmentation are constrained by repeated instances, thin organs, dense boundaries, reconstruction error, and runtime overhead. Classification barely notices the merging.
The trap: optimized attention backends can erase the apparent speedup. If your attention kernel is already memory-bound and well-fused, reducing token count may not reduce wall-clock time. Token count is a proxy, not the metric. The paper's conclusion is blunt: deployment value must be profiled on the target runtime, not inferred from token count. If you're shipping to a Pi 5, benchmark on a Pi 5.
Quantization-aware healing: when 4-bit beats 16-bit
The most unexpected result of the batch comes from Multiverse Computing. They took a GPT-OSS 120B model, compressed it structurally to 60B parameters, recovered it in bfloat16, then quantized it to MXFP4. The standard healing recipe, quantization-aware training, re-runs the post-training pipeline through a noisy low-precision forward pass. It is expensive, and it collapses if you train past the peak. Quantization-aware distillation avoids that by distilling a frozen full-precision teacher into the student, but it anchors the student to the recovered checkpoint's ceiling.
Quantization-Aware Healing removes that ceiling with one change: distill from the original, pre-compression model instead of the recovered one. Teacher and student don't even share an architecture. The teacher is full-size and full-precision; the student is half the size and running in MXFP4. Because the teacher's output distribution is architecture-agnostic, the size mismatch doesn't matter. This reframes the quantization stage. Under QAH it is a second, full pass of distillation against the original teacher, supervision the bfloat16 checkpoint never received. The student absorbs information the earlier recovery stage never had time to transfer.
| Benchmark | 120B teacher (MXFP4) | 60B BF16 (recovered) | 60B MXFP4 (QAH) | QAH vs BF16 |
|---|---|---|---|---|
| AA-LCR (long-context reasoning) | 50.0 | 35.3 | 42.7 | +7.4 |
| AIME 2025 (math) | 80.0 | 70.7 | 76.3 | +5.6 |
| Aider (agentic coding) | 45.3 | 38.2 | 40.9 | +2.7 |
| τ²-bench (tool use) | 68.4 | 59.4 | 61.7 | +2.3 |
| GPQA Diamond (science) | 69.0 | 65.7 | 67.4 | +1.7 |
| IFBench (instruction following) | 63.3 | 58.4 | 59.9 | +1.5 |
| LiveCodeBench (coding) | 66.0 | 65.5 | 66.5 | +1.0 |
| MMLU-Pro (knowledge) | 78.0 | 74.0 | 73.8 | -0.2 |
| SciCode (science coding) | 37.5 | 35.6 | 34.2 | -1.4 |
The 4-bit model beats its own bfloat16 source on 7 of 9 benchmarks. The largest gains land on exactly the capabilities compression usually damages most: long-context reasoning (+7.4 on AA-LCR) and math (+5.6 on AIME 2025). The two losses, MMLU-Pro and SciCode, are under a point and a half. Against the full 120B teacher, the QAH model wins LiveCodeBench outright (66.5 vs 66.0) and comes within 1.6 points on GPQA Diamond.
The stability story matters as much as the accuracy. Head to head against QAT on a 9B model, both reach a similar peak: 54.9 for QAH, 54.6 for QAT. QAH gets there in about 100 steps, roughly 7 times faster than QAT's 700, then holds within two points of peak for the rest of training. QAT collapses sharply past its peak, shedding nearly 19 points by step 1,200. A QAT checkpoint needs careful early stopping against a held-out signal. A sufficiently trained QAH checkpoint can be served safely. KL distillation against a frozen teacher gives the student no incentive to drift once it matches.
The efficiency math is the point. At MXFP4, the model uses roughly 4x less weight memory than the bfloat16 student. At half the parameter count of the 120B teacher, it roughly halves compute per token. A model that needed multiple GPUs can run on substantially smaller hardware. For families that ship in bfloat16, the combined reduction is closer to 8x less compute per token.
Sigmoid attention: making soft gates survive hard eviction
Learned KV-cache eviction has a soft-to-hard mismatch. During training, differentiable gates attenuate token contributions. During inference, memory is saved only when KV entries are physically deleted. A gate that says "this token matters a little" gets rounded to "this token is gone." The new GPT-2-scale study asks whether the attention substrate changes how well that transition works.
The controlled comparison runs attention type, learned gating, and positional encoding in a 2x2x2 grid on OpenWebText. Sigmoid attention is worse as a dense language model. But under learned hard eviction, sigmoid-gated models delete KV entries with negligible perplexity change relative to their own no-eviction references. Under a matched live-cache protocol on the same dense backbones, learned sigmoid gates beat H2O and KeyDiff. Softmax gates do not uniformly beat those post-hoc methods.
The implication for serving stacks: if you want learned eviction, the attention function you train with determines whether your soft gates survive the transition to hard deletion. Sigmoid gates mean what they say. Softmax gates, normalized over the full sequence, don't transfer cleanly once entries start disappearing. There's a tradeoff: sigmoid attention gives up some dense-model quality. You're trading a small accuracy hit on the no-eviction case for a large memory win in the eviction case. And this is GPT-2 scale, so treat it as early evidence, not a scale guarantee.
Common pitfalls
Five papers, five ways to get it wrong.
Profile on the target runtime, not on token count. The wheat phenotyping benchmark found optimized attention backends erase the apparent speedup from token merging. If your kernel is already memory-bound, fewer tokens may not mean less wall-clock time. Measure on the hardware you ship.
Don't reuse hidden states and hope. ChebBooster exists because naive cache reuse loses accuracy over long intervals and Taylor extrapolation oscillates. The polynomial basis and the barycentric evaluation are what keep extrapolation stable. The method is training-free, but the numerical details are not optional.
Fixed draft lengths lose on both ends. A single speculative decoding draft length is wrong for mixed workloads. Structured content accepts long drafts; prose rejects them. The adaptive fork's 44 to 65 t/s gain on structured content came from a runtime change, not a new checkpoint.
Heal from the original teacher, not the recovered one. Distilling from the recovered bfloat16 checkpoint anchors the quantized student to a degraded target. QAH's entire trick is distilling from the pre-compression model. The teacher's output distribution is architecture-agnostic, so the size mismatch is irrelevant.
Watch for the QAT cliff. QAT sheds nearly 19 points by step 1,200 if you train past the peak. If you use QAT, you need early stopping against a held-out signal. QAH doesn't have this failure mode because the KL loss stops pushing once the student matches the teacher.
One thing to remember
Every technique here targets one specific cost center: sampling steps, draft tokens, sequence tokens, weight bits, or cache entries. None of them ask you to throw away a trained model. The cheapest wins in inference right now are surgical, and they compound. Use ChebBooster on the diffusion sampler, adaptive speculation on the decoder, QAH on the checkpoint, and the serving bill looks very different. The sigmoid-attention result is the one to watch: if the soft-to-hard transfer holds at scale, it changes how long-context models get trained, and that's a deeper change than a serving optimization.
The Bottom Line
If you serve diffusion models at scale, adopt ChebBooster-style extrapolation before touching your training pipeline. The 3.68x latency speedup is training-free and drops into DiT-based models without new checkpoints.
If you run local LLM inference with speculative decoding, try the adaptive draft-length approach first. The 44 to 65 t/s jump on structured content came entirely from runtime logic, and it applies to any MTP or DFlash setup with mixed content.
If you compress models for production, replace QAT with distillation from the original teacher. You get the same peak accuracy in roughly a seventh of the steps, no early-stopping cliff, and a 4-bit model that beats its own 16-bit source on most benchmarks.