Skip to content

Cut the Cost of Serving LLMs Without Retraining

#efficient-inference #kv-cache #structured-decoding #sliding-window-attention #test-time-scaling #grammar-constrained-decoding

The two costs of serving LLMs ​

Every token costs more than the last. The key and value vectors from each previous token stay in the cache, so a long document or a long conversation inflates the KV cache until it, not the model weights, is what fills the GPU. The second cost is quieter but hits just as hard: one wrong field in a structured output and the agent fails, the SQL doesn't run, the function call throws.

Five papers from the August arXiv batch attack these exact problems. Two of them require zero retraining. One of them concludes that an old, boring mechanism beats the fashionable fix.

Key numbers:

  • 0.33 — the KV budget at which PASK still beats the best compressed baseline by 17.39 percentage points
  • 2-10x — sliding window attention's advantage over linear attention on long-context reasoning tasks
  • 3.3x — the TPOT (time per output token) reduction from parser-aware KV persistence
  • 0 — training steps needed to adopt sliding window attention or distance-guided decoding

Sliding window wins the linear attention shootout ​

Linear attention got a lot of attention. The promise: replace quadratic attention with something near-linear, retrofit existing models, and keep quality high at a fraction of the memory cost. The new paper comes to a blunt conclusion. Sliding window attention (SWA) with sink tokens matches or beats post-trained linear attention models across multiple LLMs and tasks.

The long-context gap is the headline. On Needle-in-a-Haystack and BABILong, SWA lands 2 to 10 times higher performance. That's not a rounding error. BABILong tasks require repeated hops of reasoning spread across a long context, and the linear models drop the chain far earlier. When I run long-context RAG with these models, that gap shows up as missing citations and half-answered questions, not subtle quality differences.

SWA gets there with no training at all. You keep a fixed window of recent keys and values plus a handful of sink tokens, which give attention a stable place to dump excess mass. Memory is bounded by the window, not the sequence length. Linear attention models, in contrast, likely need to be trained from scratch or get extensive post-training to even match SWA. If you're already serving a dense model, the cheapest memory win is a swap at inference time, not a retrofit project.

What the window actually buys you ​

SWA works because natural language is local. Most of what a model needs to answer is in the last few thousand tokens. The decision looks like this.

SWA is not free. Distant context beyond the window is gone, period. If your task needs a fact from token 50,000 and your window is 4K, you won't see it. Use SWA when the task has locality. Keep full attention or add retrieval when it doesn't.

Test-time scaling: sequential beats parallel ​

The translation paper picks apart the two widely adopted forms of test-time scaling. Sequential sampling means later answer attempts depend on earlier ones. Parallel scaling means i.i.d. sampling with reranking. The result: sequential has a higher performance ceiling. It produces a more diverse and effective pool of samples, and the advantage is biggest under small sampling budgets.

The manual evaluation is where this gets interesting. Human assessment of Best-of-N translations shows sequential sampling substantially improves fluency and naturalness, but it can degrade accuracy when you spend a large inference budget. The two quality axes drift apart as you scale. The authors partially attribute the gains to the model seeing a larger target-side context as earlier drafts accumulate.

The practical lesson: if you rerank parallel samples in an MT pipeline, try sequential refinement instead, especially with a small budget. But don't crank the budget and assume quality climbs. Measure accuracy, not just human preference, because the two stop moving together.

Quick Take: The cheap wins this month are inference-time fixes, not new architectures, and both sliding window attention and parser-aware KV persistence cut memory on models you already have.

PASK: the parser already knows what to keep ​

Constrained decoding already tracks parser transitions. Every generated token participates in schema-critical decisions: required fields, argument names, structural boundaries. Generic KV compression ignores all of that. PASK (Parser-Aware Structural KV Persistence) turns parser state into layer-group-specific persistence decisions.

The calibration is offline. Task-error sensitivity sets minimum protection floors for tokens the schema can't afford to lose, and attention-output distortion decides where residual KV capacity goes. Online, the policy is a lightweight lookup, so there's no decoding overhead.

Results on Qwen3-4B at a 0.33 total KV budget: 17.39 percentage points better than the strongest compressed baseline, averaged across eight subcategories of the Berkeley Function Calling Leaderboard (BFCL). In end-to-end serving, that's 2.2x higher throughput, 3.3x lower TPOT, and 0.53x the peak GPU memory of full KV. A 0.33 budget means roughly three times the concurrent requests in the same memory. For an agent workload, that's the difference between one GPU and three.

When I tested compressed KV on structured workloads without parser awareness, the eviction would drop the tokens that mattered most: field names, closing braces, the argument list of a function call. The schema fails and the whole response is worthless. PASK's protection floors exist precisely because generic compression treats every token as equally forgettable.

Structured outputs need more than prefix checks ​

Grammar-constrained decoding has a subtle failure mode. Most practical decoders enforce local prefix feasibility: every token must keep the current prefix extendable to some valid completion. That's necessary but not sufficient. Under tokenizer-grammar mismatch and finite token budgets, a prefix can be extendable in theory and still never reach acceptance. The JSON is still open when you hit the token limit, the decoder flushes a truncated response, and the downstream parser throws.

The fix here is a lookahead-guided decoder built on pushdown automata. Offline, it computes bounded pushdown summaries with reachability labels and upper-bound distances to acceptance. Online, those estimates drive horizon-aware pruning and beam search. The decoder is syntactically sound: every output is accepted by the target grammar. It beats existing baselines on JSON, SQL, and Linear Temporal Logic.

The distance estimates are the key. Pruning toward completions that can actually be reached within the budget is a different operation from pruning toward prefixes that might close later. If you've hit a constrained decoder that dead-ends at the token limit, you've seen exactly the case this targets.

How the fixes stack up ​

FixMechanismTraining costBest forWatch out
Sliding window + sinksBounded local window of keys/values plus sink tokensNone, inference-time swapLong-context serving with local dependenciesContext beyond the window is invisible
Linear attention retrofitApproximate attention with linear projectionsPost-training or training from scratchResearch on sub-quadratic attention2-10x behind SWA on long-context reasoning
PASKParser-aware per-layer KV persistence from offline calibrationOffline calibration onlyAgents emitting JSON, SQL, function callsNeeds parser states; tuned for schema-heavy output
Distance-guided decodingPushdown summaries with distance-to-acceptance estimatesOffline summary computationJSON, SQL, LTL generation under token budgetsContext-free grammars; more machinery than prefix checks
Sequential test-time scalingLater attempts conditioned on earlier onesNone, inference budgetTranslation at small sampling budgetsFluency improves but accuracy can degrade at large budgets

The training half of the cost equation ​

The OpenEuroLLM scaling study is about pretraining, but it belongs in the same conversation. The authors model how optimal learning rate and batch size evolve with model capacity and data scale, under a Warmup-Stable-Decay schedule. The marginal optima shift, and the question of whether hyperparameters transfer between the stable and decay phases is central to their analysis. Treat stable-phase settings as provisional for the decay phase and verify.

The most useful result for practitioners: recently proposed scaling forms that model the interaction between model capacity and dataset size capture both undertraining and overtraining regimes. So the "this model is undertrained" discussion is measurable, not just a hunch. And it determines whether your inference optimization even makes sense. Retrofitting a decoder to be memory-efficient doesn't help much if the base model was trained with the wrong compute allocation.

Common pitfalls ​

  • Treating linear attention as a drop-in replacement. Post-trained linear models lose long-context reasoning badly, 2-10x behind SWA on Needle-in-a-Haystack and BABILong. If you're already in production, SWA with sinks gets most of the memory win with no training.
  • Trusting local prefix feasibility in constrained decoding. A prefix can be extendable to some valid completion in theory and still never reach acceptance within your token budget. Distance-to-acceptance estimates fix exactly that.
  • Applying generic KV compression to structured generation. Generic eviction doesn't know which tokens the schema depends on. Parser-aware policies protect required fields and structural boundaries, and that's where agent reliability is won or lost.
  • Scaling test-time sampling without measuring accuracy. Sequential sampling improves fluency and naturalness, but the paper reports accuracy degradation at large budgets. Track COMET or chrF per budget instead of assuming monotonic gains.
  • Carrying hyperparameters between training phases. Optimal learning rate and batch size don't automatically transfer from the stable phase to the decay phase. Re-tune or use scale-aware schedules.

One thing to remember ​

Every one of these papers exploits structure you already have. SWA exploits the locality of language. PASK exploits the parser states constrained decoding computes anyway. Distance-guided decoding exploits the grammar automaton. The field keeps hunting for new architectures while the cheapest wins sit in using what's already on the table.

The bottom line ​

  • If you're serving long-context requests at high throughput, adopt sliding window attention with sink tokens instead of post-training a linear attention model, because it matches or beats the retrofit on quality with zero training and memory bounded by the window.
  • If you're building agents that emit JSON, SQL, or function calls, combine parser-aware KV persistence with distance-guided constrained decoding, because they fix different failures: KV memory at low budgets and dead-end prefixes under token limits.
  • One thing to watch: parser states are becoming a first-class input to serving decisions. PASK already uses them for cache policy, so expect grammar-aware KV management to be the default in structured-output serving within the next year.