Appearance
Every token an LLM serves carries a bill: the silicon time that produced it, the architecture choices baked into the model, and the API price a customer will tolerate before switching providers. This cluster of papers and industry posts attacks all three at once. A simplified FP4 pretraining recipe nets 21.2% more throughput. A streaming video framework cuts per-frame prefill latency by 52x. A pruning paper shows sparsity amplifies model bias, then fixes it. The thread connecting them is simple: the winner in this generation of AI infrastructure won't be the biggest model, it will be the one that produces the cheapest correct token.
Key numbers from this cluster
- 52.1x: per-frame prefill latency reduction from ShallowStream on streaming video
- 21.2%: token throughput gain from dropping the Hadamard transform and BF16 final-block exemption in FP4 pretraining
- 81.8%: exact match on GraphWalks Parents with Trace as State, up from 29.2% on the first pass
- 90%+: API cost reduction for a 500k-DAU customer service agent when moving from flagship to Flash-class models
Why every efficiency trick starts with memory bandwidth
The reason inference is expensive is boring: it's memory-bound. Generating one token forces hundreds of gigabytes of weights to move from HBM to compute units, and most of the time the silicon sits idle waiting for the bus. Compute is cheap, memory movement is not. So every serious optimization in this space does the same thing: move fewer bytes per token, or amortize the bytes you must move across more useful output.
Speculative decoding is the cleanest example. A small draft model guesses the next several tokens, then the big model verifies them all in one forward pass. Verification costs about the same as generating a single token, because the weights get loaded from memory either way. If the draft is right, you just produced multiple tokens for the price of one memory sweep. The speedup comes down to a single equation:
Speedup = E[L] × T_target / (T_draft + T_verify)Two levers push it: raise E[L], the expected number of accepted draft tokens, or shrink the denominator by making the draft faster and the verification cheaper.
The hardware constant that locks the draft length
NVIDIA's official inference guide, published as a model-hardware co-design spec, makes the second lever concrete. When the critical path is the attention kernel, the total tokens the verifier processes is 1 + D, where D is the draft length. GPU attention kernels group query and KV heads, and the kernel block size is bounded by the physical matrix tile, typically 128 on current architectures. That yields the constraint:
G × (1 + D) ≤ 128, so D = 128/G − 1With a GQA ratio of G=8, the optimal draft length is 15 tokens. With G=32, it's 3. Pick a draft length that doesn't fill the tile, and you pay for the whole tile anyway. Waste is waste, whether the tile is half-empty or fully used.
That constraint reverses how teams should think about speculative decoding. It used to be an engineering bolt-on applied after training. Now the draft length is a function of a number you choose at architecture design time, the query-to-KV head ratio. You don't tune the draft to the GPU, you design the model so the GPU tiles fill perfectly.
The official guide ranks the current methods, and the ranking says something about the field. EAGLE-3, the sequential drafting champion, gets demoted to a reference row because its acceptance length has been overtaken. MTP, where multi-token prediction is trained jointly with the main model, is called the best choice on GPU for models that can afford training-time coupling. DFlash brings diffuser-style parallel generation: seven draft tokens produced in one forward pass. DSpark stacks a light serial module on a parallel base, pushing acceptance length roughly 30% above EAGLE-3 and 18% above DFlash, and in DeepSeek-V4 production it runs 60% to 85% faster than MTP.
| Method | Drafting mechanism | What the guide says |
|---|---|---|
| EAGLE-3 | Sequential, token by token | Acceptance length now surpassed; reference position |
| MTP | Joint multi-token prediction during pretraining | Best choice on GPU for mainstream models |
| DFlash | Parallel, diffuser-style, 7 tokens per forward | Up to 6x lossless speedup |
| DSpark | Parallel base + light serial module | +30% acceptance vs EAGLE-3; 60-85% faster than MTP in production |
Reading the Chinese industry coverage of this guide, I wasn't the only one who noticed what the table implied about the research ecosystem. The middle four rows trace to specific teams, and the commentary around the release kept hammering on that point. The hardware vendor documented the constraints, but the recipes came from the research community. Nobody should underestimate how fast that exchange is moving. When I plugged our own GQA ratios into that constraint equation, the right draft length changed by 5x between two of our serving configs. We had been leaving a full tile's worth of throughput on the floor.
FP4 pretraining without the training wheels
Four-bit floating point has always had a representation problem. The E2M1 payload, one explicit mantissa bit and two exponent bits, covers such a narrow range of magnitudes that training collapses unless you intervene continuously. NVIDIA's Transformer Engine recipe solves it with current-tensor scaling, a randomized Hadamard transform (RHT), and bfloat16 for the final layers. It works, but it adds work outside the core FP4 matrix multiplications.
A new paper from the same batch argues most of that machinery is optional. Instead of E2M1 payloads with dynamic current-tensor scaling, pair them with unsigned E5M3 block scales. The wider exponent range of E5M3 permits periodic, rather than per-tensor, scaling. The recipe skips RHT entirely, uses FP4 in every eligible internal linear, and applies selective stochastic rounding only to backward gradients. The authors pretrained a Nemotron-H 8B model for nearly 190 billion tokens. That's a full industry-scale run, not a toy demonstration.
| Dimension | Transformer Engine | UE5M3 block-16 recipe |
|---|---|---|
| Scaling | Current-tensor, dynamic | Periodic, E5M3 unsigned block scales |
| Hadamard transform | Required (RHT) | Omitted |
| Final layers | BF16 exemption | FP4 in all eligible linears |
| Gradient rounding | Standard | Selective stochastic rounding on backward pass |
| Model-body throughput | Baseline | +21.2% native execution ablation |
The headline result is the ablation: jointly removing RHT and the BF16 final-block exemption increases measured model-body token throughput by 21.2%. Nearly a fifth more tokens per second for the same silicon, because you cut an expensive transform and keep the whole model at one precision. The final-window training loss is lower than the Transformer Engine baseline, and under each recipe's respective quantized-inference policy, held-out negative log-likelihood is also lower. Simpler recipe, better loss, faster throughput.
Quick Take: If you copied the Transformer Engine FP4 recipe wholesale, this ablation says you're paying about a fifth of your model-body throughput for insurance you may not need. E5M3 block scales get you stable FP4 pretraining with fewer moving parts.
Fewer tokens: streaming video and sparse state
The other way to cut bytes per token is to stop processing tokens you don't need. ShallowStream tackles streaming video, where a multimodal LLM re-runs full-depth prefill over every incoming frame. That's brutally expensive, and the KV cache grows in proportion to prefill depth. The insight is that you don't need the deep layers to know which frames matter. Shallow layers can encode frames and build a retrieval index at the same time, since they run cheaply and their attention scores carry enough signal to rank context.
During stream processing, ShallowStream keeps an always-on lightweight index built from shallow-layer KV cache. At query time, it scores candidate frames using shallow-layer attention scores, then applies a diversity-aware selection pass to pull the frames that actually help. The result matches the strongest existing streaming methods while cutting per-frame prefill latency by up to 52.1x and 10-second end-to-end latency by up to 11.9x. In practical terms, a 10-second response becomes under a second, and you can run that on hardware that previously couldn't keep up with a live feed.
Graph Machine takes the same philosophy into pretraining. Instead of dense layers that attend over everything, it uses sparse layers with O(n)-sized state and dynamic, pointer-like routing it calls edges. Pointer chasing sounds exotic, but the mechanism is straightforward: edges are differentiable objects that decide which entries of the state to retrieve, so the accessible state stays large while the compute stays small. The authors replaced 75% of the dense Transformer layers in Qwen3-0.6B and pretrained from scratch on 15.7B tokens. With only 2 of 4,096 tokens retrieved per KV head in each sparse layer, loss degrades only slightly; with 4, the best config marginally improves on dense. That's 0.05% of the attention work of a dense layer while keeping the bulk of its capacity.
At the extreme edge, GaLe partitions feature maps instead of tokens. A local exact representation preserves fine details; a global approximate representation keeps long-range dependencies. That split supports global operations and attention, which standard tiling can't, and it needs no retraining. On a Cortex-M33-class microcontroller, GaLe matches exact-inference ImageNet accuracy while delivering up to 65% speedup and 90% RAM reduction over patch-based inference. A microcontroller running pretrained attention models at full accuracy is a capability that mostly didn't exist last year.
Long-context: give the reader a map
There's an efficiency win hiding in how we order context. Causal transformers process information in sequence, but long-context reasoning often depends on state discovered only later in the context. Trace as State formalizes this mismatch with conditional state update tasks: for a causal processor, receiving the condition first can require exponentially less memory than receiving it last.
The practical move is almost embarrassingly simple. Take a reasoning trace, use it as a textual proxy for task state, and place it before the long-context block on a second pass. The model rereads the context with a map of where things are. The matched control, Trace Append, puts the same trace after the context. Across three models and three long-context datasets, Trace as State wins 26 of 27 model-task-metric combinations. On GraphWalks Parents, exact match lifts DeepSeek V4 Pro Preview from 29.2% on the initial pass and 43.0% with Trace Append to 81.8%. GLM-5.2 goes from 66.4% and 83.2% to a perfect 100.0%.
No training, no extra parameters. You're just changing the order of blocks on a reread pass, which means the efficiency gain is free at inference time. If your retrieval pipeline appends evidence after the prompt, this paper suggests you have the order backwards.
A related paper in the same batch fixes a training-side instability that shows up when you move away from the single residual stream. Orthogonal Hyper-Connections restricts the residual mixing matrix to the rotation group SO(n). Because a rotation can neither amplify nor attenuate the residual streams, training stays stable without forcing the streams to converge and lose their diversity. At four streams, the group is parameterized in closed form by two unit quaternions: zero extra parameters, faster construction than the iterative projection used by the doubly-stochastic variant. It's a reminder that the efficiency problem extends backward into pretraining, not just forward into serving.
Pruning has a bias problem
SparseGPT is the standard post-training pruning workhorse, but it has a side effect most papers don't measure: it amplifies existing model bias. Outputs vary more strongly with persona cues in the prompt after sparsification. If you prune a model and ship it, you can make its social biases worse without changing its perplexity.
Debias-SparseGPT fixes that with a representational debiasing term in the pruning objective, a second-order term defined over demographically contrasting inputs. Across 25%, 50%, and structured 2:4 sparsity, it consistently reduces pruning-induced bias while preserving perplexity and zero-shot accuracy. The most striking result is under 2:4 structured sparsity, the pattern that aggressively degrades model quality. There, augmenting the calibration set with long-context, content-rich examples improves both downstream performance and fairness. The bias-performance trade-off isn't a law, it's a consequence of how you build the pruning objective.
Two implications matter for practitioners. First, validate pruned models on demographic groups, not just benchmark averages. Second, when sparsity is aggressive, the calibration data is doing more work than you think. Treat it as a dataset design problem.
Flash models put the pricing on the table
The academic papers in this cluster are about physics and math. The industry reports are about the outcome: Flash-class models becoming the economic default. Google shipped Gemini 3.8 Flash in early September. Zhipu's GLM-5.3-Flash, revealed in late August as the mysterious Ox Alpha, packs 320 billion total parameters but activates only 18 billion per token. Alibaba cut Qwen3.8-Flash pricing to 0.8 yuan per million input tokens and 2.7 yuan per million output tokens. These are not small models. They're sparsely activated models with the knowledge capacity of a giant and the serving cost of a dwarf.
The architectural choices mirror the arxiv work. GLM-5.3-Flash cut layers from 92 to 45, dropped attention compute by 3.01x and KV cache memory by 4.44x using a hybrid of sparse and linear attention, and ran its 100 trillion tokens per day of blind-test traffic on a cluster of over 100,000 domestic AI accelerators. The key-value cache reduction matters more than the flop reduction, because KV cache pressure is the silent killer of long-context serving.
The unit economics are the real story. When I modeled a 500k-DAU customer service agent, the math was stark. At 500 million tokens per day and roughly 30 yuan per million tokens on a flagship model, the API bill hits 15,000 yuan per day. Switching to a Flash-class model at under 2 yuan per million tokens drops it below 1,000 yuan per day. A 90%+ cost reduction flips an application from loss-making to profitable, which is why the industry commentary around this shift keeps using the phrase "commercial settlement." MiniMax's first-half 2026 results back that up: B-side revenue of $73.9 million, up 703.1%, now over 60% of total revenue. Enterprise workloads, not consumer chat, are paying the bills.
This is the spectrum that matters now: flagship models as the slow-thinking layer, called only for complex reasoning, long-chain logic, and research-grade problems; Flash models as the fast-thinking layer absorbing 90% of tokens for routing, extraction, code completion, and high-frequency agent tasks. Around 85% of enterprise tasks don't need frontier intelligence, but they are hypersensitive to latency and price.
Common Pitfalls
Copying the Transformer Engine FP4 recipe wholesale. The ablation says you're paying 21.2% model-body throughput for RHT and the BF16 final-block exemption. If you can't change hardware, at least test the E5M3 block-scale path. It's simpler and finished with lower loss at 190B tokens.
Picking draft length without checking the tile math. The constraint D = 128/G − 1 is not a suggestion. With G=32, don't configure a draft length of 15; you'll pay for the full tile regardless. Match the draft to your GQA ratio or redesign the ratio.
Pruning and only checking perplexity. SparseGPT amplifies persona-dependent output divergence, especially at 50% and 2:4 sparsity. Measure output variance across demographic prompt cues before you ship a pruned model. If 2:4 is required, rebuild the calibration set with long-context, content-rich examples.
Running full-depth prefill on every streaming frame. ShallowStream's 52.1x prefill speedup comes from using shallow layers as the index and retrieving selectively. Naive token pruning on top of full-depth prefill still pays the cost you're trying to avoid.
**Appending retrieved evidence after the