Skip to content

Faster Diffusion: Timesteps, Tokens, Caches, and Parallelism

#diffusion-models #inference-optimization #sampling-schedules #feature-caching #parallel-sampling #diffusion-transformers

The bottleneck is the loop itself ​

A diffusion model generates an image by denoising a random latent over many steps. Each step is a full forward pass through a large network. For a 12B-parameter model like FLUX, that's 30 to 50 passes through a network that barely fits on one high-end GPU. You can't batch your way out of it, because the steps are sequential. You can't easily shrink the model, because quality drops. So the field has settled on a different strategy: make each step cheaper, or make fewer steps count for more.

Four papers from this month's arXiv batch attack exactly that. They're worth reading together, because they hit four different parts of the sampling loop, and they're explicitly designed to stack.

Four attack surfaces, one loop ​

Before diving in, here's where each method lands on the sampling loop. This mental model matters more than any single benchmark.

OYS decides which timesteps to visit. AViTS decides which tokens deserve high resolution. LinCa decides which features to recompute. PiX-MC decides how to run many steps at once. Four different questions, four different answers.

OYS: timesteps as a black-box optimization problem ​

Most sampling work optimizes the solver, not the schedule. You pick a solver like Euler or DPM-Solver++, then distribute your steps evenly, or with a hand-tuned schedule, across the denoising trajectory. The implicit assumption is that the schedule matters less than the solver.

OYS (Optimizing Your Sampling) throws that assumption out. It treats timestep selection as a black-box optimization problem and optimizes the quality metric directly with Bayesian optimization. No surrogate, no theoretical proxy for sample quality. It optimizes the thing you actually care about.

The results are hard to argue with. A 5-step OYS schedule retains 89-94% of the quality of a 50-step schedule. That's a 10x inference cost cut while keeping nearly all of the output quality. It beats both default schedules and Align Your Steps schedules on text-to-image, and it works on inpainting and other image tasks too. It also improves both simple and sophisticated samplers, so you don't need to switch solvers to benefit.

Two details make this practical. First, it requires no training, so it works on distilled models, which already have fragile schedules. Second, the Bayesian optimization is a one-time cost per model and task. You pay it once, then ship the schedule.

The catch: you need a quality metric you can evaluate, which means a reference set of prompts or images. That's a reasonable ask for a production pipeline, less so for a one-off generation script.

AViTS: don't upsample everything at once ​

Dynamic-resolution sampling is a neat trick: denoise at low resolution for the early steps, where the signal is mostly coarse structure, then upsample for the fine detail. The problem is that most implementations upsample all latent tokens at once when they switch resolutions. That's wasteful. Some tokens are already fine; others still need work.

AViTS (Adaptive Spatiotemporal Token Selection) makes the upsampling selective. It scores each token two ways: spatial importance from latent-text attention, which captures whether the token matters for the prompt, and temporal importance from how much the token's features change across diffusion steps, which captures whether the token is still evolving. It fuses the two scores and upsamples the important tokens first, deferring the rest.

The numbers: up to 6.34x FLOPs reduction on FLUX, nearly 9x on Qwen-Image-Edit and FLUX.1-Kontext-dev. FLUX is a 12B DiT, so a 6x FLOPs cut is the difference between a GPU that's saturated and one that has headroom. Combined with a distilled model, AViTS reaches 14.76x. And it's orthogonal to distillation, quantization, and feature caching, which is the polite way of saying you can stack it with all of them. Code is on GitHub if you want to reproduce the numbers.

One thing I appreciate about AViTS is that it doesn't require training. It reads attention maps and feature differences that are already being computed. The scoring cost is cheap relative to a full denoising pass.

Quick Take: These four methods don't compete with each other. They attack four different parts of the sampling loop, and the interesting gains come from stacking them.

LinCa: cache features, but learn how ​

Feature caching is the most intuitive acceleration: consecutive denoising steps produce similar intermediate features, so why recompute them? Cache the features from step t, reuse or lightly update them at step t+1. Training-free caching methods do this with uniform prediction strategies. They work at low acceleration ratios and degrade hard at high ones, because features don't all evolve at the same rate. I've seen this failure mode firsthand: push a training-free cache past 4x and the fast-changing features go stale.

LinCa's insight is that different feature components have different continuity properties. Some change slowly across steps; some change fast. A single prediction strategy can't handle both. LinCa decomposes cached features into sub-components using a lightweight invertible network, then applies a differentiated prediction order matched to how each component evolves. The invertibility guarantees lossless reconstruction back to the original feature space.

The pipeline is Decompose-Predict-Reconstruct, and it works across models: FLUX, Qwen-Image, HunyuanVideo. With less than 0.2% additional parameters, it maintains near-lossless quality at 5-7x speedup. That parameter count is the part that matters for deployment. You're adding a tiny learned predictor, not a second model. Code is on GitHub.

The tradeoff: LinCa requires training separate predictors for different models and timestep segments. It's cheap training, but it's not zero. If you're switching models weekly, that's friction. If you're serving one model at scale, it's nothing.

PiX-MC: parallelize across time, not just batch ​

The first three methods make the sequential loop cheaper. PiX-MC (Picard Proximal Monte Carlo) does something different: it makes the loop parallel.

This one targets Bayesian imaging inverse problems, not text-to-image. Think CT reconstruction, super-resolution, inpainting with a score-based prior. These problems need samples from a posterior, and the standard approach is sequential Langevin dynamics. Sequential means slow, especially for 3D volumes.

PiX-MC reformulates the sampling as a fixed-point problem and uses Picard iteration, which exposes parallelism across discretization nodes. The proximal-likelihood formulation exploits the fact that many imaging likelihoods have efficient proximal operators. The result is a sampler that runs across multiple GPUs, with multi-block and annealed variants for harder posteriors.

On a 512×512×80 sparse-view CT problem, the annealed multi-block version achieves up to 50x wall-clock speedup over a standard Langevin sampler using eight GPUs. That's the difference between a reconstruction that takes minutes and one that takes seconds. For clinical or industrial imaging, that changes what's feasible at the point of care.

The paper also establishes convergence guarantees under transparent assumptions, covering non-log-concave posteriors and imperfect learned score models. That's rare in this area, and it matters if you're using this in a regulated setting.

How the four compare ​

Here's the honest comparison. Note that the acceleration metrics measure different things, which is itself the point.

MethodAttack surfaceHeadline accelerationOverheadTraining neededBest for
OYSTimestep schedule10x cost cut (5 vs 50 steps), 89-94% quality retainedOne-time BO per model/taskNoText-to-image, distilled models
AViTSToken/resolution selection6.34x FLOPs on FLUX, ~9x on Qwen-Image-Edit, 14.76x with distillationAttention scoring at transitionsNoHigh-res DiTs
LinCaFeature caching5-7x near-lossless<0.2% paramsYes (lightweight predictors)Production serving of one model
PiX-MCTime-parallel sampling50x wall-clock on 8 GPUs (CT)Multi-GPU infraNoImaging inverse problems

The acceleration numbers aren't directly comparable. OYS cuts step count, AViTS cuts FLOPs, LinCa cuts wall-clock per step, PiX-MC cuts wall-clock via parallelism. If you're comparing them for your own workload, measure end-to-end latency and quality on your data, not the headline multipliers.

Stacking the speedups ​

The reason to read these four together is that they compose. AViTS explicitly claims orthogonality to distillation, quantization, and feature caching. LinCa is a feature caching method, so it slots in alongside AViTS. OYS optimizes the schedule, which is independent of what happens inside each step. PiX-MC is a different sampling approach entirely, so it's the odd one out, but for imaging problems it replaces the sequential loop rather than optimizing it.

A realistic stacked scenario for a text-to-image DiT: OYS gets you to 5 steps, AViTS cuts the high-resolution FLOPs by 6-9x, LinCa caches features across those 5 steps. That's a 10x step reduction times a 6x FLOPs reduction, minus the overhead of scoring and prediction. The papers don't report the combined number, but the individual results suggest the ceiling is much higher than any single method.

Key Numbers10x: inference cost cut from a 5-step OYS schedule, retaining 89-94% of 50-step quality. 14.76x: AViTS combined with a distilled model on Qwen-Image-Edit. <0.2%: additional parameters LinCa adds for its learned feature predictors. 50x: wall-clock speedup for PiX-MC on an 8-GPU CT reconstruction problem.

Common Pitfalls ​

Five things trip people up when they try to accelerate diffusion sampling.

  1. Uniform schedules waste steps. The default evenly-spaced schedule is rarely optimal. Early steps and late steps do different work, and a 5-step uniform schedule smears quality across steps that don't need equal attention. OYS exists because this is a measurable, fixable problem.

  2. Upsampling everything at resolution transitions. When you switch from low to high resolution, the naive move is to upsample the whole latent. You pay high-resolution FLOPs for tokens that are already converged. AViTS's scoring is cheap, and it's what turns a uniform high-res pass into one that only touches the tokens that need it.

  3. Using one prediction strategy for all cached features. Training-free feature caching falls apart at high acceleration ratios because some feature components evolve slowly and others fast. A uniform predictor either over-updates, wasting compute, or under-updates, losing quality. LinCa's decomposition exists for this reason.

  4. Benchmarking FLOPs instead of wall-clock. FLOPs reductions look great on paper, but feature caching and token selection methods move memory traffic, and memory bandwidth is often the real bottleneck on modern GPUs. A 6x FLOPs reduction can be a 2x latency reduction if you're bandwidth-bound. My team has hit this more than once. Measure end-to-end latency.

  5. Assuming the schedule optimizer transfers across models. OYS's Bayesian optimization is per model and per task. A schedule that's optimal for FLUX text-to-image won't be optimal for FLUX inpainting, and it definitely won't transfer to a distilled variant. Re-run the optimization when your model or task changes.

One Thing to Remember ​

The common thread across all four papers is that they treat the sampling loop as something to be optimized, not as a fixed procedure. The days of running 50 steps with a uniform schedule and calling it done are over. The cheapest wins require no retraining: OYS for the schedule, AViTS for the resolution transitions. If you're serving a single model at scale, LinCa's learned predictors pay for themselves quickly.

The Bottom Line ​

If you're building a text-to-image pipeline and want a drop-in speedup with no retraining, start with OYS-style schedule optimization. A 5-step schedule that keeps 90% of the quality is the easiest 10x you'll find.

If you're running high-resolution DiTs and care about FLOPs or latency, adopt selective token upsampling like AViTS. It composes with distillation and quantization, so it won't block future optimizations.

If you're doing Bayesian imaging inverse problems on multi-GPU infrastructure, PiX-MC is the one to watch. A 50x wall-clock speedup on CT reconstruction changes what's feasible in clinical settings.

One thing to watch: these methods are designed to stack, and the combined results will be bigger than any single number here. Expect published 15-20x end-to-end speedups for DiT pipelines within the next year, and plan your serving infrastructure for that.