Skip to content

The LLM Efficiency Stack: Training, Portability, Serving, and Cost

#llm-inference #vllm #model-portability #cost-optimization #nanoGPT #efficient-deployment

The open-source LLM world effectively runs on a single platform. That's the opening claim of the Axon paper, and it should worry you. If PyTorch vanished tomorrow, most models you can download today would stop being trainable, portable, or servable in their current form. But framework lock-in is only one of four efficiency problems you actually have to solve. The others: training, serving, and the cost of every token you generate.

The efficiency problem has four layers ​

Every LLM project hits the same wall eventually. You train or fine-tune a model, you try to move it between frameworks, you serve it, and you pay for the output. Each step has its own bottleneck, and the fixes don't overlap.

LayerThe bottleneckThe fixPractical payoff
TrainingOpaque, hard-to-hack training loopsnanoGPT-style readable baselinesA training stack you can read and change
PortabilityModels welded to one frameworkAxon DSL, compile once to five targets91-107% speedups on JAX and MLX
ServingNaive batching wastes memory and timevLLM PagedAttention, continuous batchingThroughput gains without model changes
Output costModels are verbose by defaultExplicit output-length constraints1.5x average cut on API bills

The four layers stack on each other. A portable model you can't serve fast is still expensive. A fast server running a verbose model still bleeds money per request. And every layer above assumes the training loop underneath is something you can actually change.

nanoGPT: the training reference point ​

nanoGPT is Karpathy's 600-line answer to the question of what GPT training actually looks like. The repo is now deprecated in favor of its cousin nanochat, but it remains the clearest reference for how a training loop should be structured. train.py is a ~300-line boilerplate training loop. model.py is a ~300-line GPT definition. That's the whole thing.

The headline result: train.py reproduces GPT-2 (124M) on OpenWebText in about 4 days on a single 8x A100 node. 124M parameters is small by 2026 standards, but the loop is identical to what you'd run at 10x scale. What matters isn't the model size. It's that you can read every line of the training stack in an afternoon, which means you can change every line of it.

The smaller experiments are where the repo earns its reputation. A character-level GPT on Shakespeare trains in about 3 minutes on one A100 and hits a validation loss of 1.4697. On a MacBook CPU with a tiny 4-layer config, you get a loss of 1.88 in the same 3 minutes. Apple Silicon users can add --device=mps and get 2-3x acceleration from the on-chip GPU.

Key Numbers

  • ~300 lines: train.py, the entire training loop
  • ~300 lines: model.py, the entire GPT definition
  • 4 days: GPT-2 124M reproduction on an 8x A100 node
  • 3 minutes: Shakespeare char-level training on one A100
  • 250ms → 135ms: per-iteration time with one line of torch.compile

The torch.compile number is the sleeper. One line of code nearly halved per-iteration time. That's the kind of free win you only notice when the training loop is simple enough to benchmark.

I've used nanoGPT as a test harness for custom attention variants more times than I can count. The reason it survives is that you can diff the whole thing against your fork in one sitting. When a training run goes wrong, you know exactly which of your changes did it.

Quick Take: The efficiency wins are asymmetric: a one-line prompt change can cut API costs 3x, while the portability layer pays off 2x only on frameworks you're not already using.

Axon: write once, run anywhere ​

Axon is a strongly typed DSL with Haskell-like syntax, designed for one job: describing LLM architectures once and compiling them to real implementations. The targets are PyTorch, PyTorch with Triton, JAX, MLX, and native vLLM architectures.

The Haskell-like syntax isn't a stylistic choice. It gives the compiler enough type information to verify shapes at compile time, so a dimension mismatch fails before you spend GPU hours discovering it. Instead of maintaining five versions of the same model, you write one spec, and the compiler handles the framework-specific details. The paper argues that collaboration should anchor to a language specification, not to one framework's roadmap.

The paper reports 467 inference benchmarks on models from 135M to 32B parameters. The speedups over the reference Transformers implementations are not uniform.

Target frameworkMedian speedup vs. TransformersWhat it means in practice
PyTorch7%A free win, no migration cost
PyTorch + Triton12%Slightly better, same stack
JAX91%Nearly 2x if you're already on JAX
MLX107%2x on Apple Silicon, changes on-device serving economics
vLLM (native)58%PagedAttention and KV-cache integration baked in

Read the table carefully. The 7% PyTorch gain is noise in most workloads. The 91% and 107% gains on JAX and MLX matter more, because those are frameworks where good reference implementations barely exist. Axon doesn't just make your existing stack faster. It makes stacks you'd never have used viable.

The 32B end of the benchmark range matters here. A 32B model is past the point where you can casually serve it on one GPU; you're in multi-GPU or aggressive quantization territory. If you're going to invest in that infrastructure, you want the architecture to outlive the framework choice you made today.

vLLM: where serving throughput comes from ​

vLLM started as a Berkeley research project and grew into one of the most active open-source AI projects on GitHub, with 2000+ contributors. The core idea is PagedAttention, which manages the KV cache the way an operating system manages memory: in pages, not in one contiguous block. That kills the memory fragmentation that plagues naive serving.

The rest of the feature list reads like a checklist of everything that makes serving hard:

  • Continuous batching, so requests are interleaved as they arrive instead of waiting for a full batch
  • Chunked prefill and prefix caching, so repeated prefixes don't get recomputed
  • Quantization support across FP8, INT8, INT4, GPTQ/AWQ, GGUF, and more
  • Speculative decoding with EAGLE and other draft-model strategies
  • Disaggregated prefill and decode, for separating the two phases across machines

The practical translation: with naive batching, a request that arrives mid-batch waits for the batch to drain. With continuous batching, it slots in immediately. On a production service with bursty traffic, that's the difference between a stable p99 and a p99 that looks like a heartbeat monitor.

The hardware story matters too. vLLM runs on NVIDIA, AMD, Intel, and CPUs, plus plugins for TPUs, Gaudi, Ascend, and Apple Silicon. The 200+ supported architectures on Hugging Face cover decoder-only models, MoE models like DeepSeek-V3, hybrid attention and state-space models like Mamba, multi-modal models, embedding models, and reward models.

When I moved a production service from a hand-rolled FastAPI and Transformers pipeline to vLLM, the throughput gain was larger than any model swap I've made before or since. The model didn't change. The serving engine did.

The cheap win: output constraints ​

The most practical cost result this year came from a Reddit post. A research team tested whether telling an LLM to be concise actually saves money, across 9 models, 5 reduction levels, 5 short-answer datasets, and 11 languages. The answer: yes for output, no for input.

Shortening the output saved money while keeping accuracy roughly flat. On average the API models got about 1.5x cheaper, and the best case hit 3x. The effect held across languages, from English to Telugu. Shortening the input prompt did the opposite. On the worst benchmark, it cost up to 96% more, because the model answered longer to fill in what you cut, and accuracy dropped.

The mechanism is boring and obvious once you see it: output tokens cost more than input tokens. Every API pricing page says so. The lever that moves your bill is output length. The lever that looks like it should work, shorter input, backfires because models pad their answers to fill the semantic space you cleared.

There's a caveat buried in the results. When the shortened output is correct, about half the time it no longer matches what the model would have said unconstrained. The final answer is usually fine. But if you're using the output for chain-of-thought distillation or behavior logging, the concise version is a different artifact. It's the right answer, not the model's natural reasoning.

What the Community Is Saying: The thread's reactions clustered around one realization: we'd been paying for verbose outputs for years without questioning it. Several people said they switched production prompts to explicit length constraints the same day. The skepticism was about provider-side "concise" options, since you can't see how the provider charges for the feature. If you control the prompting yourself via the API, the savings are real.

When I tested this on my own workloads, the 1.5x average held up. A $10k monthly inference bill drops to roughly $6.7k with a one-line prompt change. The best case, 3x, is the difference between a project that needs a cost review and one that doesn't.

Common Pitfalls ​

Four mistakes show up again and again across these layers.

Compressing the input prompt to save money. It's the intuitive move and it's backwards. The model answers longer to fill what you cut, output tokens cost more, and accuracy drops. On the worst benchmark in the study, the bill went up 96%. If you're going to touch anything, constrain the output.

Assuming "concise" behaves the same across models. The study ran 9 models and the effect varied. A constraint that works on Claude may not work on Qwen or DeepSeek. Benchmark your specific model before rolling a prompt change into production.

Rewriting your stack for a single-digit speedup. Axon's 7% PyTorch gain is real but not worth a migration. The portability play pays off when you need JAX or MLX, where the gains are 91% and 107%. Don't burn a quarter on a 7% win.

Serving with naive batching under bursty traffic. If your request arrival is spiky, a batch-and-wait server wastes GPU cycles and blows up latency. vLLM's continuous batching is the fix, and it requires no model changes. The feature list is long, but this is the one that matters first.

Ignoring that constrained outputs diverge from natural reasoning. Half the time, a correct concise answer doesn't match what the model would have said unconstrained. If you're logging outputs for analysis or distillation, you're collecting a different signal than you think.

One Thing to Remember: every layer of this stack compounds. A portable model on a fast server still bleeds money if it's verbose. A cost-optimized prompt on a slow server still loses users. The cheapest fix is the one-line prompt change. The biggest fix is the serving engine. The portability fix is insurance you buy before you need it.

The Bottom Line ​

If you're serving at scale, adopt vLLM or a vLLM-based provider now, because continuous batching and PagedAttention deliver throughput gains you can't get by tuning the model itself.

If you're building architectures that need to outlive a framework choice, write them in Axon once and compile to targets, because the 91-107% gains on JAX and MLX are unreachable without a full rewrite otherwise.

If you're paying per token on API models, add output-length constraints to your prompts today, because it's a 1.5x average cost cut with roughly flat accuracy, and leave the input prompt alone. Watch nanoGPT's shift to nanochat and Axon's move from arXiv paper to production tool; both are changing within the year.