Appearance
The golden age started with a shortage
Two years ago, the standard answer to any local-inference complaint was to buy more GPU and stop tweaking. The hardware shortage ended that. I've watched the local-LLM threads pivot from hardware shopping to source-level optimization, and the shift has paid off in measurable ways. Last week I spent an afternoon benchmarking forked llama.cpp builds on a Strix-class APU with Qwen 3.8 Flash Next (Q38FN). The results are the best argument the community has made in years: 52 tokens per second decode, which streams roughly ten times faster than you can read, and 1,300 tokens per second prefill. The prefill number is the one that impressed me. A 10,000-token prompt, a dense paper or a large file, is fully processed in about eight seconds before generation even starts.
The baseline deserves context. Stock llama.cpp on the same hardware runs about 26 tok/s decode and roughly 220 tok/s prefill; those are implied by the 2x and 5-6x gains in the benchmark reports. The two halves of the improvement are worth keeping separate. On the engine side, the forked llama.cpp builds and the halogen-flash-server project added kernels tuned for Strix Halo's unified-memory layout. On the model side, Q38FN's Engram architecture is the other half of the story: the model is small, and the community reports it's smarter than its size suggests.
| Metric | Stock llama.cpp | Community fork | Gain |
|---|---|---|---|
| Decode | ~26 tok/s | 52 tok/s | 2x |
| Prefill | ~220 tok/s | 1,300 tok/s | 5-6x |
The practical implication is simple: a generation that took ten minutes last year takes five, and prompt processing that took a minute now takes about ten seconds. Neither change required new silicon. It's the same machine, running software that now understands the machine.
Key numbers
- 52 tok/s decode on the community fork, 2x over stock and roughly 10x the pace of human reading.
- 1,300 tok/s prefill, a 5-6x gain: a 10K-token document is ingested in about eight seconds.
- 1.8 TB/s vs 1.2 TB/s: the 5090's bandwidth advantage over the M5 Ultra, which maps directly to decode speed.
There's a reason these posts read like nostalgia. The early web ran on forums, IRC, and shared scripts, and the people who came up in that era understood the stack from bare metal up. The convenience platforms that followed optimized that literacy away into endless feeds. The local-LLM scene reversed the slide: scarce compute forces you to learn. When you have too little, you dig in; when you have too much, you scroll. Right now there isn't enough to go around, and it's producing exactly the kind of end-to-end thinking the field needed.
The hardware question everyone keeps asking
The same scarcity drives a specific question through every hardware thread: sell the RTX 5090 for a Mac Studio M5 Ultra with 96GB? I've nearly made this swap myself. I can get $5,000 for the 5090, and the Mac runs $5,499 before tax. On paper it reads like a downgrade: the 5090's memory bandwidth is 1.8 TB/s against the M5 Ultra's 1.2 TB/s. On paper is where the mistake lives.
Decode speed in llama.cpp scales with memory bandwidth, so any model that fits in 32GB will stream tokens about 1.5x faster on the 5090. The qualifier is the whole argument. The M5 Ultra fits a 70B-class model at Q4, around 40GB of weights, with room left over for a long context. The 5090 doesn't come close. For coding, that's the difference between a model that reads a whole repository and a model that sees a slice of it.
Context is the hidden multiplier. A 128K context window isn't free; its KV cache eats VRAM in gigabytes, and on 32GB you're permanently trading model size against context length. At 96GB you stop budgeting. That single line is what pushes me toward the Mac: the 5090 buys speed on models I can already fit, while the Mac buys the ability to run models I currently can't.
But the decision isn't only bandwidth and capacity. The 5090 keeps you in the CUDA world: vLLM, TensorRT-LLM, exllama loaders, and the full PyTorch stack for fine-tuning. A Mac points you at llama.cpp and MLX, and that's most of your software universe. If the plan includes server-style batching or real fine-tuning, keep the 5090 and eat the resale loss. If the plan is llama.cpp with a large context window, the Mac is the right call, and I'd probably make it.
Everyone who hits out-of-memory errors daily swears by unified memory. Everyone who values token speed or CUDA tooling won't touch a Mac. The recurring warning in the threads is about software, not silicon: a machine you can't run your stack on is an expensive paperweight, whatever its gigabytes.
Quick Take: In the local-LLM world right now, hardware buying is a capacity decision first and a speed decision second, with one exception: if you depend on CUDA tooling, that ecosystem wins over both.
Local VLMs: pick by tier, not by leaderboard
The VLM recommendation threads on r/LocalLLaMA enforce a structure for a reason. Open-weights only, memory footprint classification, a real description of the use case, and prompt plus tooling details. The rules exist because VLM benchmarks are untrustworthy, the tooling is immature, and inference is stochastic; run the same prompt twice and you get different answers. The footprint taxonomy the threads use: S is under 8GB, M is 8 to 32GB, L is 32 to 64GB, XL is 64 to 128GB, Unlimited is beyond 128GB.
| Tier | Memory | Model class that fits at Q4 | What people run it for |
|---|---|---|---|
| S | under 8GB | 2-4B | OCR, screenshots, quick captions, laptop iGPUs |
| M | 8-32GB | 7-20B | Document QA, chart reading, the daily workhorse |
| L | 32-64GB | 30-40B | High-res images, long documents, heavier MoE VLMs |
| XL | 64-128GB | 70B+ | Multi-image batches, video frames, long context |
| Unlimited | over 128GB | 100B+ or full precision | Batch processing, evaluation, research |
The pattern across the best answers is that people don't run one VLM; they run a portfolio. I do the same now: a model under 8GB handles screenshots and phone captures because it costs almost nothing per call, and a mid-tier model reads documents and charts. The budget split matters more than picking the single "best" model, because best-on-leaderboard is usually a benchmark artifact.
The tier math gets concrete fast. An 8B VLM at Q4 sits around 5GB, squarely in S tier, and runs on a 4060 or even an integrated GPU. A 32B at Q4 is roughly 20GB, an M-tier daily driver. A 70B at Q4 is about 40GB, which lands in L or XL depending on the context you need. Each step up buys resolution and multi-image handling, not just parameter count; the KV cache for high-resolution vision tokens is brutal, and that's another reason the tiers exist.
The open-weights rule keeps the recommendations honest: a model you can't download and run is just marketing. And the thread's demand for detail is the right medicine for the evaluation problem. Benchmarks lie, so people test on their own documents with their own prompts, several runs each. That's why the good answers share full setups, not just model names.
Open weights are binaries: hashes, licenses, and run status
The more models you download, the more the registry layer matters, and it's finally getting proper engineering attention. The Hugging Bay is the example people keep pointing at: a searchable index of open models, license comparison, inspectable source records, and hosted files with published SHA-256 hashes. Each artifact page also shows whether the file actually runs, which sounds mundane until you've spent an evening debugging a loader error from a checkpoint corrupted in transit or converted with outdated scripts.
Open weights are binaries, and binaries are a trust problem. A quantized checkpoint from a repo without a hash is an act of faith: you're betting the file matches its model card, that the Q4 conversion used the right tool, and that nobody tampered with the upload. A published SHA-256 hash changes that. You can verify a download byte-for-byte before loading it, and the check takes seconds, not hours.
License comparison is the part people skip and regret later. "Open weights" is not one legal bucket. Some licenses bar commercial use, some require reproducing license text in derivatives, some allow anything short of misrepresenting the source. For hobby inference none of this matters. For a product, it's the difference between shipping and reworking your model stack. Comparing licenses before you download takes minutes; discovering a restriction after release takes much longer.
The registry also has a tool surface. Agents can query it over MCP or plain REST, which is the bridge to the coordination pattern below.
Agent coordination: the shared blackboard
Local inference used to be one model, one machine, one prompt. The newer pattern is several models cooperating across machines, and it needs a shared memory. The Hugging Bay's blackboard is a public scratchpad where agents post asks, results, notes, and messages under a task key, then read what other agents posted. It's exposed as an MCP tool at /api/mcp or as a plain REST POST to /api/v1/blackboard/put, with no account and no API key. That matters for local stacks: a coordination layer without auth overhead fits scripts, cron jobs, and agent loops.
What the community is saying about multi-agent local pipelines is consistent: the bottleneck is coordination between models, not the models themselves. People running several models across machines describe the same failure mode, duplicated work and lost context between steps. The blackboard is a direct answer to that complaint.
The design constraints are tuned for the job. Text messages cap at 16,000 characters, enough for a long document chunk or a detailed result; the full value fits in 64 KiB of JSON. Messages carry a kind, a task, optional topics, and a parent_id for threading. The practical payoff is de-duplication. If two agents both OCR the same screenshot because the pipeline has no shared memory, you're burning the scarce compute the community is trying to economize. Written once under a task key, read by everyone who needs it, the blackboard turns that into a single cheap call.
A tiny S-tier model can join this loop as easily as a big one: the limits are small enough that even a model on an iGPU can answer as an agent. The parent_id field turns the board into a threaded log, so a chain of cooperating models leaves a trace you can debug. That's the piece most local pipelines I've seen are missing: clean bookkeeping between agents.
Common pitfalls
The same lessons repeat in every thread, so here are the ones I keep re-learning.
Rating hardware on bandwidth alone. The 1.8 TB/s versus 1.2 TB/s comparison reads as decisive until the model doesn't fit. Capacity decides what you can run; bandwidth decides how fast. Check capacity first.
Benchmarking stock llama.cpp on a chip with community forks. On Strix Halo, the official build is a starting point, not a reference. The forks hit 52 tok/s decode and 1,300 tok/s prefill on Q38FN, 2x decode and 5-6x prefill over stock. Before you spend a weekend tuning, check whether specialized kernels exist for your hardware.
Treating the advertised context window as a budget. A 32B model at Q4 with a 128K context generates a KV cache measured in tens of gigabytes. Pick context by task. For coding, 32K usually covers the files you actually need, and the saved memory goes to model size.
Hoarding fp16 weights. A 70B checkpoint at fp16 is roughly 140GB; the same model at Q4_K_M is about 40GB. Most local tasks can't measure a quality difference between Q4 and Q8, let alone fp16. The gigabytes you save become context and batch size.
Skipping hash and license checks. A download without a SHA-256 check can silently be a broken artifact. A license you read as permissive when it isn't can become a product you can't ship. Both checks take minutes.
One thing to remember
The through-line of this moment is that software is catching up to the constraints. Inference forks are faster, small models are smarter per byte, registries are adding provenance, and coordinated agents are turning scattered machines into one pipeline. If you bought hardware six months ago, sit on it. The same machine is capable of more today than it was at purchase, and it will be capable of more in six months, because the stack around it is improving faster than the silicon is.
The Bottom Line: three calls to make right now
If your local coding workflow is dominated by out-of-memory errors and truncated context, move to a 96GB unified-memory machine and accept the decode penalty: stepping from a 30B-class model to a 70B changes output quality more than a 1.5x token rate does.
If your models fit in 32GB and you rely on CUDA tooling (vLLM, TensorRT-LLM, PyTorch fine-tuning), keep the 5090. The $5,000 resale is tempting, but the M5 Ultra costs more and runs your software universe slower per token.
One thing to watch: provenance and coordination are becoming API surfaces. Within six months, expect SHA-256 verification, license comparison, and MCP blackboards to be defaults in the local stack rather than extras. Build your download and pipeline scripts around those interfaces now.