Skip to content

Qwen3.8-27B at Q3_xxs: What a 16GB GPU Can Run Now

#qwen3.8-27b #quantization #local-llm #gguf #nvfp4 #inference

The 16GB GPU just got more interesting ​

I usually avoid Q3 quants. Too many bad experiences with models that degraded into mush below Q4. But there's no 35B-class MoE with a small active-parameter count released yet, so I tried Qwen3.8-27B at Q3_xxs on my 16GB RTX 4060 Ti.

It one-shot multiple serious coding tasks, producing fully working games and web apps on the first pass. The same tasks had either failed outright or taken hours of prompting and feedback with Qwen 3.6 35b. I don't use LLMs in agentic workflows, just plain text generation, but I need real coding help sometimes. This delivered.

Fully in VRAM, it runs at 30-35 t/s. That's the same speed I got from a higher-quant 35B model offloaded to RAM. Older dense models like Gemma 3 27B and Mistral Small 24B manage 13-17 t/s at best on the same hardware.

That's the story here. A dense 27B model at a quant most people avoid is outperforming bigger models at higher quants, and it fits in 16GB.

Why Qwen3.8-27B survives aggressive quantization ​

The architecture helps. This is a dense 27B model, not an MoE. Dense models quantize more predictably because every token routes through the same parameters. There's no router to destabilize at low precision, and no expert weights to lose. It also beat the higher Q4-Q5 MoEs I've tried, which tracks with that.

The training mix matters more. Qwen3.8-27B is code-heavy, and code is structured and repetitive. Syntax, function names, and API patterns survive low precision far better than the fuzzy semantics of open-ended conversation. That explains the model's weird profile: it one-shots serious math and logic tasks, then fails basic sorting or counting in casual chat. The low quant is likely responsible for the silly failures. The code-maxxed training explains the serious wins.

The quant menu ​

The quantized variants for this model have multiplied fast. GGUF for llama.cpp, MLX for Apple Silicon, NVFP4 for Blackwell, and a 1-bit quant that exists mostly as a punchline. Here's what each level gets you:

QuantHardwareWhat you get
Q3_xxs16GB cards, full VRAM30-35 t/s, strong on coding, occasional logic slips
Q4_0 / Q4_K_M16-24GB cardsThe safe default, reference 4-bit quality
NVFP4RTX 50 seriesSame footprint as Q4, 50% faster prefill
Q6_K24GB cardsQuality ceiling, slowest of the group
1-bit8GB cardsBrain damage. Fits, but nonsense output

Quick Take: the quantization floor for capable local models just dropped a full level, and Qwen3.8-27B is the model proving it.

There's also a whole uncensored variant scene. JonathanColetti's GGUF, orcarouter's MLX for Macs, OBLITERATUS's OBLITERATED re-tune, and a free HF Space endpoint from victor with 79 likes if you just want to test the model without downloading anything.

NVFP4 changes the speed equation ​

The most interesting quant is a speed story, not a size story. akopytko's NVFP4 GGUF is Blackwell-native and prefill-optimized. NVFP4 is NVIDIA's 4-bit floating-point format, native to the Blackwell architecture. It's not a generic re-quant; it maps directly onto the hardware's native format, which is why prefill is so much faster.

On an RTX 5090 32GB, it hits 6250 t/s prefill at pp2048. That's 50% faster than a Q4_0 quant at the same memory footprint, and 4-7% faster than other NVFP4 quants like unsloth's.

Prefill is the bottleneck for long prompts. Agents, RAG pipelines, code review: they all push thousands of tokens in before the model emits a single response token. Cutting prefill from 4130 to 6250 t/s drops time-to-first-token by about a third. In an interactive agent loop, that's the difference between feeling responsive and feeling slow.

MTP draft heads: free speed at the quant level ​

The NVFP4 GGUF also ships a quantized MTP draft head, and with the recommended settings it runs 15% faster than without it. MTP, or multi-token prediction, is a training-time technique where the model learns to predict several future tokens at once. At inference, the draft head generates candidates and the main model verifies them in a single pass. When the draft is right, you emit multiple tokens per forward pass.

HauhauCS's "Aggressive-MTP" GGUF pushes the same idea harder. More aggressive draft heads mean more tokens per pass when they're right, but the quality risk goes up when the draft gets too confident and the main model has to reject more.

What it's actually like to run this thing ​

The Q3_xxs experience holds up over a week of use. Generation stays at 30-35 t/s until context gets long, then it settles at 21-22 t/s. Still usable, but you notice it. The model's quirks don't go away: it will one-shot a tricky logic problem and then fail to sort a small list correctly in casual conversation. If you're using it for code, that trade is easy to accept.

I also ran the unsloth 1-bit quant on an 8GB card. It's brain damage. The output is nonsense, and it gave me a good laugh. But the fact that a 27B model fits in 8GB at all says something about how far quantization has come, even if the 1-bit level is a meme that no one should ship.

When my Claude Pro subscription expired, I switched my daily coding to Qwen3.8-27b on a 5090m 24GB running pi. Everything I was doing in claudecode still gets done. The one downside: claudecode left my GPU free, so now I have to plan around the GPU being busy. That's a real cost of local inference that nobody puts in the benchmark tables.

Key numbers

30-35 t/s: sustained generation on a 16GB RTX 4060 Ti with Q3_xxs, fully in VRAM. 6250 t/s: NVFP4 prefill on an RTX 5090, 50% faster than Q4_0 at the same footprint. 21-22 t/s: generation speed at long context. Still faster than older dense 27B models at their best. 15%: the MTP speedup from the quantized draft head with recommended settings.

The deployment stack around it ​

The model ecosystem is only half the story. Osmantic's ODS project is a one-command installer that turns a PC, Mac, or Linux box into a private AI server. It wires together llama-server, Open WebUI, n8n, ComfyUI, Qdrant, Whisper, and a dozen other services. The interesting part is the model catalog: the installer detects your GPU, assigns a hardware tier, and picks the right GGUF for your memory envelope.

TierHardwareDefault pickContext
08GB RAM, CPU-onlyQwen3.5 2B Q4_K_M8K
18GB discrete VRAM (RTX 4060, 3060 12GB)Qwen3.5 9B Q4_K_M32K
212GB discrete VRAM (RTX 4070-class)Phi-4 14B Q4_K_M16K
324GB discrete VRAM (RTX 4090, A6000)Qwen3.5 27B Q4_K_M32K
448GB discrete VRAM (A6000 Ada, L40S)DeepSeek R1 Distill 70B Q4_K_M32K

The catalog is a snapshot, and it's already behind the community. The 24GB tier picks Qwen3.5 27B, but people are running Qwen3.8-27B at Q3_xxs on 16GB cards manually. The automated stacks will catch up. Model selection is becoming a solved problem: you no longer need to shop quants by hand, the installer does it for you.

Common pitfalls ​

Carrying over "Q3 is always garbage" from older models is the most common mistake. Test per model family. Qwen3.8-27B at Q3_xxs outperforms bigger models at Q4-Q5, but that's because of its code-heavy training. The next model that survives Q3 this well may not exist yet.

Running NVFP4 on pre-Blackwell GPUs is a close second. NVFP4 is Blackwell-native. On an RTX 30 or 40 series, it won't be faster than Q4_0, and it may not load at all. Check your GPU generation before downloading.

Using an MTP draft head with default sampler settings wastes the speedup. The draft head only pays off with the settings the quant author documents. Default samplers can ignore the draft or slow you down. And "Aggressive" MTP variants can degrade output quality when the draft gets too confident.

Judging a quant by perplexity instead of task performance leads you astray. I saw weird failures at sorting and counting in casual chat, but the same model one-shot serious coding tasks. Benchmark the tasks you actually run.

Building on someone's free endpoint is a trap. The victor HF Space with 79 likes is fine for a quick test. It's not a production API. Download the GGUF or MLX variant and run it locally.

One thing to remember: the model is no longer the bottleneck in local AI. A 27B model at Q3_xxs on a 16GB card, an NVFP4 quant that cuts prefill in half, a one-command installer that picks the right model for your hardware. The bottleneck is your workflow, not your VRAM.

What this means ​

If you're on a 16GB card, run Qwen3.8-27B at Q3_xxs or Q4 instead of squeezing in an MoE. Dense models quantize more predictably, and this one's code-heavy training survives low precision better than anything else at the size.

If you're on an RTX 50 series, switch to the NVFP4 GGUF with the quantized MTP draft head. Same memory footprint as Q4_0, 50% faster prefill, 15% faster MTP with the right settings. That's the difference between a sluggish agent loop and a responsive one.

If you're setting up a local box for someone who isn't an ML engineer, use a stack like ODS that picks the model for the hardware. Manual quant shopping is becoming a hobby, not a requirement. One thing to watch: NVFP4 and quantized draft heads are moving fast, and prefill-speed gains like these will be the default within a few model generations.