Which quantization should I pick?
Quantization stores model weights at fewer bits to save VRAM. The trade-off is a small quality drop that's usually invisible at Q4_K_M and up. Here's the decision tree, plus a lookup table for every model in the catalogue on a 24 GB RTX 4090.
Decision tree
- Do you have >= 2 × params in GB of VRAM? Run fp16. No quantization trade-off.
- Do you have ≈ 1 × params in GB? Q8_0 — near-lossless with half the memory of fp16.
- Do you have ≈ 0.7 × params in GB? Q5_K_M — quality drop under 1 % on most benchmarks.
- Do you have ≈ 0.6 × params in GB? Q4_K_M. This is the mainstream default and what Ollama defaults to.
- Anything less? Try a smaller model at Q4 first — it'll almost always beat a heavily-quantized larger one at real workloads.
Quantization levels compared
| Level | Bits / weight | GB per B params | Typical quality loss | When to use |
|---|---|---|---|---|
| fp16 | 16 | 2.00 | None (baseline). | Production inference on H100/H200 where memory isn't the bottleneck. Fine-tuning. |
| fp8 | 8 | 1.00 | Negligible on H100+ hardware. | Production hosted inference. Doubles throughput vs fp16 with no measurable quality drop on Hopper/Blackwell. |
| Q8_0 | 8 | 1.00 | ≈ 0.1 % on academic benchmarks. | Highest-quality local option. Use when you have the VRAM headroom (2× the params in GB). |
| Q6_K | 6.6 | 0.83 | ≈ 0.3 % typical. | Sweet spot between Q8 and Q5. Rarely the default, but useful when Q5 feels too aggressive. |
| Q5_K_M | 5.5 | 0.69 | ≈ 1 % typical. | One notch above the mainstream default. Choose when you have 30–40 % more VRAM than Q4 needs. |
| Q4_K_M | 4.85 | 0.60 | ≈ 1–3 % typical. | The mainstream local-inference default. 4× smaller than fp16 with modest quality cost. |
| Q3_K_M | 3.9 | 0.49 | ≈ 3–6 %; visible on hard prompts. | Emergency only — when the model won't fit at Q4. Consider dropping to a smaller model instead. |
| Q2_K | 2.6 | 0.33 | ≈ 6–12 %; clearly degraded. | Testing only. Don't run this in production. |
Recommended quantization on a 24 GB RTX 4090
Assumes 25 % of VRAM reserved for KV cache + runtime overhead. Real-world numbers vary with context length; leave slack for long contexts.
Related: hardware guide · quantization glossary.