OSAIM
Open Source AI Models

Which quantization should I pick?

Quantization stores model weights at fewer bits to save VRAM. The trade-off is a small quality drop that's usually invisible at Q4_K_M and up. Here's the decision tree, plus a lookup table for every model in the catalogue on a 24 GB RTX 4090.

Decision tree

  1. Do you have >= 2 × params in GB of VRAM? Run fp16. No quantization trade-off.
  2. Do you have ≈ 1 × params in GB? Q8_0 — near-lossless with half the memory of fp16.
  3. Do you have ≈ 0.7 × params in GB? Q5_K_M — quality drop under 1 % on most benchmarks.
  4. Do you have ≈ 0.6 × params in GB? Q4_K_M. This is the mainstream default and what Ollama defaults to.
  5. Anything less? Try a smaller model at Q4 first — it'll almost always beat a heavily-quantized larger one at real workloads.

Quantization levels compared

LevelBits / weightGB per B paramsTypical quality lossWhen to use
fp16162.00None (baseline).Production inference on H100/H200 where memory isn't the bottleneck. Fine-tuning.
fp881.00Negligible on H100+ hardware.Production hosted inference. Doubles throughput vs fp16 with no measurable quality drop on Hopper/Blackwell.
Q8_081.00≈ 0.1 % on academic benchmarks.Highest-quality local option. Use when you have the VRAM headroom (2× the params in GB).
Q6_K6.60.83≈ 0.3 % typical.Sweet spot between Q8 and Q5. Rarely the default, but useful when Q5 feels too aggressive.
Q5_K_M5.50.69≈ 1 % typical.One notch above the mainstream default. Choose when you have 30–40 % more VRAM than Q4 needs.
Q4_K_M4.850.60≈ 1–3 % typical.The mainstream local-inference default. 4× smaller than fp16 with modest quality cost.
Q3_K_M3.90.49≈ 3–6 %; visible on hard prompts.Emergency only — when the model won't fit at Q4. Consider dropping to a smaller model instead.
Q2_K2.60.33≈ 6–12 %; clearly degraded.Testing only. Don't run this in production.

Recommended quantization on a 24 GB RTX 4090

Assumes 25 % of VRAM reserved for KV cache + runtime overhead. Real-world numbers vary with context length; leave slack for long contexts.

Related: hardware guide · quantization glossary.