Quantization

Detailed guide to model compression methods: GGUF K-quants, GPTQ, AWQ, MLX — and how to choose the right one.

🇷🇺 Russian version: quantization.ru.md


← Local models · Benchmarks →


Contents

  1. Why quantize?
  2. GGUF and K-quants
  3. Quality comparison
  4. Quantization by engine
  5. How to estimate RAM
  6. Tools for quantization
  7. Further reading
  8. Whats next

1. Why quantize?

Without quantization, large language models simply do not fit in most consumer devices:


2. GGUF and K-quants

GGUF (GPT-Generated Unified Format) is the main format for llama.cpp and related tools (Ollama, LM Studio, etc.).

K-quants in GGUF

These are special quantization schemes developed for llama.cpp that use different bit widths for different weight types based on importance.

Type Avg bits/weight Description When to use
Q2_K 2.0 Very low quality Only when extremely memory constrained
Q3_K_S 2.5 Low quality Very limited resources
Q3_K_M 3.0 Low-medium quality Very tight fit
Q4_0 4.0 Satisfactory Basic 4-bit uniform
Q4_K_S 4.3 Good quality Balance of size and quality
Q4_K_M 4.8 Excellent quality Recommended default
Q5_0 5.0 Very good When you need a bit more quality
Q5_K_S 5.3 Excellent quality Superior balance
Q5_K_M 5.5 Excellent + For demanding tasks
Q6_K 6.0 High quality For code, math, precise calculations
Q8_0 8.0 Near FP16 When memory allows and max quality needed
F16 16.0 Maximum quality Experiments, baseline comparison

3. Quality comparison

Example: 7B model

Method Size Perplexity (lower = better) Relative speed
F16 (FP16) ~13.5 GB Baseline (5.12) 1.0×
Q8_0 ~10.4 GB 5.15 (~+0.6%) ~0.9×
Q6_K ~7.8 GB 5.20 (~+1.6%) ~0.7×
Q5_K_M ~6.2 GB 5.25 (~+2.5%) ~0.6×
Q4_K_M ~4.9 GB 5.30 (~+3.5%) ~0.5×
Q4_0 ~4.5 GB 5.45 (~+6.4%) ~0.5×
Q3_K_M ~3.6 GB 5.80 (~+13.3%) ~0.4×
Q2_K ~2.3 GB 6.50 (~+27.0%) ~0.3×

Quick rule


4. Quantization by engine

Ollama / llama.cpp

mlx-lm (Apple Silicon)

GPTQ

AWQ


5. How to estimate RAM

For GGUF models

  1. Check the .gguf file size (e.g., 7.2 GB)
  2. Add ~1-2 GB for KV cache (depends on context length)
  3. Add 0.5-1 GB for framework overhead
  4. Total: file_size + 1.5-3.0 GB should be less than available RAM

Example: 6.8 GB model + 2.0 GB = 8.8 GB fits in 16 GB RAM with headroom.

For MLX formats

Similar but typically 10-20% less overhead due to more efficient storage on Apple Silicon.


6. Tools for quantization

llama.cpp

MLX

HuggingFace Transformers + bitsandbytes


7. Further reading


8. Whats next

If you want Go to
See benchmarks of different quantizations on Mac benchmarks/apple-silicon.md
Compare tools by quantization support tools.md
Choose a model for your task models.md
Back to navigation README.md

In section: getting-started · running-models · models · catalog · quantization · memory-and-context · tools · advanced-setup · troubleshooting · apple-silicon
Related sections: Zero Level · AI Agents · Use Cases
Navigation: ← Local Models · ↑ Back to main · 🇷🇺 Русский