Quantization Comparison
Compare FP16, INT8, INT4, GGUF, AWQ and GPTQ for quality, speed and compatibility
Quantization Method Comparison
| Method | Quality loss (perplexity ↑) | Inference speed | Memory reduction | Hardware compatibility | Best use case |
|---|---|---|---|---|---|
| FP16 | None (baseline)(0%) | 1× (baseline) | — | All GPUs / CPUs / Apple Silicon | Quality-first training & inference baseline with ample VRAM |
| INT8 | Near-lossless(<1%) | ~1.6× | ↓ 47% | NVIDIA / AMD GPUs, some CPU instruction sets | Production serving that must stay close to FP16 quality |
| INT4 | Noticeable degradation(3–8%) | ~2.4× | ↓ 72% | Newer NVIDIA GPUs (TensorRT-LLM etc.) | Maximum throughput where some quality loss is acceptable |
| GGUF Q8_0 | Near-lossless(<1%) | ~1.3× | ↓ 47% | CPU, Apple Silicon, edge devices like Raspberry Pi | Running locally on CPU / Mac without quality loss |
| GGUF Q4_K_M | Slight degradation(1–3%) | ~1.8× | ↓ 70% | CPU, Apple Silicon, phones (llama.cpp / ollama) | First choice for laptops, home machines and mobile |
| AWQ (W4A16) | Slight degradation(1–3%) | ~2.2× | ↓ 72% | NVIDIA GPUs (well supported by vLLM / TensorRT-LLM) | High-throughput API serving on GPU servers |
| GPTQ (4-bit) | Slight degradation(1–4%) | ~2.0× | ↓ 72% | NVIDIA / AMD GPUs, mature ecosystem (ExLlama, vLLM) | When ready-made GPTQ checkpoints are already available |
Figures are rule-of-thumb estimates: real usage varies with framework, group size and context length — measure before committing.
Decision Guide: Which Should I Use?
Model Size Calculator
| Method | Weights | Total | Saved vs FP16 |
|---|---|---|---|
| FP16 | 14.0 GB | 14.0 GB | — |
| INT8 | 7.44 GB | 7.44 GB | 6.56 GB (↓ 47%) |
| INT4 | 3.94 GB | 3.94 GB | 10.1 GB (↓ 72%) |
| GGUF Q8_0 | 7.44 GB | 7.44 GB | 6.56 GB (↓ 47%) |
| GGUF Q4_K_M | 4.20 GB | 4.20 GB | 9.80 GB (↓ 70%) |
| AWQ (W4A16) | 3.94 GB | 3.94 GB | 10.1 GB (↓ 72%) |
| GPTQ (4-bit) | 3.94 GB | 3.94 GB | 10.1 GB (↓ 72%) |