VRAM Estimator
Estimate GPU memory requirements based on model size, quantization and context length
Model architecture (affects KV cache)
Estimate
| Model weights | 14.2 GB |
| KV cache | 0.22 GB |
| Runtime overhead | 1.44 GB |
| Total VRAM required | 15.9 GB |
Fits on common GPUs
- RTX 3060 (12 GB)No
- RTX 4070 Ti (16 GB)Tight (<10% headroom)
- RTX 3090 / 4090 (24 GB)Yes
- RTX A6000 (48 GB)Yes
- A100 / H100 (80 GB)Yes
- H200 / B200 (141 GB)Yes
Estimates are for capacity planning only; real usage varies with the inference engine (vLLM, llama.cpp, etc.), CUDA context and activations. The KV cache uses the byte width of the selected quantization.