← Back to all tools中文

VRAM Estimator

Estimate GPU memory requirements based on model size, quantization and context length

Model architecture (affects KV cache)

Estimate

Model weights14.2 GB
KV cache0.22 GB
Runtime overhead1.44 GB
Total VRAM required15.9 GB

Fits on common GPUs

Estimates are for capacity planning only; real usage varies with the inference engine (vLLM, llama.cpp, etc.), CUDA context and activations. The KV cache uses the byte width of the selected quantization.