LLM Deployment Checklist
Interactive phased checklist for LLM deployment (hardware, software, model, security)
Progress overview
Overall progress 0/30 (0%) · Critical items 0/10
Critical items remain — not ready to ship
Hardware
0/6 (0%)
Software
0/6 (0%)
Model
0/6 (0%)
Security
0/7 (0%)
Monitoring
0/5 (0%)
Checked state is saved automatically in your browser
Checklist
1. Hardware0/6 (0%)
- GPU VRAM capacity checkCriticalVRAM must hold model weights (params × bytes per param) + KV cache (grows with concurrency and context length) + framework overhead. Reserve 10–20% headroom.
- GPU count & interconnect topologyConfirm the model fits on one GPU; otherwise plan tensor parallelism across cards and check NVLink / PCIe bandwidth for bottlenecks.
- Sufficient system RAMSystem RAM should fit at least one full copy of the weights (loading/conversion passes through CPU memory); CPU offloading needs more.
- Storage capacity & speedUse NVMe SSDs for weights, logs and temp files. Load speed drives cold-start time; keep free space ≥ 2× model size.
- Network bandwidth estimateEstimate ingress/egress at peak concurrency (streaming tokens × connections), and pull speed from the model registry / object storage.
- Power & cooling headroomOn-prem: verify PSU wattage, UPS and cooling can sustain GPUs at full load. Cloud: confirm quotas and zone capacity.
2. Software0/6 (0%)
- OS & kernel tuningUse a supported Linux distro; tune file descriptor limits, shared memory (/dev/shm) and network sysctl as needed.
- GPU driver versionCriticalInstall an NVIDIA driver matching the CUDA toolchain; nvidia-smi must list all GPUs. Record the driver version in the deploy doc.
- CUDA / cuDNN compatibilityCriticalVerify the CUDA runtime matches the inference engine's support matrix — version mismatch is a top cause of failed deploys.
- Inference engine selectionCriticalPick vLLM / SGLang / TensorRT-LLM / TGI / llama.cpp etc. per use case; confirm support for the model architecture, quantization format and required features (streaming, tool calling).
- Pinned runtime environmentPin Python and dependency versions via a Docker image or lockfile; scan the image for vulnerabilities and push to a private registry.
- Performance benchmarkLoad-test on real hardware: time to first token (TTFT), throughput (tokens/s), max concurrency. Proceed only when the SLA is met.
3. Model0/6 (0%)
- Model format compatibilityCriticalConfirm the weight format (safetensors / GGUF / AWQ / GPTQ) is natively supported by the engine; avoid silent fallback to a slow path.
- Quantization & quality regressionChoose a precision (FP16 / BF16 / INT8 / INT4) and run quality regression on your eval set to confirm acceptable accuracy loss.
- Model license reviewCriticalConfirm the license permits your use (commercial, redistribution, fine-tune commercialization); watch extra terms and MAU limits in Llama, Qwen, etc.
- Version pinning & rollbackRecord the exact model version / commit hash and weight checksums; keep the previous version available for fast rollback.
- Tokenizer & chat templateVerify tokenizer config and chat template match training — mismatches visibly degrade output. Test special-character splitting for multilingual use.
- Context length configurationSet max context per business need; for extended context verify RoPE scaling / YaRN settings and measure long-input quality and VRAM usage.
4. Security0/7 (0%)
- API authenticationCriticalAll external endpoints require auth (API key / OAuth / mTLS); never expose an unauthenticated inference port. Issue per-caller credentials.
- Rate limiting & quotasSet per-user / per-key QPS, concurrency and daily token quotas so one caller can't saturate the GPUs.
- Network isolationCriticalKeep inference services inside a private network / VPC behind a gateway or reverse proxy; security groups and firewalls allow only required ports.
- Input guardrailsCap input length; deploy prompt-injection detection and malicious-content filtering where needed to prevent abuse.
- Output filtering & redactionDetect and redact sensitive data (PII, secrets) in outputs; add content-safety filtering for end-user-facing scenarios.
- Secrets managementStore API keys and HF tokens in KMS / Vault / injected env vars — never in images, repos or logs.
- Audit loggingLog caller identity, timestamps and usage for audit (avoid storing raw sensitive prompts, or encrypt and access-control them).
5. Monitoring0/5 (0%)
- Structured request loggingCriticalLog request id, latency, input/output token counts, model version and error code per request for debugging and billing.
- Metrics collectionCollect GPU utilization, VRAM usage, throughput, queue depth, TTFT/TPOT (Prometheus + DCGM or the engine's built-in metrics).
- Alerting rulesCriticalAlert on error-rate spikes, latency over threshold, OOM, GPU loss and disk filling up; route notifications to the on-call channel.
- Monitoring dashboardBuild a Grafana dashboard showing QPS, latency distribution, GPU status and quota usage; rehearse one incident triage before launch.
- Cost & usage trackingTrack token usage and GPU-hours per tenant / project; regularly check unit cost against budget.