← Back to all tools中文

LLM Deployment Checklist

Interactive phased checklist for LLM deployment (hardware, software, model, security)

Progress overview

Overall progress 0/30 (0%) · Critical items 0/10

Critical items remain — not ready to ship

Hardware
0/6 (0%)
Software
0/6 (0%)
Model
0/6 (0%)
Security
0/7 (0%)
Monitoring
0/5 (0%)

Checked state is saved automatically in your browser

Checklist

1. Hardware0/6 (0%)

  • GPU VRAM capacity checkCriticalVRAM must hold model weights (params × bytes per param) + KV cache (grows with concurrency and context length) + framework overhead. Reserve 10–20% headroom.
  • GPU count & interconnect topologyConfirm the model fits on one GPU; otherwise plan tensor parallelism across cards and check NVLink / PCIe bandwidth for bottlenecks.
  • Sufficient system RAMSystem RAM should fit at least one full copy of the weights (loading/conversion passes through CPU memory); CPU offloading needs more.
  • Storage capacity & speedUse NVMe SSDs for weights, logs and temp files. Load speed drives cold-start time; keep free space ≥ 2× model size.
  • Network bandwidth estimateEstimate ingress/egress at peak concurrency (streaming tokens × connections), and pull speed from the model registry / object storage.
  • Power & cooling headroomOn-prem: verify PSU wattage, UPS and cooling can sustain GPUs at full load. Cloud: confirm quotas and zone capacity.

2. Software0/6 (0%)

  • OS & kernel tuningUse a supported Linux distro; tune file descriptor limits, shared memory (/dev/shm) and network sysctl as needed.
  • GPU driver versionCriticalInstall an NVIDIA driver matching the CUDA toolchain; nvidia-smi must list all GPUs. Record the driver version in the deploy doc.
  • CUDA / cuDNN compatibilityCriticalVerify the CUDA runtime matches the inference engine's support matrix — version mismatch is a top cause of failed deploys.
  • Inference engine selectionCriticalPick vLLM / SGLang / TensorRT-LLM / TGI / llama.cpp etc. per use case; confirm support for the model architecture, quantization format and required features (streaming, tool calling).
  • Pinned runtime environmentPin Python and dependency versions via a Docker image or lockfile; scan the image for vulnerabilities and push to a private registry.
  • Performance benchmarkLoad-test on real hardware: time to first token (TTFT), throughput (tokens/s), max concurrency. Proceed only when the SLA is met.

3. Model0/6 (0%)

  • Model format compatibilityCriticalConfirm the weight format (safetensors / GGUF / AWQ / GPTQ) is natively supported by the engine; avoid silent fallback to a slow path.
  • Quantization & quality regressionChoose a precision (FP16 / BF16 / INT8 / INT4) and run quality regression on your eval set to confirm acceptable accuracy loss.
  • Model license reviewCriticalConfirm the license permits your use (commercial, redistribution, fine-tune commercialization); watch extra terms and MAU limits in Llama, Qwen, etc.
  • Version pinning & rollbackRecord the exact model version / commit hash and weight checksums; keep the previous version available for fast rollback.
  • Tokenizer & chat templateVerify tokenizer config and chat template match training — mismatches visibly degrade output. Test special-character splitting for multilingual use.
  • Context length configurationSet max context per business need; for extended context verify RoPE scaling / YaRN settings and measure long-input quality and VRAM usage.

4. Security0/7 (0%)

  • API authenticationCriticalAll external endpoints require auth (API key / OAuth / mTLS); never expose an unauthenticated inference port. Issue per-caller credentials.
  • Rate limiting & quotasSet per-user / per-key QPS, concurrency and daily token quotas so one caller can't saturate the GPUs.
  • Network isolationCriticalKeep inference services inside a private network / VPC behind a gateway or reverse proxy; security groups and firewalls allow only required ports.
  • Input guardrailsCap input length; deploy prompt-injection detection and malicious-content filtering where needed to prevent abuse.
  • Output filtering & redactionDetect and redact sensitive data (PII, secrets) in outputs; add content-safety filtering for end-user-facing scenarios.
  • Secrets managementStore API keys and HF tokens in KMS / Vault / injected env vars — never in images, repos or logs.
  • Audit loggingLog caller identity, timestamps and usage for audit (avoid storing raw sensitive prompts, or encrypt and access-control them).

5. Monitoring0/5 (0%)

  • Structured request loggingCriticalLog request id, latency, input/output token counts, model version and error code per request for debugging and billing.
  • Metrics collectionCollect GPU utilization, VRAM usage, throughput, queue depth, TTFT/TPOT (Prometheus + DCGM or the engine's built-in metrics).
  • Alerting rulesCriticalAlert on error-rate spikes, latency over threshold, OOM, GPU loss and disk filling up; route notifications to the on-call channel.
  • Monitoring dashboardBuild a Grafana dashboard showing QPS, latency distribution, GPU status and quota usage; rehearse one incident triage before launch.
  • Cost & usage trackingTrack token usage and GPU-hours per tenant / project; regularly check unit cost against budget.

Actions