Inference Engine Comparison
Compare vLLM, SGLang, TensorRT-LLM and llama.cpp features and use cases
Quick Selector Quiz
Answer 3 questions to get a recommendation
1. What hardware do you have?
2. What matters most to you?
3. What is your deployment scale?
Core Metrics
| Engine | Throughput | Latency (5=low) | Ease of setup (5=easy) | Hardware support | Quantization | Best for |
|---|---|---|---|---|---|---|
| vLLM | ●●●●● | ●●●●● | ●●●●● | NVIDIA GPU, AMD GPU, CPU, TPU | FP8, INT8, INT4 (AWQ), GPTQ, GGUF (limited) | High-concurrency production serving with an OpenAI-compatible API |
| SGLang | ●●●●● | ●●●●● | ●●●●● | NVIDIA GPU, AMD GPU | FP8, INT8, INT4 (AWQ), GPTQ | Complex prompt pipelines, structured output and multi-turn prefix reuse |
| TensorRT-LLM | ●●●●● | ●●●●● | ●●●●● | NVIDIA GPU | FP8, INT8, INT4 (AWQ), NVFP4 | NVIDIA-only shops chasing maximum performance with engineering capacity |
| llama.cpp | ●●●●● | ●●●●● | ●●●●● | CPU, NVIDIA GPU, AMD GPU, Apple Silicon | GGUF Q2, GGUF Q4, GGUF Q5, GGUF Q6, GGUF Q8 | Local / edge devices, personal computers, CPU or Mac inference |
Feature Matrix
✓ Supported · ~ Partial · ✗ Not supported
| Feature | vLLM | SGLang | TensorRT-LLM | llama.cpp |
|---|---|---|---|---|
| Continuous batching | ✓ | ✓ | ✓ | ~ |
| OpenAI-compatible API | ✓ | ✓ | ~ | ✓ |
| Prefix caching | ✓ | ✓ | ✓ | ~ |
| Structured / constrained output (JSON) | ✓ | ✓ | ~ | ✓ |
| Tensor parallelism (multi-GPU) | ✓ | ✓ | ✓ | ~ |
| Speculative decoding | ✓ | ✓ | ✓ | ✓ |
| Multimodal (vision) models | ✓ | ✓ | ✓ | ~ |
| LoRA adapter hot-loading | ✓ | ✓ | ✓ | ~ |
| CPU-only inference | ~ | ✗ | ✗ | ✓ |
| Apple Silicon (Metal) | ✗ | ✗ | ✗ | ✓ |
| AMD GPU (ROCm) | ✓ | ✓ | ✗ | ✓ |
| Single pip install | ✓ | ✓ | ✗ | ✓ |
Decision Flowchart
Start: where will you run the model?
│
├─ CPU only / Mac / edge device?
│ └─ Yes → llama.cpp (GGUF quantization, single binary)
│
└─ GPU server available?
│
├─ Non-NVIDIA (AMD ROCm)?
│ └─ Yes → vLLM or SGLang (TensorRT-LLM is NVIDIA-only)
│
└─ NVIDIA GPU?
│
├─ Need maximum performance and can invest engineering effort?
│ └─ Yes → TensorRT-LLM (with Triton Inference Server)
│
└─ No → what is the main workload?
│
├─ Multi-turn chat / structured output / heavy prefix reuse
│ └─ → SGLang (RadixAttention advantage)
│
└─ General high-concurrency API serving
└─ → vLLM (most mature ecosystem, safe default)