vLLM
Free tierEasy, fast, and cheap LLM serving for everyone
Free tier available·All audiences·API available·Open source
Key strengths
State-of-the-art serving throughputPagedAttention for efficient KV memory managementContinuous batching and chunked prefillBroad quantization support (FP8, INT8, INT4, GPTQ/AWQ, GGUF, etc.)OpenAI-compatible API server with Anthropic Messages API and gRPC supportSupports 200+ model architectures on HuggingFaceDistributed inference with tensor, pipeline, data, expert, and context parallelismSpeculative decoding (n-gram, suffix, EAGLE)Multi-LoRA supportBroad hardware support (NVIDIA, AMD, x86/ARM/PowerPC CPUs, TPUs, Intel Gaudi, and more)
Free tier + paid plans
Berkeley, United States
Founded 2023
Self-hostable
No ratings yet
vLLM — Developer Documentation Notes
vLLM is open source and self-hostable, with an API layer designed for straightforward integration into existing inference pipelines.
API Surface
- OpenAI-compatible REST API — swap vLLM in as a drop-in replacement for OpenAI endpoints with no client-side code changes.
- Anthropic Messages API — native support for the Anthropic message format.
- gRPC — available for high-performance, low-overhead inter-service communication.
Deployment Options
vLLM can be deployed via:
- Docker — official container images for reproducible environments.
- Kubernetes — Helm charts available; integrates with KServe, KubeRay, KubeAI, and KAITO for orchestrated serving.
- Ray Serve / Anyscale — distributed serving on Ray clusters.
- Cloud platforms — SkyPilot, Modal, RunPod, Cerebrium, Nebius, Hugging Face Inference Endpoints, and more.
- NVIDIA Triton / NVIDIA Dynamo — for enterprise inference serving pipelines.
Model & Hardware Configuration
- Load any of 200+ Hugging Face model architectures directly.
- Configure quantization (FP8, INT8, INT4, GPTQ, AWQ, GGUF) at load time to match hardware constraints.
- Enable distributed inference by selecting parallelism strategy (tensor, pipeline, data, expert, or context) via configuration flags.
- Attach multiple LoRA adapters at runtime for multi-tenant fine-tuned serving.
Ecosystem Integrations
Works out of the box with LangChain, LlamaIndex, Haystack, AutoGen, LiteLLM, BentoML, Dify, AnythingLLM, Lobe Chat, Open WebUI, Streamlit, Llama Stack, AIBrix, dstack, and more.
