Throughput king. Cut our inference bill almost in half.

vLLM
High-throughput, memory-efficient LLM inference and serving
About vLLM
vLLM is an inference and serving engine for large language models (LLMs). It is intended for developers and teams deploying models across different hardware platforms.
The engine provides a drop-in OpenAI-compatible API and supports a range of hardware, including NVIDIA CUDA GPUs, AMD ROCm GPUs, AWS Neuron accelerators, Google Cloud TPUs, Apple Silicon, and others. PagedAttention, advanced scheduling, and continuous batching are used to maximize throughput and GPU utilization.
vLLM can be installed with Python or Docker. The quick-start instructions require Python 3.10 or newer, with Python 3.12 or newer recommended; stable and nightly builds are available.
Fast and boring in the best way. PagedAttention still magic.
Similar tools
ChatGPT
— AI assistant for answers, writing, images, work, and codeAI assistant for answers, writing, images, work, and code

Claude
— AI assistant for research, coding, writing, and everyday workAI assistant for research, coding, writing, and everyday work

GitHub Copilot
— AI coding assistance across editors, GitHub, and the CLIAI coding assistance across editors, GitHub, and the CLI
Cursor
— AI coding agent for building softwareAI coding agent for building software

Claude Code
— An AI coding agent for terminal, IDE, web, and SlackAn AI coding agent for terminal, IDE, web, and Slack
Grok
— AI assistant for chat, image creation, coding, and web answersAI assistant for chat, image creation, coding, and web answers