vibehacker

vLLM

High-throughput, memory-efficient LLM inference and serving

by vLLM communityCoding & Dev ToolsAgents & Automation Jun 1, 2023
vLLM cover
vLLM screenshot 2vLLM screenshot 3vLLM screenshot 4

About vLLM

vLLM is an inference and serving engine for large language models (LLMs). It is intended for developers and teams deploying models across different hardware platforms.

The engine provides a drop-in OpenAI-compatible API and supports a range of hardware, including NVIDIA CUDA GPUs, AMD ROCm GPUs, AWS Neuron accelerators, Google Cloud TPUs, Apple Silicon, and others. PagedAttention, advanced scheduling, and continuous batching are used to maximize throughput and GPU utilization.

vLLM can be installed with Python or Docker. The quick-start instructions require Python 3.10 or newer, with Python 3.12 or newer recommended; stable and nightly builds are available.

Used vLLM?

Log in to write a review.

Similar tools

View all