StackMap
Subscribe
Explore / vLLM
vllm-project

vLLM

High-throughput, memory-efficient inference and serving engine for LLMs.

89,493 20,944 Python Apache-2.0updated today
View on GitHubDispute this mapping →
README.md

vLLM is a fast inference and serving engine for LLMs, using PagedAttention for high throughput — the production choice when you outgrow single-request local inference.

pip install vllm

Continue your stack

What teams reach for next — and why each earns a place beside vLLM. Ranked by curator confidence.