StackMap
Subscribe
Explore / vLLM
vllm-project

vLLM

High-throughput, memory-efficient inference and serving engine for LLMs.

92,864 22,758 Python Apache-2.0updated 3 days ago
View on GitHubDispute this mapping →

Continue your stack

What teams reach for next — and why each earns a place beside vLLM. Ranked by curator confidence.

pairs wellpairs wellpairs wellpairs wellalternativealternativealternativealternativebuilt withbuilt withbuilt withbuilt withModel-Optimizerlitellmspeech-to-speechai-avatar-systemFreeTokenOllamasieairllmLMCacheagentic-apillm-dkvcachedvLLM
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md1 min read

vLLM is a fast inference and serving engine for LLMs, using PagedAttention for high throughput — the production choice when you outgrow single-request local inference.

pip install vllm