
An open, programmable decision layer for models and compute.
Documentation | Playground | Blog | Publications | Hugging Face | Slack
About
Intelligence beyond any one model.
Give your agent harness one API for many models. vLLM Semantic Router selects or combines models for each call, guided by your policy.
Your harness keeps the agent loop, tools, and task state. The Router chooses among configured backends across local, private, and cloud compute.
| Dimension | Fragmented today | With vLLM SR |
|---|---|---|
| Models | Different models excel at different tasks. | Select or combine models. |
| Compute | Hardware varies in speed and capacity. | Choose among configured backends. |
| Location | Edge, private, and cloud. | Keep calls within approved locations. |
| Preference | Priorities change by task. | Set quality, latency, and cost priorities. |
Getting Started
Install
curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable
For pip, uv, or agent-driven installation, see the Installation Guide.
Connect your agent harness
Point your harness at the Router's inference endpoint. Use a public model ID such as vllm-sr/auto.
Follow Connect an agent harness for setup and compatibility.
Online playground
Try the online playground at https://app.vllm-sr.ai/playground.
Credentials:
- Username:
love@vllm-sr.ai - Password:
vllm-sr-read
Latest News
- [2026/10/06] Decision 2.0 reached #1 on Hugging Face Trending Collections.
- [2026/10/06] Vela 2.0: Towards Open Foundation Routing Models
- [2026/09/24] vLLM Semantic Router v0.4 Hermes: Many Models, One Improving System
- [2026/09/22] Introducing Decision 1.0: Open Decision Foundation Models
- [2026/09/18] Introducing Vela 1.0
- [2026/08/24] Find Your Focus: How to Join and Work Together
- [2026/08/05] [LettuceDetect v2 in Semantic Router: Generative Hallucination Detection as a vLLM Endpoint](https://vllm-sr.ai/blog/lettucedetect-v2-generative-hallucinat