StackMap
Subscribe
Explore / semantic-router
vllm-project

semantic-router

vLLM's programmable decision layer: one endpoint for your agent harness that selects or combines models per call by policy — quality, latency, cost, location — across local and cloud backends.

6,075 1,023 Go Apache-2.0updated yesterday
View on GitHubDispute this mapping →
Curator's take

Use it when you serve several models — especially self-hosted vLLM next to cloud APIs — and want policy-driven per-request selection with guardrails, caching and hallucination checks in the path. Heavyweight for a single-provider app or a laptop: Kubernetes-grade infra with its own routing models. Just want quota failover for coding CLIs? That's a coding gateway's job, not this.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside semantic-router. Ranked by curator confidence.

pairs wellpairs wellalternativealternativealternativevLLMllm-drouterplanoLLMRoutersemantic-router
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md1 min read
vLLM Semantic Router

An open, programmable decision layer for models and compute.

Documentation | Playground | Blog | Publications | Hugging Face | Slack

vllm-project%2Fsemantic-router | Trendshift Decision 2.0 — #1 on Hugging Face Trending Collections, October 6, 2026

Main GitHub Release Go Ask DeepWiki


About

Intelligence beyond any one model.

Give your agent harness one API for many models. vLLM Semantic Router selects or combines models for each call, guided by your policy.

Your harness keeps the agent loop, tools, and task state. The Router chooses among configured backends across local, private, and cloud compute.

Dimension Fragmented today With vLLM SR
Models Different models excel at different tasks. Select or combine models.
Compute Hardware varies in speed and capacity. Choose among configured backends.
Location Edge, private, and cloud. Keep calls within approved locations.
Preference Priorities change by task. Set quality, latency, and cost priorities.

Explore how it works →

Getting Started

Install

curl -fsSL https://vllm-sr.ai/install.sh | bash -s -- --channel stable

For pip, uv, or agent-driven installation, see the Installation Guide.

Connect your agent harness

Point your harness at the Router's inference endpoint. Use a public model ID such as vllm-sr/auto.

Follow Connect an agent harness for setup and compatibility.

Online playground

Try the online playground at https://app.vllm-sr.ai/playground.

Credentials:

  • Username: love@vllm-sr.ai
  • Password: vllm-sr-read

Latest News