StackMap
Subscribe

router vs semantic-router

Weave's Go model router: drop-in proxy speaking Anthropic, OpenAI and Gemini APIs that scores each request with an on-box ONNX embedder (Avengers-Pro clusters) and routes to the cheapest fit model. — versus — vLLM's programmable decision layer: one endpoint for your agent harness that selects or combines models per call by policy — quality, latency, cost, location — across local and cloud backends.

The curated verdict

Both are drop-in model routers that classify each request and send it to the best-fit model; Weave's is a single Go proxy with an on-box embedder, vLLM SR a policy-programmable cluster-scale layer.

routersemantic-router
Stars5.6k6.1k
Forks1531.0k
LanguageGoGo
LicenseApache-2.0Apache-2.0
Last activity3 days agoyesterday
Topicsgatewaygateway, local
Curated connections65

router — the curator's take

Choose it when the goal is *smart* routing, not just multi-provider fan-out: an in-process embedder classifies each action and picks a model per request in under 50ms, with a `/v1/route` endpoint to inspect the decision and OTLP traces out of the box. `npx @weave-os/router` wires Claude Code, Codex, opencode or pi in one step; self-hosting needs Postgres (and Pub/Sub for multi-replica). Don't reach for it if you want a plain gateway (litellm), free-tier pooling (omniroute), or to route through subscription CLIs (cliproxyapi) — and note the Elastic License v2, which is not OSI open source, and that the frictionless quickstart is the hosted service.

semantic-router — the curator's take

Use it when you serve several models — especially self-hosted vLLM next to cloud APIs — and want policy-driven per-request selection with guardrails, caching and hallucination checks in the path. Heavyweight for a single-provider app or a laptop: Kubernetes-grade infra with its own routing models. Just want quota failover for coding CLIs? That's a coding gateway's job, not this.