router vs semantic-router
Weave's Go model router: drop-in proxy speaking Anthropic, OpenAI and Gemini APIs that scores each request with an on-box ONNX embedder (Avengers-Pro clusters) and routes to the cheapest fit model. — versus — vLLM's programmable decision layer: one endpoint for your agent harness that selects or combines models per call by policy — quality, latency, cost, location — across local and cloud backends.
Both are drop-in model routers that classify each request and send it to the best-fit model; Weave's is a single Go proxy with an on-box embedder, vLLM SR a policy-programmable cluster-scale layer.
| router | semantic-router | |
|---|---|---|
| Stars | 5.6k | 6.1k |
| Forks | 153 | 1.0k |
| Language | Go | Go |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 3 days ago | yesterday |
| Topics | gateway | gateway, local |
| Curated connections | 6 | 5 |
router — the curator's take
Choose it when the goal is *smart* routing, not just multi-provider fan-out: an in-process embedder classifies each action and picks a model per request in under 50ms, with a `/v1/route` endpoint to inspect the decision and OTLP traces out of the box. `npx @weave-os/router` wires Claude Code, Codex, opencode or pi in one step; self-hosting needs Postgres (and Pub/Sub for multi-replica). Don't reach for it if you want a plain gateway (litellm), free-tier pooling (omniroute), or to route through subscription CLIs (cliproxyapi) — and note the Elastic License v2, which is not OSI open source, and that the frictionless quickstart is the hosted service.
semantic-router — the curator's take
Use it when you serve several models — especially self-hosted vLLM next to cloud APIs — and want policy-driven per-request selection with guardrails, caching and hallucination checks in the path. Heavyweight for a single-provider app or a laptop: Kubernetes-grade infra with its own routing models. Just want quota failover for coding CLIs? That's a coding gateway's job, not this.