plano vs semantic-router
AI-native Envoy-based proxy for agentic apps: agent orchestration via a 4B routing model, smart LLM routing, guardrail filter chains and zero-code OTEL traces. Rust, framework-agnostic. — versus — vLLM's programmable decision layer: one endpoint for your agent harness that selects or combines models per call by policy — quality, latency, cost, location — across local and cloud backends.
Both sit between agents and models with learned routing and guardrail chains; plano is an Envoy-based proxy with a 4B routing model, SR a vLLM-native decision layer with its own routing models.
| plano | semantic-router | |
|---|---|---|
| Stars | 7.1k | 6.1k |
| Forks | 488 | 1.0k |
| Language | Rust | Go |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 4 days ago | yesterday |
| Topics | gateway, orchestration | gateway, local |
| Curated connections | 9 | 5 |
plano — the curator's take
Reach for it when multi-agent code is drowning in hidden middleware — intent routing, provider quirks, guardrail hooks, tracing glue. Plano moves all of that out-of-process: agents are plain OpenAI-compatible HTTP servers in any language, orchestration is YAML plus a purpose-built 4B router model. NOT for a quick single-agent demo (it adds an infra hop and Envoy operational surface), and note the catch: the hosted Plano-Orchestrator LLM is free-tier only — production means running the routing models yourself or getting API keys. If you only need provider unification, LiteLLM is lighter.
semantic-router — the curator's take
Use it when you serve several models — especially self-hosted vLLM next to cloud APIs — and want policy-driven per-request selection with guardrails, caching and hallucination checks in the path. Heavyweight for a single-provider app or a laptop: Kubernetes-grade infra with its own routing models. Just want quota failover for coding CLIs? That's a coding gateway's job, not this.