StackMap
Subscribe

plano vs semantic-router

AI-native Envoy-based proxy for agentic apps: agent orchestration via a 4B routing model, smart LLM routing, guardrail filter chains and zero-code OTEL traces. Rust, framework-agnostic. — versus — vLLM's programmable decision layer: one endpoint for your agent harness that selects or combines models per call by policy — quality, latency, cost, location — across local and cloud backends.

The curated verdict

Both sit between agents and models with learned routing and guardrail chains; plano is an Envoy-based proxy with a 4B routing model, SR a vLLM-native decision layer with its own routing models.

planosemantic-router
Stars7.1k6.1k
Forks4881.0k
LanguageRustGo
LicenseApache-2.0Apache-2.0
Last activity4 days agoyesterday
Topicsgateway, orchestrationgateway, local
Curated connections95

plano — the curator's take

Reach for it when multi-agent code is drowning in hidden middleware — intent routing, provider quirks, guardrail hooks, tracing glue. Plano moves all of that out-of-process: agents are plain OpenAI-compatible HTTP servers in any language, orchestration is YAML plus a purpose-built 4B router model. NOT for a quick single-agent demo (it adds an infra hop and Envoy operational surface), and note the catch: the hosted Plano-Orchestrator LLM is free-tier only — production means running the routing models yourself or getting API keys. If you only need provider unification, LiteLLM is lighter.

semantic-router — the curator's take

Use it when you serve several models — especially self-hosted vLLM next to cloud APIs — and want policy-driven per-request selection with guardrails, caching and hallucination checks in the path. Heavyweight for a single-provider app or a laptop: Kubernetes-grade infra with its own routing models. Just want quota failover for coding CLIs? That's a coding gateway's job, not this.