mesh-llm vs Personal-AI-Router
Distributed LLM inference in Rust: pool GPUs across machines into one OpenAI-compatible endpoint — local fit first, mesh routing, and stage splits for models too large for any single box. — versus — NVIDIA's local inference router: pairs home machines running Ollama or LM Studio behind OpenAI-, Anthropic- and Ollama-compatible endpoints, sending each request to the best available node.
Both turn several machines into one local endpoint. mesh-llm pools GPUs and splits model stages across boxes to fit bigger models; PAIR routes whole requests to one node and never shards.
| mesh-llm | Personal-AI-Router | |
|---|---|---|
| Stars | 3.4k | 1.5k |
| Forks | 420 | 257 |
| Language | Rust | Go |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 4 days ago | 5 days ago |
| Topics | local | local, gateway |
| Curated connections | 7 | 4 |
mesh-llm — the curator's take
The 'LLM for the people' play: friends or a homelab pool mid-range GPUs and serve models none of them could run alone, with public meshes discoverable via Nostr. The Skippy stage-split design is genuinely clever. Experimental distributed systems — expect rough edges, and never treat a public mesh as private infrastructure. One box that fits your model? Just run Ollama.
Personal-AI-Router — the curator's take
Use PAIR when you own two or more capable machines and your local workload is concurrent — multi-agent runs, parallel coding sessions — and you want one endpoint that fans requests out by model availability and load, with a desktop app that installs Ollama or LM Studio for you. The Anthropic Messages endpoint means clients that speak Claude's API can hit local models too. Know the limit it states itself: each request runs on one node — no pooled VRAM, no sharding a model across boxes — so it won't let you run a model bigger than your biggest machine (that's mesh-llm's job). One machine? Just run Ollama. It also only manages Ollama and LM Studio, not vLLM or llama.cpp servers.