[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:mesh-llm":3},"\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FMesh-LLM\u002Fmesh-llm\u002FHEAD\u002Fdocs\u002Fmesh-llm-wordmark.png\" alt=\"Mesh LLM\" width=\"420\" \u002F>\n\u003C\u002Fp>\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FMesh-LLM\u002Fmesh-llm\u002FHEAD\u002Fmesh.png\" alt=\"Mesh LLM web console\" \u002F>\u003C\u002Fp>\n\u003Cp>Mesh LLM pools GPUs and memory across machines and exposes the result as one\nOpenAI-compatible API at \u003Ccode>http:\u002F\u002Flocalhost:9337\u002Fv1\u003C\u002Fcode>. Start one node, add more\nnodes later, and let the mesh decide whether a model runs locally, routes to a\npeer, or uses Skippy stage splits for models that are too large for one box.\u003C\u002Fp>\n\u003Ch2>Quick start\u003C\u002Fh2>\n\u003Cp>Install the latest release executable:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">curl -fsSL https:\u002F\u002Fraw.githubusercontent.com\u002FMesh-LLM\u002Fmesh-llm\u002Fmain\u002Finstall.sh | bash\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>On Windows, use PowerShell:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-powershell\">irm https:\u002F\u002Fraw.githubusercontent.com\u002FMesh-LLM\u002Fmesh-llm\u002Fmain\u002Finstall.ps1 | iex\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Finish setup:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">mesh-llm setup\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>On Windows PowerShell, use \u003Ccode>mesh-llm.exe setup\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>To remove an executable install later, preview the cleanup first:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">mesh-llm uninstall --dry-run\nmesh-llm uninstall --yes\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Uninstall preserves \u003Ccode>~\u002F.mesh-llm\u003C\u002Fcode> configuration and identity data unless you\nexplicitly pass \u003Ccode>--purge-config\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>Join the public mesh and start serving:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">mesh-llm serve --auto\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>That command chooses a backend flavor, downloads a suitable model if needed,\njoins the best discovered public mesh, starts the local API on port \u003Ccode>9337\u003C\u002Fcode>, and\nstarts the web console on port \u003Ccode>3131\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Cp>Check available models:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">curl -s http:\u002F\u002Flocalhost:9337\u002Fv1\u002Fmodels | jq '.data[].id'\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Send an OpenAI-compatible request:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">curl http:\u002F\u002Flocalhost:9337\u002Fv1\u002Fchat\u002Fcompletions \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\"model\":\"GLM-4.7-Flash-Q4_K_M\",\"messages\":[{\"role\":\"user\",\"content\":\"hello\"}]}'\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>For server deployments, add \u003Ccode>--headless\u003C\u002Fcode> to hide the web UI while keeping the\nmanagement API on the \u003Ccode>--console\u003C\u002Fcode> port:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">mesh-llm serve --auto --headless\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>Pick the workflow you need\u003C\u002Fh2>\n\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Goal\u003C\u002Fth>\n\u003Cth>Command\u003C\u002Fth>\n\u003Cth>Full guide\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Try the public mesh\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve --auto\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FMESHES.md\" rel=\"nofollow ugc noopener\">docs\u002FMESHES.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Start a private mesh\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve --model Qwen3-8B-Q4_K_M\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FMESHES.md\" rel=\"nofollow ugc noopener\">docs\u002FMESHES.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Publish your own mesh\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve --model Qwen3-8B-Q4_K_M --publish\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FMESHES.md\" rel=\"nofollow ugc noopener\">docs\u002FMESHES.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Join by invite token\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve --join &lt;token&gt;\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FMESHES.md\" rel=\"nofollow ugc noopener\">docs\u002FMESHES.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Run an API-only client\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm client --auto\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FMESHES.md\" rel=\"nofollow ugc noopener\">docs\u002FMESHES.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Run a big model with splits\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve --model hf:\u002F\u002Fmeshllm\u002F&lt;repo&gt;@&lt;rev&gt; --split\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FSKIPPY_SPLITS.md\" rel=\"nofollow ugc noopener\">docs\u002FSKIPPY_SPLITS.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Attach a Flash-MoE SSD backend\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm serve\u003C\u002Fcode> with \u003Ccode>[[plugin]] name = \"flash-moe\"\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002Fplugins\u002Fflash-moe.md\" rel=\"nofollow ugc noopener\">docs\u002Fplugins\u002Fflash-moe.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Fan out one prompt to every model in the mesh\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>curl ... -d '{\"model\":\"mesh\", ...}'\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002Fdesign\u002FMOA_GATEWAY.md\" rel=\"nofollow ugc noopener\">docs\u002Fdesign\u002FMOA_GATEWAY.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Use Goose, OpenCode, Claude Code, or Pi\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>mesh-llm goose\u003C\u002Fcode>, \u003Ccode>mesh-llm opencode\u003C\u002Fcode>, \u003Ccode>mesh-llm claude\u003C\u002Fcode>, \u003Ccode>mesh-llm pi\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002Fdocs\u002FAGENTS.md\" rel=\"nofollow ugc noopener\">docs\u002FAGENTS.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Build or contribute\u003C\u002Ftd>\n\u003Ctd>\u003Ccode>just build\u003C\u002Fcode>\u003C\u002Ftd>\n\u003Ctd>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMesh-LLM\u002Fmesh-llm\u002Fblob\u002FHEAD\u002FCONTRIBUTING.md\" rel=\"nofollow ugc noopener\">CONTRIBUTING.md\u003C\u002Fa>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\n\u003Ch2>How the mesh works\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Single-machine fit first.\u003C\u002Fstrong> If one node can host the full model, it serves\nthe model locally without stage traffic.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Mesh routing.\u003C\u002Fstrong> Every node exposes the same \u003Ccode>\u002Fv1\u003C\u002Fcode> API. Requests are routed\nby the \u003Ccode>model\u003C\u002Fcode> field to the peer that can serve that model.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Owner-control plane.\u003C\u002Fstrong> Operator config and inventory actions use an\nadditive \u003Ccode>mesh-llm-control\u002F1\u003C\u002Fcode> lane with explicit endpoint bootstrap, while\npublic mesh join, gossip, routing, and inference stay on the public mesh\nplane for mixed-version compatibility.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Skippy stage splits.\u003C\u002Fstrong> Large dense models can load as package-backed layer\nstages. The coordinator plans contiguous layer ranges, starts downstream\nstages first, waits for readiness, then publishes the stage-0 route.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Layer packages.\u003C\u002Fstrong> Package repositories contain \u003Ccode>model-package.json\u003C\u002Fcode> plus\nGGUF fragments so peers fetch only the pieces needed for their assigned stage.\u003C\u002Fli>\n\u003Cli>**Public disc\u003C\u002Fli>\n\u003C\u002Ful>\n",1784564568347]