[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:freetoken":3},"\u003Cdiv align=\"center\">\n  \u003Cpicture>\n    \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFlashML-org\u002FFreeToken\u002Fmain\u002Fassets\u002Ffreetoken-logo-dark.svg\">\u003C\u002Fsource>\n    \u003Csource media=\"(prefers-color-scheme: light)\" srcset=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFlashML-org\u002FFreeToken\u002Fmain\u002Fassets\u002Ffreetoken-logo-light.svg\">\u003C\u002Fsource>\n    \u003Cimg alt=\"FreeToken\" src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFlashML-org\u002FFreeToken\u002Fmain\u002Fassets\u002Ffreetoken-logo.svg\" width=\"65%\" \u002F>\n  \u003C\u002Fpicture>\n\u003C\u002Fdiv>\u003Cp align=\"center\">\n| \u003Ca href=\"https:\u002F\u002Fwww.flashml.ai\u002F\" rel=\"nofollow ugc noopener\">\u003Cb>Download\u003C\u002Fb>\u003C\u002Fa> | \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.16157\" rel=\"nofollow ugc noopener\">\u003Cb>Paper\u003C\u002Fb>\u003C\u002Fa> | \u003Ca href=\"https:\u002F\u002Fjoin.slack.com\u002Ft\u002Fflashml\u002Fshared_invite\u002Fzt-3zpdh5j10-9dwTXrgLiqpVxizhA9KVbA\" rel=\"nofollow ugc noopener\">\u003Cb>Developer Slack\u003C\u002Fb>\u003C\u002Fa> | \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002FxzwSnMdsX\" rel=\"nofollow ugc noopener\">\u003Cb>Community Discord\u003C\u002Fb>\u003C\u002Fa> | \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken\u002Fblob\u002Fmain\u002Fassets\u002Ffreetoken-wechatgroup.png\" rel=\"nofollow ugc noopener\">\u003Cb>Community WeChat\u003C\u002Fb>\u003C\u002Fa> |\n\u003C\u002Fp>\u003Cp>Unlock datacenter-class intelligence on the hardware you already own — Run 290B+ frontier MoE models locally on your gaming PC at blistering interactive speeds.\u003C\u002Fp>\n\u003Ch2>About\u003C\u002Fh2>\n\u003Cp>FreeToken is an edge-native Mixture-of-Experts (MoE) serving engine designed for running frontier-scale open-weight models on personal and consumer hardware. It treats heterogeneous edge resources—GPUs, CPUs, host memory, and interconnects—as a unified, elastic inference platform. Its core features include:  \u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Fast Edge-Native Runtime\u003C\u002Fstrong>: Provides efficient MoE serving with bandwidth-adaptive CPU–GPU co-execution ($q^\\star$ policy), full-layer double-buffered prefill streaming, global LRU expert caching, graph-compatible execution, and the FTW fast weight format.  \u003C\u002Fli>\n\u003Cli>\u003Cstrong>Semantic-Aware Caching\u003C\u002Fstrong>: Features semantic anchor checkpoints for recurrent state and KV caches, allowing agentic context edits (e.g., tool calls, thinking blocks) to avoid redundant context recomputation.  \u003C\u002Fli>\n\u003Cli>\u003Cstrong>Elastic Memory Management\u003C\u002Fstrong>: Supports dynamic, runtime VRAM re-allocation between expert caches and KV memory without engine restarts or weight reloading.  \u003C\u002Fli>\n\u003Cli>\u003Cstrong>Broad MoE &amp; Ecosystem Support\u003C\u002Fstrong>: Supports frontier open-weight MoE models (e.g., DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2) across various parameter scales and quantization formats (e.g., MXFP4, NVFP4, FP8, BF16), with Anthropic\u002FOpenAI-compatible APIs for seamless integration with real-world coding and tool-calling agents (e.g., Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness). \u003C\u002Fli>\n\u003Cli>\u003Cstrong>Diverse Consumer Hardware\u003C\u002Fstrong>: Scales across consumer laptops, gaming desktops, and workstation GPUs, with native support for NVIDIA RTX 30, RTX 40, and RTX 50 series GPUs.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Getting Started\u003C\u002Fh2>\n\u003Ch3>Desktop app\u003C\u002Fh3>\n\u003Cp>Download FreeToken for Windows or Linux at \u003Ca href=\"https:\u002F\u002Fwww.flashml.ai\u002F\" rel=\"nofollow ugc noopener\">flashml.ai\u003C\u002Fa>. It sets the engine up for you and gives you a GUI for running models, chatting, and tuning the engine.\u003C\u002Fp>\n\u003Cdiv align=\"center\">\n  \u003Cimg alt=\"FreeToken Desktop\" src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFlashML-org\u002FFreeToken\u002Fmain\u002Fassets\u002Fdesktop-console.png\" width=\"92%\" \u002F>\n\u003C\u002Fdiv>\u003Ch3>CLI\u003C\u002Fh3>\n\u003Cp>Install FreeToken with \u003Ca href=\"https:\u002F\u002Fdocs.astral.sh\u002Fuv\u002F\" rel=\"nofollow ugc noopener\">uv\u003C\u002Fa> (recommended) or pip:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">uv pip install \"freetoken[accel]\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Or build from source:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">git clone https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken.git &amp;&amp; cd FreeToken\nuv venv &amp;&amp; source .venv\u002Fbin\u002Factivate\nuv pip install -e \".[accel]\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>For More details:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken\u002Fblob\u002Fmain\u002Fdocs\u002Finstall.md\" rel=\"nofollow ugc noopener\">Install FreeToken\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken\u002Fblob\u002Fmain\u002Fdocs\u002Fquickstart.md\" rel=\"nofollow ugc noopener\">Quick start\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken\u002Fblob\u002Fmain\u002Fdocs\u002Fmodels.md\" rel=\"nofollow ugc noopener\">Supported models\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFlashML-org\u002FFreeToken\u002Fblob\u002Fmain\u002Fdocs\u002Fcli.md\" rel=\"nofollow ugc noopener\">CLI reference\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Citation\u003C\u002Fh2>\n\u003Cp>If you use FreeToken for your research, please cite our \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.16157\" rel=\"nofollow ugc noopener\">paper\u003C\u002Fa>:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bibtex\">@article{yang2026freetoken,\n  title={FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution},\n  author={Yang, Shuo and Fan, Xiaoze and Pan, Melis\n\u003C\u002Fcode>\u003C\u002Fpre>\n",1787581366392]