[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:bigmoeonedge":3},"\u003Cp align=\"center\">\n  \u003Cpicture>\n    \u003Csource media=\"(prefers-color-scheme: dark)\" srcset=\"docs\u002Fassets\u002Flogo\u002Fhorizontal-ink.svg\">\u003C\u002Fsource>\n    \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FHelldez\u002Fbigmoeonedge\u002FHEAD\u002Fdocs\u002Fassets\u002Flogo\u002Fhorizontal.svg\" width=\"440\" alt=\"BigMoeOnEdge\" \u002F>\n  \u003C\u002Fpicture>\n\u003C\u002Fp>\u003Cp align=\"center\">\u003Cb>Run Mixture-of-Experts models bigger than your device's RAM. On a phone, on a PC, CPU only.\u003C\u002Fb>\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FHelldez\u002FBigMoeOnEdge\u002Freleases\u002Flatest\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Fv\u002Frelease\u002FHelldez\u002FBigMoeOnEdge\" alt=\"Latest release\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FHelldez\u002FBigMoeOnEdge\u002Factions\u002Fworkflows\u002Fci.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002FHelldez\u002FBigMoeOnEdge\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"CI\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FHelldez\u002Fbigmoeonedge\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Flicense\u002FHelldez\u002FBigMoeOnEdge\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003C\u002Fp>\u003Chr \u002F>\n\u003Cp>A Mixture-of-Experts model is made of many small \"experts\", and each generated token only uses a\nfew of them. BigMoeOnEdge takes that literally: it keeps the small always-needed part of the model\nat hand and reads just the experts each token asks for, directly from flash storage, at the moment\nthey are needed. The rest of the model stays on disk. That is what lets a model several times\nbigger than your RAM generate text on an ordinary phone, losslessly: the output is byte-identical\nto running the same model fully resident.\u003C\u002Fp>\n\u003Cp>It is built \u003Cstrong>on top of llama.cpp's public API\u003C\u002Fstrong>, not as a fork. Every quantization format,\ntokenizer and chat template llama.cpp supports works out of the box, because llama.cpp itself is\ndoing that part: MXFP4 and Q4_K_M stream through the same code. Supporting a new MoE architecture\nis one row in a registry, and following a new llama.cpp release is a routine submodule bump.\u003C\u002Fp>\n\u003Cp>The most extreme thing it can do today: \u003Cstrong>DeepSeek V4 Flash 0731\u003C\u002Fstrong>, a 284B-parameter MoE\n(~91 GB on disk at 2-bit expert quantization), generating on a phone with 12 GB of RAM at about\n\u003Cstrong>1 tok\u002Fs\u003C\u002Fstrong>. More than seven times more model than memory, streamed from flash as the three shard\nfiles Hugging Face ships, with no merge step and no PC in the loop.\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FHelldez\u002Fbigmoeonedge\u002FHEAD\u002Fdocs\u002Fassets\u002Fhero-dsv4.gif\" width=\"360\" alt=\"DeepSeek V4 Flash 0731 (284B, ~91 GB) generating in the demo app on a 12 GB phone, with live tok\u002Fs and telemetry\" \u002F>\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cem>DeepSeek V4 Flash 0731: 284B parameters, ~91 GB on disk, on a 12 GB phone.\n0.94 tok\u002Fs in the demo app, real time.\u003C\u002Fem>\u003C\u002Fp>\u003Cp>It is not one model, either. Below: three of them, one after another on the same phone, each past\nwhat it should be able to hold.\u003C\u002Fp>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Ff899b93f-c7c4-4ce9-9fb0-5ed1bae13761\" rel=\"nofollow ugc noopener\">https:\u002F\u002Fgithub.com\u002Fuser-attachments\u002Fassets\u002Ff899b93f-c7c4-4ce9-9fb0-5ed1bae13761\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cem>Left to right: gpt-oss-120b (~60 GB), Qwen3-30B-A3B (18.5 GB), Gemma-4-26B-A4B (17 GB),\nrecorded in the demo app on a 12 GB phone, real time, not sped up.\u003C\u002Fem>\u003C\u002Fp>\u003Ch2>Table of contents\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"#why-this-exists\" rel=\"nofollow ugc noopener\">Why this exists\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#try-it-on-your-phone\" rel=\"nofollow ugc noopener\">Try it on your phone\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#features\" rel=\"nofollow ugc noopener\">Features\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#supported-models\" rel=\"nofollow ugc noopener\">Supported models\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#benchmarks\" rel=\"nofollow ugc noopener\">Benchmarks\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#quickstart\" rel=\"nofollow ugc noopener\">Quickstart\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#how-it-works\" rel=\"nofollow ugc noopener\">How it works\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#documentation\" rel=\"nofollow ugc noopener\">Documentation\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#prior-art\" rel=\"nofollow ugc noopener\">Prior art\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#license\" rel=\"nofollow ugc noopener\">License\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Why this exists\u003C\u002Fh2>\n\u003Cp>The models people actually want to talk to keep growing faster than the RAM in the devices they\ncarry. MoE models offer a way out, because most of their weights sit idle on any given token, but\nevery mainstream runtime still insists on holding (or paging) the whole file in memory. So a 20 GB\nmodel on a 12 GB phone either refuses to load or crawls while the OS frantically swaps.\u003C\u002Fp>\n\u003Cp>BigMoeOnEdge treats flash storage as part of the memory hierarchy instead. Three situations where\nthat changes what your device can run:\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Models far past RAM.\u003C\u002Fstrong> A ~60 GB model on a 12 GB phone cannot be resident, full stop. Streamed,\nit runs at usable speed. This is the headline case, but it is not the only one.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Models just past RAM.\u003C\u002Fstrong> An 18 to 22 GB model on a 12 GB phone is where the ordinary way of\nloading (mmap) tur\u003C\u002Fp>\n",1787530081802]