BigMoeOnEdge vs kimi-k3-in-c
Run MoE models bigger than your RAM: keep the always-needed weights resident and stream each token's experts from flash — a 284B model on a 12 GB phone, CPU only, byte-identical output. — versus — Kimi K3 (2.78T-parameter MoE) inference in portable C99 on one CPU: the dense trunk stays in RAM, 4-bit experts stream from disk, and output is byte-identical from 8 GB to 224 GB.
Same insight, resident trunk plus experts streamed from flash with byte-identical output; bigmoeonedge builds on llama.cpp so many models work, kimi-k3-in-c is a from-scratch engine for one model.
| BigMoeOnEdge | kimi-k3-in-c | |
|---|---|---|
| Stars | 595 | 8.8k |
| Forks | 63 | 1.4k |
| Language | C++ | C |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 3 days ago | 9 days ago |
| Topics | local | local |
| Curated connections | 7 | 3 |
BigMoeOnEdge — the curator's take
The clever insight is that a Mixture-of-Experts model already refuses to use most of itself per token, so flash storage can be part of the memory hierarchy instead of a swap disaster. Built on llama.cpp's public API rather than as a fork, so every quantization, tokenizer and chat template works because llama.cpp is still doing that part, and following upstream is a submodule bump. The underrated case isn't the 284B stunt: it's the 18-22 GB model on a 12 GB phone, where ordinary mmap becomes a fault storm that kills other apps, and streaming turns it into a stable 5 tok/s. Set expectations: ~1 tok/s at the extreme, CPU-only with no GPU or NPU path, MoE architectures only, and a 91 GB download is still a 91 GB download.
kimi-k3-in-c — the curator's take
Read it as the clearest from-scratch walkthrough of MoE memory layout around, and as proof of where the floor is: a 1.56 TB checkpoint answering correctly in 8.24 GB of RAM, because routed experts stay packed in 4-bit on disk and are multiplied in place. Part II of the README builds every component step by step. As a way to use Kimi K3 it is not practical: about 26 s per token on an 8 GB laptop and still 5.6 s with 128 GB+, CPU-only, one model, and you need 1.56 TB of fast NVMe for the checkpoint. For daily local work, run a model that fits your machine in Ollama; for big MoE at usable speed, colibri or freetoken.