StackMap
Subscribe
Explore / kimi-k3-in-c
FareedKhan-dev

kimi-k3-in-c

Kimi K3 (2.78T-parameter MoE) inference in portable C99 on one CPU: the dense trunk stays in RAM, 4-bit experts stream from disk, and output is byte-identical from 8 GB to 224 GB.

8,831 1,435 C Apache-2.0updated 9 days ago
View on GitHubDispute this mapping →
Curator's take

Read it as the clearest from-scratch walkthrough of MoE memory layout around, and as proof of where the floor is: a 1.56 TB checkpoint answering correctly in 8.24 GB of RAM, because routed experts stay packed in 4-bit on disk and are multiplied in place. Part II of the README builds every component step by step. As a way to use Kimi K3 it is not practical: about 26 s per token on an 8 GB laptop and still 5.6 s with 128 GB+, CPU-only, one model, and you need 1.56 TB of fast NVMe for the checkpoint. For daily local work, run a model that fits your machine in Ollama; for big MoE at usable speed, colibri or freetoken.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside kimi-k3-in-c. Ranked by curator confidence.

alternativealternativealternativecolibriBigMoeOnEdgeairllmkimi-k3-in-c
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md2 min read

kimi-k3-in-c

A 2.78-trillion-parameter model. One CPU. 8 GB of RAM.

Kimi K3 inference in portable C99.
No BLAS. No framework. No GPU.

CI License C99 Platform Version

2.78T
parameters
1.56 TB
checkpoint on disk
8.24 GB
peak RSS, measured
176 KB
the whole engine
0
GPUs

The same 2.78-trillion-parameter model, the same answer, on whatever machine you own.
More memory only buys speed:

the machine you have RAM time per token what is going on
an ordinary laptop 8 GB 26.5 s the whole model streams off the disk on every step
a high-end laptop 32 GB 24.2 s some of the model now sits in memory
a desktop 64 GB 19.8 s more of it sits in memory
a heavy workstation 128 GB+ 5.6 s the model fits entirely in memory, the disk wait is gone

Same short prompt at every size, and the output is byte-identical from the smallest machine to the largest; only the clock changes. One machine, 124 cores, fast NVMe drive: the first three rows still read the model from disk each step, so a slower drive is slower there, while the 128 GB+ row keeps everything in memory and no longer waits on the disk. On that same machine v1.0.0 made the math per token about 8× lighter, a follow-up question in a chat 3.9× faster, and long prompts about half as costly. (A token is roughly a short word-piece; the two runnable demos below are the original captures on a slower drive, so their clock reads a little higher.) Full data in docs/data/.


I am open to AI research roles and PhD positions. CV.



$ ./bin/k3 ~/k3model --trunk ~/k3trunk --preset laptop \
           --tok ~/k3model --prompt "The capital of France is" --gen 8 --incremental

--- generated text ---
 Paris.",
+            "The Eiffel
----------------------
8 tokens in 261.5 s, 32.69 s/token average
PEAK RSS for the whole run: 8.24 GB

Slow, and answering correctly, in 8.24 GB, from a checkpoint of 1.56 TB. This particular batch command deliberately asks for a raw continuation. The official Kimi K3 checkpoint is also chat-capable; use the XTML REPL below when you want answers and multi-turn history. Give the same batch request more memory and the answer does not change, only the clock:

$ ./bin/k3 ~/