colibri vs kimi-k3-in-c
Pure-C, zero-dep MoE runtime that runs GLM-5.2 (744B) on a 25GB-RAM consumer box by streaming experts from disk — VRAM/RAM/NVMe as one tiered hierarchy, never touching precision. — versus — Kimi K3 (2.78T-parameter MoE) inference in portable C99 on one CPU: the dense trunk stays in RAM, 4-bit experts stream from disk, and output is byte-identical from 8 GB to 224 GB.
Both are pure-C runtimes that stream MoE experts from disk to run frontier models in consumer RAM; colibri targets GLM-5.2 (744B) at usable speed with a VRAM/RAM/NVMe hierarchy, kimi-k3-in-c pushes a 2.78T model into 8 GB, CPU only.
| colibri | kimi-k3-in-c | |
|---|---|---|
| Stars | 38k | 8.8k |
| Forks | 4.2k | 1.4k |
| Language | C | C |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 3 days ago | 9 days ago |
| Topics | local | local |
| Curated connections | 6 | 3 |
colibri — the curator's take
Use it when you want a frontier-scale open MoE (GLM-5.2, 744B) on hardware that can't hold it, and you care about fidelity: forward pass is token-exact against the transformers oracle, and placement only ever changes speed, never semantics. The learning cache pins the experts YOUR workload actually routes, so it genuinely gets faster with use. NOT a general runner — it's a one-model engine; for everyday multi-model local use pick ollama. And expect 0.05–6.8 tok/s depending on tier residency: this is patience-ware for correctness fanatics, not latency-sensitive serving.
kimi-k3-in-c — the curator's take
Read it as the clearest from-scratch walkthrough of MoE memory layout around, and as proof of where the floor is: a 1.56 TB checkpoint answering correctly in 8.24 GB of RAM, because routed experts stay packed in 4-bit on disk and are multiplied in place. Part II of the README builds every component step by step. As a way to use Kimi K3 it is not practical: about 26 s per token on an 8 GB laptop and still 5.6 s with 128 GB+, CPU-only, one model, and you need 1.56 TB of fast NVMe for the checkpoint. For daily local work, run a model that fits your machine in Ollama; for big MoE at usable speed, colibri or freetoken.