StackMap
Subscribe

airllm vs kimi-k3-in-c

Layer-by-layer inference that runs 70B models on a 4GB GPU — no quantization required; 405B on 8GB, DeepSeek-V3 671B on ~12GB. One AutoModel line for most open model families. — versus — Kimi K3 (2.78T-parameter MoE) inference in portable C99 on one CPU: the dense trunk stays in RAM, 4-bit experts stream from disk, and output is byte-identical from 8 GB to 224 GB.

The curated verdict

Both trade speed for memory to run models far larger than the machine; airllm loads layer by layer on a small GPU, kimi-k3-in-c streams 4-bit experts on CPU alone.

airllmkimi-k3-in-c
Stars35k8.8k
Forks3.7k1.4k
LanguageJupyter NotebookC
LicenseApache-2.0Apache-2.0
Last activity3 days ago9 days ago
Topicslocallocal
Curated connections83

airllm — the curator's take

The trick is elegant and the tradeoff is brutal, and you should know both: only one layer lives on the GPU at a time, so VRAM scales with layer size instead of model size — that's how 671B fits on a hobbyist card — but every token streams the whole model from disk, so generation runs at seconds-per-token. Use it for batch/offline jobs where 'it fits' beats 'it's fast', for poking at frontier-scale open models on hardware you own, or with block-wise 4/8-bit compression for a ~3x claw-back. NOT a chat or serving solution: Ollama is the fits-in-VRAM daily driver, vLLM the throughput server. Apple-silicon Macs supported via MLX. README carries the author's sponsor/affiliate links — the library stands on its own.

kimi-k3-in-c — the curator's take

Read it as the clearest from-scratch walkthrough of MoE memory layout around, and as proof of where the floor is: a 1.56 TB checkpoint answering correctly in 8.24 GB of RAM, because routed experts stay packed in 4-bit on disk and are multiplied in place. Part II of the README builds every component step by step. As a way to use Kimi K3 it is not practical: about 26 s per token on an 8 GB laptop and still 5.6 s with 128 GB+, CPU-only, one model, and you need 1.56 TB of fast NVMe for the checkpoint. For daily local work, run a model that fits your machine in Ollama; for big MoE at usable speed, colibri or freetoken.