[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:colibri":3},"\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fassets\u002Fcolibri.svg\" width=\"500\" alt=\"colibrì — tiny engine, immense model\" \u002F>\n\u003C\u002Fp>\u003Cp align=\"center\">\n  English · \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FJustVugg\u002Fcolibri\u002Fblob\u002FHEAD\u002FREADME.zh-TW.md\" rel=\"nofollow ugc noopener\">繁體中文\u003C\u002Fa>\n\u003C\u002Fp>\u003Cp>\u003Cstrong>Tiny engine, immense model.\u003C\u002Fstrong> Run \u003Cstrong>GLM-5.2 (744B-parameter MoE)\u003C\u002Fstrong> on a consumer machine with ~25 GB of RAM — in pure C, with zero dependencies, by streaming experts from disk.\u003C\u002Fp>\n\u003Cp>Colibrì is a lightweight, quality-preserving MoE runtime that treats VRAM, RAM,\nand storage as one managed memory hierarchy. Insufficient fast memory may reduce\nspeed, but the default policy \u003Cstrong>never silently changes model precision or router\nsemantics\u003C\u002Fstrong>.\u003C\u002Fp>\n\u003Cpre>\u003Ccode>$ .\u002Fcoli chat\n  🐦 colibrì v1.0 — GLM-5.2 · 744B MoE · int4 · streaming CPU\n  ✓ ready in 32s · resident 9.9 GB\n  › ciao!\n  ◆ Ciao! 😊 Come posso aiutarti oggi?\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Ch2>See it running\u003C\u002Fh2>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Fcolibri-dashboard.png\" width=\"900\" alt=\"colibrì web dashboard — live metrics, hardware panel, expert tiers\" \u002F>\n\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cem>The web dashboard (\u003Ccode>.\u002Fcoli web\u003C\u002Fcode>): a 744B model at \u003Cstrong>4 tok\u002Fs, TTFT 1.6 s, disk 0\u003C\u002Fstrong> —\nfull expert residency on 6× RTX 5090, with live token metrics, the per-turn time breakdown,\nthe VRAM\u002FRAM\u002Fdisk tier bar and the live mini-brain in the corner.\u003C\u002Fem>\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Fcolibri-brain.png\" width=\"900\" alt=\"the Brain page — 19,456 experts as a live cortex\" \u002F>\n\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cem>The \u003Cstrong>Brain\u003C\u002Fstrong> page: all 19,456 experts as a living cortex — colour is the storage tier,\nbrightness is routing heat, and every expert routed in a turn flashes white. Hovering shows the expert's\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FJustVugg\u002Fcolibri\u002Fissues\u002F175\" rel=\"nofollow ugc noopener\">measured topic affinity\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Fcolibri-atlas.png\" width=\"900\" alt=\"the Atlas page — the measured expert atlas as a 3-D galaxy\" \u002F>\n\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Cem>The \u003Cstrong>Atlas\u003C\u002Fstrong> page: the \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FJustVugg\u002Fcolibri\u002Fissues\u002F175\" rel=\"nofollow ugc noopener\">measured expert atlas\u003C\u002Fa>\nas a 3-D galaxy — 13,260 characterised experts, 1,041 replicated specialists clustering by topic\n(poetry, law, Chinese, SQL…). Position is measured routing affinity, not a learned embedding. Drag to spin.\u003C\u002Fem>\u003C\u002Fp>\u003Ch2>The idea\u003C\u002Fh2>\n\u003Cp>A 744B Mixture-of-Experts model activates only ~40B parameters per token — and\nonly ~11 GB of those change from token to token (the routed experts):\u003C\u002Fp>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Fsparse.png\" width=\"880\" alt=\"only ~5.4% of parameters are active per token\" \u002F>\n\u003C\u002Fp>\u003Cp>So the model doesn't need to \u003Cem>fit\u003C\u002Fem> in fast memory — it needs to be \u003Cstrong>placed\u003C\u002Fstrong>:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>the \u003Cstrong>dense part\u003C\u002Fstrong> (attention, shared experts, embeddings — ~17B params) stays\n\u003Cstrong>resident in RAM at int4\u003C\u002Fstrong> (~9.9 GB);\u003C\u002Fli>\n\u003Cli>the \u003Cstrong>19,456 routed experts\u003C\u002Fstrong> (75 MoE layers × 256 + the MTP head, ~19 MB each\nat int4) live \u003Cstrong>on disk\u003C\u002Fstrong> (~370 GB) and are \u003Cstrong>streamed on demand\u003C\u002Fstrong>, with a\nper-layer LRU cache, a learned pinned hot-store, and an optional VRAM tier.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>The engine is a single C file (\u003Ccode>c\u002Fglm.c\u003C\u002Fcode>) plus small headers. No BLAS, no Python\nat runtime, no GPU required.\u003C\u002Fp>\n\u003Ch2>How it works\u003C\u002Fh2>\n\u003Ch3>The per-token path\u003C\u002Fh3>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Ftoken-path.png\" width=\"880\" alt=\"route → union → place → overlap → learn\" \u002F>\n\u003C\u002Fp>\u003Cp>Every layer of every token walks the same five steps. The design goal is that\n\u003Cstrong>placement only ever decides speed\u003C\u002Fstrong> — the router's decisions and the weights'\nprecision are the same whether an expert answered from VRAM or from disk.\u003C\u002Fp>\n\u003Ch3>One memory hierarchy instead of one memory requirement\u003C\u002Fh3>\n\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FJustVugg\u002Fcolibri\u002FHEAD\u002Fdocs\u002Fmedia\u002Ftiers.png\" width=\"880\" alt=\"VRAM \u002F RAM \u002F NVMe three-tier expert residency\" \u002F>\n\u003C\u002Fp>\u003Cp>The same engine spans the whole range: on a 25 GB laptop everything streams from\ndisk (slow but correct); on a large host the entire expert set becomes resident\n(\u003Ccode>CUDA_EXPERT_GB=auto PIN_GB=all\u003C\u002Fcode>) and disk drops out of the decode path\nentirely. Between the tiers sits a \u003Cstrong>learning cache\u003C\u002Fstrong>: the engine records which\nexperts \u003Cem>your\u003C\u002Fem> workload routes to (`.coli_usa\u003C\u002Fp>\n",1784564567516]