[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:kimi-k3-in-c":3},"\u003Cdiv align=\"center\">\u003Ch1>kimi-k3-in-c\u003C\u002Fh1>\u003Ch3>A 2.78-trillion-parameter model. One CPU. 8 GB of RAM.\u003C\u002Fh3>\u003Cp>Kimi K3 inference in portable C99.\u003Cbr \u002F>No BLAS. No framework. No GPU.\u003C\u002Fp>\u003Cp>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Factions\u002Fworkflows\u002Fci.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fgithub\u002Factions\u002Fworkflow\u002Fstatus\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Fci.yml?branch=main&amp;style=flat-square&amp;label=CI\" alt=\"CI\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Fblob\u002FHEAD\u002FLICENSE\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-Apache--2.0-blue?style=flat-square\" alt=\"License\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Fblob\u002FHEAD\u002FMakefile\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FC99-portable-lightgrey?style=flat-square\" alt=\"C99\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"#requirements\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fplatform-Linux%20x86--64-lightgrey?style=flat-square\" alt=\"Platform\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Fblob\u002FHEAD\u002FCHANGELOG.md\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fversion-1.0.0-brightgreen?style=flat-square\" alt=\"Version\" \u002F>\u003C\u002Fa>\n\u003C\u002Fp>\u003Ctable>\n\u003Ctr>\n\u003Ctd align=\"center\">\u003Cb>2.78T\u003C\u002Fb>\u003Cbr \u002F>\u003Csub>parameters\u003C\u002Fsub>\u003C\u002Ftd>\n\u003Ctd align=\"center\">\u003Cb>1.56 TB\u003C\u002Fb>\u003Cbr \u002F>\u003Csub>checkpoint on disk\u003C\u002Fsub>\u003C\u002Ftd>\n\u003Ctd align=\"center\">\u003Cb>8.24 GB\u003C\u002Fb>\u003Cbr \u002F>\u003Csub>peak RSS, measured\u003C\u002Fsub>\u003C\u002Ftd>\n\u003Ctd align=\"center\">\u003Cb>176 KB\u003C\u002Fb>\u003Cbr \u002F>\u003Csub>the whole engine\u003C\u002Fsub>\u003C\u002Ftd>\n\u003Ctd align=\"center\">\u003Cb>0\u003C\u002Fb>\u003Cbr \u002F>\u003Csub>GPUs\u003C\u002Fsub>\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftable>\u003Cp>\u003Cb>The same 2.78-trillion-parameter model, the same answer, on whatever machine you own.\u003C\u002Fb>\u003Cbr \u002F>More memory only buys speed:\u003C\u002Fp>\u003Ctable>\n\u003Ctr>\n\u003Cth align=\"left\">the machine you have\u003C\u002Fth>\n\u003Cth align=\"right\">RAM\u003C\u002Fth>\n\u003Cth align=\"right\">time per token\u003C\u002Fth>\n\u003Cth align=\"left\">what is going on\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd align=\"left\">an ordinary laptop\u003C\u002Ftd>\n\u003Ctd align=\"right\">8 GB\u003C\u002Ftd>\n\u003Ctd align=\"right\">\u003Cb>26.5 s\u003C\u002Fb>\u003C\u002Ftd>\n\u003Ctd>the whole model streams off the disk on every step\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd align=\"left\">a high-end laptop\u003C\u002Ftd>\n\u003Ctd align=\"right\">32 GB\u003C\u002Ftd>\n\u003Ctd align=\"right\">\u003Cb>24.2 s\u003C\u002Fb>\u003C\u002Ftd>\n\u003Ctd>some of the model now sits in memory\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd align=\"left\">a desktop\u003C\u002Ftd>\n\u003Ctd align=\"right\">64 GB\u003C\u002Ftd>\n\u003Ctd align=\"right\">\u003Cb>19.8 s\u003C\u002Fb>\u003C\u002Ftd>\n\u003Ctd>more of it sits in memory\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd align=\"left\">a heavy workstation\u003C\u002Ftd>\n\u003Ctd align=\"right\">128 GB+\u003C\u002Ftd>\n\u003Ctd align=\"right\">\u003Cb>5.6 s\u003C\u002Fb>\u003C\u002Ftd>\n\u003Ctd>the model fits entirely in memory, the disk wait is gone\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftable>\u003Cp>\u003Csub>Same short prompt at every size, and the output is \u003Cb>byte-identical\u003C\u002Fb> from the smallest machine to the largest; only the clock changes. One machine, 124 cores, fast NVMe drive: the first three rows still read the model from disk each step, so a slower drive is slower there, while the 128 GB+ row keeps everything in memory and no longer waits on the disk. On that same machine v1.0.0 made the math per token about \u003Cb>8×\u003C\u002Fb> lighter, a follow-up question in a chat \u003Cb>3.9×\u003C\u002Fb> faster, and long prompts about \u003Cb>half\u003C\u002Fb> as costly. (A token is roughly a short word-piece; the two runnable demos below are the original captures on a slower drive, so their clock reads a little higher.) Full data in \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002Fblob\u002FHEAD\u002Fdocs\u002Fdata\u002F\" rel=\"nofollow ugc noopener\">docs\u002Fdata\u002F\u003C\u002Fa>.\u003C\u002Fsub>\u003C\u002Fp>\n\u003Chr \u002F>\u003Cp>\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFareedKhan-dev\u002Fkimi-k3-in-c\u002FHEAD\u002Fdocs\u002Fimages\u002Fpatrick_pray.png\" height=\"44\" align=\"middle\" alt=\"\" \u002F>\n  \u003Ci>I am open to AI research roles and PhD positions. \u003Ca href=\"https:\u002F\u002Fdrive.google.com\u002Ffile\u002Fd\u002F1yW5xHDS6Mr9ByrkCgVve85OqF4UOPv9K\u002Fview?usp=sharing\" rel=\"nofollow ugc noopener\">CV\u003C\u002Fa>.\u003C\u002Fi>\n\u003C\u002Fp>\u003Chr \u002F>\u003C\u002Fdiv>\u003Cbr \u002F>\u003Cpre>\u003Ccode class=\"language-console\">$ .\u002Fbin\u002Fk3 ~\u002Fk3model --trunk ~\u002Fk3trunk --preset laptop \\\n           --tok ~\u002Fk3model --prompt \"The capital of France is\" --gen 8 --incremental\n\n--- generated text ---\n Paris.\",\n+            \"The Eiffel\n----------------------\n8 tokens in 261.5 s, 32.69 s\u002Ftoken average\nPEAK RSS for the whole run: 8.24 GB\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Slow, and answering correctly, in 8.24 GB, from a checkpoint of 1.56 TB. This particular\nbatch command deliberately asks for a raw continuation. The official Kimi K3 checkpoint\nis also chat-capable; use the XTML REPL below when you want answers and multi-turn history.\nGive the same batch request more memory and the answer does not change, only the clock:\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-console\">$ .\u002Fbin\u002Fk3 ~\u002F\n\u003C\u002Fcode>\u003C\u002Fpre>\n",1790887801719]