mini-AGI
mini-AGI - is a continual learning byte-level language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads. It stores its weights as ordinary files on disk and pages them onto the card as it needs them, so the parameter count is bounded by free disk space rather than by VRAM. It grows new capacity while training when it runs short, prunes what nothing asks for, and reads through exactly the same code path it serves on. Targeted at a PC or laptop with at least an 8 GB VRAM GPU on the board.
NOTE: as of now this is a small toy-level model. Do not expect a frontier level capabilities. This is rather a small experiment to show, that continual learning from the single stream of data without catastrophic forgetting is possible. Furthermore it is possible on a modest hardware. Which means that almost everyone could train their own version of the model (or simply continue training this one) exactly as they see it fit. And the capabilities would be bounded by the actual hardware, scale and quality of the data available and the amount of time one willing to spend on training the model.
The name is a half-joke and not a statement of the current capabilities of the model, but rather the potential and traits it has. Continual single-stream learning and natural size and processing adaptation to the available resources exactly what I myself expect from an AGI to have. It has just a small toy-level memory footprint and hence it is “mini-AGI”.
Here is how min-run dashboard looks like. The model is pointed to the corpus to constantly read and learn from.
History - here is the samples from the whole training run history so far. You can inspect them yourself to see how the model improved over the course of training/reading the corpus.
The "final" weights are not published yet. The run is still reading its first pass over the corpus, the weights go up once it has been through all of it, which is several weeks away at the current rate. If you would like to play with undertrained weights as they are right now, you can find the most recent (1 oct 2026) snapshot here: Volotat/mini-AGI-undertrained
Graph of the whole run so far

Every sample round of the run to date: 1,717.1M characters over 2,462 evaluations.
Current quality of samples the model generates
The round with the lowest held-out loss so far - 0.6235 nats at 1,682.8M characters. Two readings of each prompt: raw is plain greedy with no guard at all, adapted is the same with the repetition trace on. The whole history is in runs/samples.txt.
==============================================================================
step 822,034 1682.8M of 7,880M characters (21.36%) 264 min 397 experts
context 4,096 characters of 4,096 reading 1,391 char/s writing 34.6 char/s still gaining +0.0018 deep into it
grad norm 1.67 against a clip of 1 clipping
train loss 0.5076 lr 1.84e-05 evidence t +2.49 over 65.7 (effect +0.0553) rate x0.062
held-out loss 0.6235 +/-0.0281 nats 0.8995 bits/char perplexity 1.87 gap +0.1159
arithmetic 0.620 chat 0.590 chat_hermes 0.875 chess 0.401 code 0.522 reasoning 0.540 stories 0.410 wikipedia 1.031
repeats 24% of 8-grams, greedy with no guard
==============================================================================
--- stories ---
prompt: 'Once upon a time, there was a little boy named Tom. One day he '
[raw] repeated 8-grams 12%
went to the park to play. He saw a big slide and wanted to go down it. But he was scared and sad. He wanted to go down the slide and slide d
[adapted] repeated 8-grams 2%
wa