StackMap
Subscribe
Explore / mini-AGI
volotat

mini-AGI

Continual-learning byte-level LM trained from scratch on one 8 GB GPU: weights page in from disk, capacity grows and prunes itself while it keeps reading a single data stream.

1,231 200 Python MITupdated yesterday
View on GitHubDispute this mapping →
Curator's take

A research toy, and its author says so: read it for how continual single-stream learning without catastrophic forgetting can run on a laptop GPU — disk-paged experts, growth and pruning mid-training. Don't use it as a model: weights are undertrained and unreleased, raw outputs loop. Want to learn LLM training end to end? train-llm-from-scratch is the clearer textbook.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside mini-AGI. Ranked by curator confidence.

alternativetrain-llm-from-scratchmini-AGI
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md3 min read

mini-AGI

mini-AGI - is a continual learning byte-level language model that assembles its own architecture, trains from scratch on a single 8 GB VRAM GPU, and keeps learning from everything it reads. It stores its weights as ordinary files on disk and pages them onto the card as it needs them, so the parameter count is bounded by free disk space rather than by VRAM. It grows new capacity while training when it runs short, prunes what nothing asks for, and reads through exactly the same code path it serves on. Targeted at a PC or laptop with at least an 8 GB VRAM GPU on the board.

NOTE: as of now this is a small toy-level model. Do not expect a frontier level capabilities. This is rather a small experiment to show, that continual learning from the single stream of data without catastrophic forgetting is possible. Furthermore it is possible on a modest hardware. Which means that almost everyone could train their own version of the model (or simply continue training this one) exactly as they see it fit. And the capabilities would be bounded by the actual hardware, scale and quality of the data available and the amount of time one willing to spend on training the model.

The name is a half-joke and not a statement of the current capabilities of the model, but rather the potential and traits it has. Continual single-stream learning and natural size and processing adaptation to the available resources exactly what I myself expect from an AGI to have. It has just a small toy-level memory footprint and hence it is “mini-AGI”.

dashboard Here is how min-run dashboard looks like. The model is pointed to the corpus to constantly read and learn from.

History - here is the samples from the whole training run history so far. You can inspect them yourself to see how the model improved over the course of training/reading the corpus.

The "final" weights are not published yet. The run is still reading its first pass over the corpus, the weights go up once it has been through all of it, which is several weeks away at the current rate. If you would like to play with undertrained weights as they are right now, you can find the most recent (1 oct 2026) snapshot here: Volotat/mini-AGI-undertrained

Graph of the whole run so far

training progress

Every sample round of the run to date: 1,717.1M characters over 2,462 evaluations.

Current quality of samples the model generates

The round with the lowest held-out loss so far - 0.6235 nats at 1,682.8M characters. Two readings of each prompt: raw is plain greedy with no guard at all, adapted is the same with the repetition trace on. The whole history is in runs/samples.txt.

==============================================================================
step 822,034   1682.8M of 7,880M characters (21.36%)   264 min   397 experts
context 4,096 characters of 4,096   reading 1,391 char/s   writing 34.6 char/s   still gaining +0.0018 deep into it
grad norm 1.67 against a clip of 1   clipping
train loss 0.5076   lr 1.84e-05   evidence t +2.49 over 65.7 (effect +0.0553)   rate x0.062
held-out loss 0.6235 +/-0.0281 nats   0.8995 bits/char   perplexity 1.87   gap +0.1159
  arithmetic 0.620   chat 0.590   chat_hermes 0.875   chess 0.401   code 0.522   reasoning 0.540   stories 0.410   wikipedia 1.031
repeats 24% of 8-grams, greedy with no guard
==============================================================================

--- stories ---
prompt: 'Once upon a time, there was a little boy named Tom. One day he '
[raw]  repeated 8-grams 12%
went to the park to play. He saw a big slide and wanted to go down it. But he was scared and sad. He wanted to go down the slide and slide d
[adapted]  repeated 8-grams 2%
wa