StackMap
Subscribe

mini-AGI vs train-llm-from-scratch

Continual-learning byte-level LM trained from scratch on one 8 GB GPU: weights page in from disk, capacity grows and prunes itself while it keeps reading a single data stream. — versus — The full LLM pipeline hand-written in plain PyTorch — tokens, transformer, pretraining, then SFT, reward model, PPO, DPO, GRPO. No trl, no peft: read every algorithm, train on one GPU.

The curated verdict

Both train your own small LM from scratch on a single GPU; train-llm-from-scratch teaches the standard pretrain→SFT→RL pipeline, mini-AGI experiments with never-ending continual learning instead.

mini-AGItrain-llm-from-scratch
Stars1.2k12k
Forks2001.7k
LanguagePythonPython
LicenseMITMIT
Last activityyesterday9 days ago
Topicstraining, localtraining
Curated connections13

mini-AGI — the curator's take

A research toy, and its author says so: read it for how continual single-stream learning without catastrophic forgetting can run on a laptop GPU — disk-paged experts, growth and pruning mid-training. Don't use it as a model: weights are undertrained and unreleased, raw outputs loop. Want to learn LLM training end to end? train-llm-from-scratch is the clearer textbook.

train-llm-from-scratch — the curator's take

The textbook that runs: every stage from raw text to an aligned reasoning-style model, each algorithm implemented by hand in readable PyTorch — including the post-training alphabet (SFT → reward model → PPO/DPO → GRPO) that most tutorials wave at. If you want to understand what TRL actually does under its trainer classes, this is the fastest honest path, and it fits on a single GPU at the 13M–1B scale. NOT for production: hand-rolled training code at toy scale is the point, not the product — when you're done learning, ship with TRL or LLaMA-Factory. From the author of all-agentic-architectures, same runnable-textbook DNA.