[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:train-llm-from-scratch":3},"\u003Cp>\u003Cimg src=\"https:\u002F\u002Fcdn-images-1.medium.com\u002Fmax\u002F5200\u002F1*r99Hq3YBd5FTTWLNYKKvPw.png\" alt=\"main image\" \u002F>\u003C\u002Fp>\n\u003Cdiv align=\"center\">\n\u003Ch1>Train LLM From Scratch\u003C\u002Fh1>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FPython-3.9%2B-blue\" alt=\"Python\" \u002F> \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-MIT-green\" alt=\"License\" \u002F> \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FContributions-Welcome-blue\" alt=\"Contributions\" \u002F> \u003Ca href=\"https:\u002F\u002Ffareedkhan-dev.github.io\u002Ftrain-llm-from-scratch\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDocs-Available-success\" alt=\"Docs\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>\u003Cstrong>I am Looking for a PhD position in AI\u003C\u002Fstrong>. \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FFareedKhan-dev\" rel=\"nofollow ugc noopener\">GitHub\u003C\u002Fa>\u003C\u002Fp>\n\u003C\u002Fdiv>\u003Cp>I implemented a transformer model from scratch using PyTorch, based on the paper \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F1706.03762\" rel=\"nofollow ugc noopener\">Attention is All You Need\u003C\u002Fa>. You can use my scripts to train your own \u003Cstrong>billion\u003C\u002Fstrong> or \u003Cstrong>million\u003C\u002Fstrong> parameter LLM using a single GPU.\u003C\u002Fp>\n\u003Cp>This started as a pretraining tutorial. It now goes all the way from raw text to an aligned, reasoning style model, with every algorithm hand written in plain PyTorch (no \u003Ccode>trl\u003C\u002Fcode>, no \u003Ccode>peft\u003C\u002Fcode>, no \u003Ccode>transformers\u003C\u002Fcode>). The whole journey is one idea repeated: turn text into numbers, predict the next token, then keep changing the data and the loss until the model does what we want.\u003C\u002Fp>\n\u003Cp>\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FFareedKhan-dev\u002Ftrain-llm-from-scratch\u002FHEAD\u002Fimages\u002F00_pipeline.png\" alt=\"From raw text to an aligned reasoning model\" \u002F>\u003C\u002Fp>\n\u003Cp>Here is the path we will walk, end to end:\u003C\u002Fp>\n\u003Cpre>\u003Ccode>raw text  -&gt;  tokens  -&gt;  a Transformer  -&gt;  next-token loss  -&gt;  a base model\nbase model  -&gt;  SFT  -&gt;  Reward Model  -&gt;  {PPO, DPO}  -&gt;  GRPO  -&gt;  evaluation and chat\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>Below is the output of a trained 13 million parameter LLM, just so you can see where the small end of this starts:\u003C\u002Fp>\n\u003Cpre>\u003Ccode>In ***1978, The park was returned to the factory-plate that\nthe public share to the lower of the electronic fence that\nfollow from the Station's cities. The Canal of ancient Western\nnations were confined to the city spot. The villages were directly\nlinked to cities in China that revolt that the US budget and in\nOdambinais is uncertain and fortune established in rural areas.\n\u003C\u002Fcode>\u003C\u002Fpre>\n\n\u003Ch2>Table of Contents\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"#who-this-is-for\" rel=\"nofollow ugc noopener\">Who this is for\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#prerequisites-and-training-time\" rel=\"nofollow ugc noopener\">Prerequisites and Training Time\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#setup\" rel=\"nofollow ugc noopener\">Setup\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#code-structure\" rel=\"nofollow ugc noopener\">Code Structure\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-1-preparing-the-data\" rel=\"nofollow ugc noopener\">Step 1: Preparing the Data\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-2-the-model-built-from-small-pieces\" rel=\"nofollow ugc noopener\">Step 2: The Model, Built From Small Pieces\u003C\u002Fa>\u003Cul>\n\u003Cli>\u003Ca href=\"#multi-layer-perceptron-mlp\" rel=\"nofollow ugc noopener\">Multi Layer Perceptron (MLP)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#single-head-attention\" rel=\"nofollow ugc noopener\">Single Head Attention\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#multi-head-attention\" rel=\"nofollow ugc noopener\">Multi Head Attention\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#the-transformer-block\" rel=\"nofollow ugc noopener\">The Transformer Block\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#the-full-transformer\" rel=\"nofollow ugc noopener\">The Full Transformer\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-3-pretraining-the-base-model\" rel=\"nofollow ugc noopener\">Step 3: Pretraining the Base Model\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-4-generating-text\" rel=\"nofollow ugc noopener\">Step 4: Generating Text\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-5-post-training-turning-a-base-model-into-an-assistant\" rel=\"nofollow ugc noopener\">Step 5: Post-Training, Turning a Base Model Into an Assistant\u003C\u002Fa>\u003Cul>\n\u003Cli>\u003Ca href=\"#sft-supervised-fine-tuning\" rel=\"nofollow ugc noopener\">SFT (Supervised Fine-Tuning)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#the-reward-model\" rel=\"nofollow ugc noopener\">The Reward Model\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#dpo-orpo-and-kto\" rel=\"nofollow ugc noopener\">DPO, ORPO and KTO\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#ppo\" rel=\"nofollow ugc noopener\">PPO\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#grpo--rlvr\" rel=\"nofollow ugc noopener\">GRPO \u002F RLVR\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-6-evaluation\" rel=\"nofollow ugc noopener\">Step 6: Evaluation\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#step-7-talking-to-the-model\" rel=\"nofollow ugc noopener\">Step 7: Talking to the Model\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#the-streamlit-control-panel\" rel=\"nofollow ugc noopener\">The Streamlit Control Panel\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#the-documentation-site\" rel=\"nofollow ugc noopener\">The Documentation Site\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#run-the-whole-thing\" rel=\"nofollow ugc noopener\">Run the Whole Thing\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"#whats-next\" rel=\"nofollow ugc noopener\">What's Next\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Who this is for\u003C\u002Fh2>\n\u003Cp>I tried to write this so one page works for very different readers:\u003C\u002Fp>\n\u003Cul>\n\u003Cli>If you are a \u003Cstrong>student\u003C\u002Fstrong>, read top to bottom. Every block of code comes after a plain explanation of what it does and why, and most blocks are followed by the output you should expect.\u003C\u002Fli>\n\u003Cli>If you are a \u003Cstrong>developer\u003C\u002Fstrong>, the commands and file paths are all here. You can copy, run, and read the referenced source files directly.\u003C\u002Fli>\n\u003Cli>If you are a \u003Cstrong>researcher\u003C\u002Fstrong>, the post-training half is the interesting part: SFT, a Bradley-Terry reward model, PPO with GAE, DPO\u002FORPO\u002FKTO, and GRPO, all from scratch on the same small Transformer, trained on real public datasets.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>Ev\u003C\u002Fp>\n",1784749656235]