[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:soup":3},"\u003Cp align=\"center\">\n  \u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FMakazhanAlpamys\u002Fsoup\u002FHEAD\u002Fsoup.png\" alt=\"Soup\" width=\"280\" \u002F>\n\u003C\u002Fp>\u003Ch1>Soup\u003C\u002Fh1>\u003Cp align=\"center\">\n  \u003Cstrong>Fine-tune and post-train LLMs in one command. No SSH, no config hell.\u003C\u002Fstrong>\n\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Ftrysoup.dev\" rel=\"nofollow ugc noopener\">Website\u003C\u002Fa> ·\n  \u003Ca href=\"#quick-start\" rel=\"nofollow ugc noopener\">Quick Start\u003C\u002Fa> ·\n  \u003Ca href=\"#configuration\" rel=\"nofollow ugc noopener\">Config\u003C\u002Fa> ·\n  \u003Ca href=\"#documentation\" rel=\"nofollow ugc noopener\">Docs\u003C\u002Fa> ·\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002Fsoup\u002Fblob\u002FHEAD\u002Fdocs\u002Fcommands.md\" rel=\"nofollow ugc noopener\">Commands\u003C\u002Fa> ·\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002Fsoup\u002Fblob\u002FHEAD\u002Fdocs\u002Fmodels.md\" rel=\"nofollow ugc noopener\">Models\u003C\u002Fa> ·\n  \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F8RgVbFA6Zq\" rel=\"nofollow ugc noopener\">Discord\u003C\u002Fa>\n\u003C\u002Fp>\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fsoup-cli\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fsoup-cli?color=blue\" alt=\"PyPI\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fpepy.tech\u002Fproject\u002Fsoup-cli\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpepy\u002Fdt\u002Fsoup-cli?color=blue\" alt=\"Downloads\" \u002F>\u003C\u002Fa>\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.10%2B-blue\" alt=\"Python 3.10+\" \u002F>\n  \u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Flicense-Apache--2.0-blue\" alt=\"Apache-2.0 License\" \u002F>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fendpoint?url=https:\u002F\u002Fgist.githubusercontent.com\u002FMakazhanAlpamys\u002F65fdc943f85f3b2c46ecddb415c2b779\u002Fraw\u002Fsoup_tests.json\" alt=\"Tests\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002FSoup\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"CI\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Ftrysoup.dev\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fwebsite-trysoup.dev-blue\" alt=\"Website\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fdiscord.gg\u002F8RgVbFA6Zq\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDiscord-join-5865F2?logo=discord&amp;logoColor=white\" alt=\"Discord\" \u002F>\u003C\u002Fa>\n  \u003Ca href=\"https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21771064\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FDOI-10.5281%2Fzenodo.21771064-blue?logo=zenodo&amp;logoColor=white\" alt=\"DOI: 10.5281\u002Fzenodo.21771064\" \u002F>\u003C\u002Fa>\n\u003C\u002Fp>\u003Chr \u002F>\n\u003Cp>Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">pip install \"soup-cli[train]\"   # add [train] to fine-tune; bare `soup-cli` is the light CLI\nsoup init --template chat\nsoup train\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>Fine-tune an 8B model on a 4 GB laptop GPU.\u003C\u002Fstrong> Layer streaming keeps the frozen base out of\nVRAM and feeds it to the GPU one decoder layer at a time. Measured on an RTX 3050 Laptop 4 GB:\nLlama-3.1-8B-Instruct + NF4 at \u003Cstrong>119.6 tok\u002Fs, 3.32 GB peak\u003C\u002Fstrong> — bit-exact against a normal\nresident run. Opt-in (\u003Ccode>stream_layers: true\u003C\u002Fcode>) and still BETA —\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002Fsoup\u002Fblob\u002FHEAD\u002Fdocs\u002Fperformance-and-quantization.md#layer-streaming-beta-v0720-nf4-v0722-disk--wider-archs-v0723-preference-losses-v0724\" rel=\"nofollow ugc noopener\">how it works\u003C\u002Fa> ·\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002FMakazhanAlpamys\u002Fsoup\u002Fblob\u002FHEAD\u002Fbenchmarks\u002F\" rel=\"nofollow ugc noopener\">all measurements\u003C\u002Fa> · \u003Ca href=\"https:\u002F\u002Fdoi.org\u002F10.5281\u002Fzenodo.21771064\" rel=\"nofollow ugc noopener\">paper\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp align=\"center\">\n  \u003Ca href=\"https:\u002F\u002Fyoutu.be\u002FT1LCErE943E\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002FMakazhanAlpamys\u002Fsoup\u002FHEAD\u002Fdocs\u002Fassets\u002Flayer-streaming.gif\" alt=\"soup train pre-flight for Llama-3.1-8B on a 4 GB card: a 3.60 GB base store pinned in RAM across 32 layers and two 113 MB VRAM buffers, then a measured peak of 3.32 GB at 119.6 tok\u002Fs, stopping short of the 4 GB line\" \u002F>\u003C\u002Fa>\u003Cbr \u002F>\n  \u003Csub>Llama-3.1-8B-Instruct + NF4, LoRA, batch 1, seq 512 on an RTX 3050 Laptop 4 GB — \u003Cb>3.32 GB peak, 119.6 tok\u002Fs\u003C\u002Fb>. \u003Ca href=\"https:\u002F\u002Fyoutu.be\u002FT1LCErE943E\" rel=\"nofollow ugc noopener\">Full video (90s)\u003C\u002Fa>\u003C\u002Fsub>\n\u003C\u002Fp>\u003Ch2>Why Soup?\u003C\u002Fh2>\n\u003Cp>Training LLMs is still painful. Even experienced teams spend 30-50% of their time fighting\ninfrastructure instead of improving models. Soup fixes that.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>Zero SSH.\u003C\u002Fstrong> Never SSH into a broken GPU box again.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>One config.\u003C\u002Fstrong> A simple YAML file is all you need.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Auto everything.\u003C\u002Fstrong> Batch size, GPU detection, quantization — handled.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Works locally.\u003C\u002Fstrong> Train on your own GPU with QLoRA. No cloud required.\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>What's New\u003C\u002Fh2>\n\u003Cp>\u003Cstrong>v0.72.4 — align on a laptop: DPO, ORPO, SimPO and KTO over layer streaming.\u003C\u002Fstrong> Layer\nstreaming keeps the frozen base out of VRAM and feeds it to the GPU one decoder layer at\na time. It used to support supervised fine-tuning only; now it runs the preference\nlosses too.\u003C\u002Fp>\n\u003Cul>\n\u003Cli>\u003Cstrong>DPO's reference model is free.\u003C\u002Fstrong> DPO needs a reference to co\u003C\u002Fli>\n\u003C\u002Ful>\n",1786232078815]