[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"readme:sia":3},"\u003Ch1>SIA (Self-Improving AI)\u003C\u002Fh1>\n\u003Cp>\u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.27276\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FarXiv-2605.27276-b31b1b.svg\" alt=\"arXiv\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fgithub.com\u002Fhexo-ai\u002Fsia\u002Factions\u002Fworkflows\u002Fci.yml\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fgithub.com\u002Fhexo-ai\u002Fsia\u002Factions\u002Fworkflows\u002Fci.yml\u002Fbadge.svg\" alt=\"CI\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fsia-agent\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fv\u002Fsia-agent.svg\" alt=\"PyPI version\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fpypi.org\u002Fproject\u002Fsia-agent\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fpypi\u002Fdm\u002Fsia-agent.svg\" alt=\"PyPI downloads\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fopensource.org\u002Flicenses\u002FMIT\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002FLicense-MIT-green.svg\" alt=\"License: MIT\" \u002F>\u003C\u002Fa>\n\u003Ca href=\"https:\u002F\u002Fwww.python.org\u002Fdownloads\u002F\" rel=\"nofollow ugc noopener\">\u003Cimg src=\"https:\u002F\u002Fimg.shields.io\u002Fbadge\u002Fpython-3.11+-blue.svg\" alt=\"Python 3.11+\" \u002F>\u003C\u002Fa>\u003C\u002Fp>\n\u003Cp>Official implementation of \u003Ca href=\"https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.27276\" rel=\"nofollow ugc noopener\">\u003Cstrong>SIA: Self Improving AI with Harness &amp; Weight Updates\u003C\u002Fstrong>\u003C\u002Fa> (Hebbar et al., 2026) — a self-improving loop where a language-model agent updates both the harness and the weights of a task-specific agent. The paper reports a 56.6% gain on LawBench, 91.9% runtime reduction on GPU kernels, and 502% improvement on single-cell RNA denoising over baseline.\u003C\u002Fp>\n\u003Cp>SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model \u002F Agent) on a benchmark task.\u003C\u002Fp>\n\u003Cblockquote>\n\u003Cp>\u003Cstrong>Just want to try it?\u003C\u002Fstrong> Skip to \u003Ca href=\"#run-sia-locally-with-built-in-tasks\" rel=\"nofollow ugc noopener\">Run SIA locally\u003C\u002Fa>.\u003C\u002Fp>\n\u003C\u002Fblockquote>\n\u003Ch2>Introduction Videos\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fwww.loom.com\u002Fshare\u002Fbe0534bc818d408bab937033c6457ec9\" rel=\"nofollow ugc noopener\">SIA setup\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fwww.loom.com\u002Fshare\u002F5b1dc2dc858b4493b4b348f0b88d5b9e\" rel=\"nofollow ugc noopener\">SIA Runs Visualizer\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Architecture\u003C\u002Fh2>\n\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhexo-ai\u002Fsia\u002FHEAD\u002Fdocs\u002Fflow.png\" alt=\"SIA orchestration flow\" width=\"720\" \u002F>\u003C\u002Fp>\n\u003Cp align=\"center\">\u003Ci>Control flow between Meta, Target, and Feedback agents over successive generations.\u003C\u002Fi>\u003C\u002Fp>\u003Cp>SIA operates by coordinating three main types of AI agents that work together to continuously improve task performance:\u003C\u002Fp>\n\u003Ch3>Glossary\u003C\u002Fh3>\n\u003Col>\n\u003Cli>\u003Cstrong>Meta-Agent\u003C\u002Fstrong>: Reads the task description and generates an initial Target Agent tailored to the task.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Target \u002F Task Specific Agent\u003C\u002Fstrong>: Attempts to complete the task and records its actions and results.\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Feedback\u002FImprovement Agent\u003C\u002Fstrong>: Reviews the Target Agent's performance logs, identifies improvements, and updates the Target Agent accordingly.\u003C\u002Fli>\n\u003C\u002Fol>\n\u003Cp>This iterative process allows the system to autonomously refine and enhance its ability to solve scientific tasks.\u003C\u002Fp>\n\u003Ch2>Benchmark Results\u003C\u002Fh2>\n\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhexo-ai\u002Fsia\u002FHEAD\u002Fdocs\u002Fmlebench.png\" alt=\"MLE Bench Results\" width=\"720\" \u002F>\u003Cbr \u002F>\u003Ci>OpenAI MLE-Bench Hard: a gauntlet of real Kaggle ML competitions where agents must write, run, and iterate full ML pipelines. SIA ranks #1 across all generations tested.\u003C\u002Fi>\u003C\u002Fp>\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhexo-ai\u002Fsia\u002FHEAD\u002Fdocs\u002Flawbench.png\" alt=\"LawBench Results\" width=\"720\" \u002F>\u003Cbr \u002F>\u003Ci>LawBench: predict the criminal charge from Chinese court case descriptions across 191 charge categories. SIA-W+H reaches 70.1% Top-1 accuracy, beating the prior SOTA of 45%.\u003C\u002Fi>\u003C\u002Fp>\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhexo-ai\u002Fsia\u002FHEAD\u002Fdocs\u002Ftrimul_cuda.png\" alt=\"TriMul CUDA Results\" width=\"720\" \u002F>\u003Cbr \u002F>\u003Ci>AlphaFold-3 TriMul Triton Kernel: implement and optimize the Triangle Multiplicative Update as a Triton kernel, preserving correctness while hitting H100 latency targets. SIA-W+H achieves 14x speedup over baseline.\u003C\u002Fi>\u003C\u002Fp>\u003Cp align=\"center\">\u003Cimg src=\"https:\u002F\u002Fraw.githubusercontent.com\u002Fhexo-ai\u002Fsia\u002FHEAD\u002Fdocs\u002Fdenoising.png\" alt=\"Denoising Results\" width=\"720\" \u002F>\u003Cbr \u002F>\u003Ci>scRNA-seq Denoising: impute missing gene expression values in single-cell RNA sequencing data. SIA-W+H scores 0.289 MSE\u003Csub>norm\u003C\u002Fsub>, surpassing the prior SOTA of 0.240.\u003C\u002Fi>\u003C\u002Fp>\u003Chr \u002F>\n\u003Ch2>Run SIA locally with built-in tasks\u003C\u002Fh2>\n\u003Cp>SIA ships with four built-in tasks: \u003Ccode>gpqa\u003C\u002Fcode>, \u003Ccode>lawbench\u003C\u002Fcode>, \u003Ccode>longcot-chess\u003C\u002Fcode>, \u003Ccode>spaceship-titanic\u003C\u002Fcode>.\u003C\u002Fp>\n\u003Ch3>Install\u003C\u002Fh3>\n\u003Cp>Pick the agent impl that matches the LLMs you want to run.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>Claude agent impl\u003C\u002Fstrong> (Claude Agent SDK, Claude models only):\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">python3 -m venv .venv &amp;&amp; source .venv\u002Fbin\u002Factivate\npip install 'sia-agent[claude]'\nexport ANTHROPIC_API_KEY=\"...\"\n\u003C\u002Fcode>\u003C\u002Fpre>\n\u003Cp>\u003Cstrong>OpenHands agent impl\u003C\u002Fstrong> (multi-provider — Gemini, OpenAI, Anthropic, etc.):\u003C\u002Fp>\n\u003Cpre>\u003Ccode class=\"language-bash\">python3 -m venv .venv &amp;&amp; s\n\u003C\u002Fcode>\u003C\u002Fpre>\n",1784828823927]