AutoHarness
Self-Learning Skills for Claude Code
autoharness is a self-learning skill layer for Claude Code. It learns skills from your real sessions, merges same-scenario ones instead of stacking near-duplicates, updates them in use, and prunes any that stop getting used — so the layer stays clean on its own, touching only the skills it wrote itself.
Same model, different harness — 42% → 78% on CORE-Bench (HAL). The harness does much of the work (swyx's Big Model vs Big Harness), yet it's still rebuilt by hand every model generation. autoharness bets one slice of it — the skill layer — can maintain itself.
| Learns from real work | Each episode is distilled into a skill from the session you were already having — no separate data-collection or replay loop. It fires on its own once a session has done enough work; /learn distills on demand when you want a lesson kept now. |
| Groups, doesn't just pile up | A new episode doesn't always add a skill — the reflector compares it against what's there and folds same-scenario skills into one, so the layer consolidates by category instead of accreting near-duplicates. A fold records which skill absorbed which, so a merge is never mistaken for a death. |
| Keeps its own library in view | Every session opens with a grouped index of the skills it wrote, so recall doesn't depend on the host happening to surface them. The host's native recall is left exactly as it was; the index is added on top. |
| Validated in use, not on a benchmark | A skill survives by being adhered to in later turns (loads over the requests it was available for), not a held-out score. No oracle on the active path, and no tokens spent on a dedicated eval. |
| Only its own skills | Touches only the skills it generated through this plugin — everything else, whether you wrote it or installed it, is left completely alone. |
| Evidence kept for later | Every create/update logs its scenario and decision to a per-skill ledger — the raw material to build a benchmark from real usage if you ever want one. |
Install
Requires Python 3.11+ as the python3 on your PATH — autoharness runs entirely as Python
(zero third-party dependencies); its hooks and MCP server won't fire without it. The hooks resolve
bare python3, so an older interpreter earlier on your PATH (Xcode ships 3.9.6 at
/usr/bin/python3) turns every hook off for the session; autoharness says so on stderr.
Type these in the Claude Code input box.
/plugin marketplace add tigerless-labs/autoharness
/plugin install autoharness@autoharness
Then run /reload-plugins (or restart Claude Code).
Zero config. It now watches your sessions and lands learned skills into .claude/skills/ in the
background. Cadence and lifecycle thresholds are tunable — see Configuration.
Nothing to invoke, but one entry point exists when you want it: /learn distills the session
you're in right now — say it after working something out and the lesson goes through the same
proposal-and-validation chain the background pass uses.
MCP server naming. The .mcp.json registers the server as stage_skill, but agent
definitions reference the fully-qualified name mcp__plugin_autoharness_stage_skill__stage_skill.
This translation is automatic: the plugin runt