StackMap
Subscribe

autoharness vs SkillOpt

Self-learning skill layer for Claude Code: distills skills from your real sessions, merges same-scenario ones, updates them in use and prunes the unused — touching only skills it wrote. — versus — Microsoft's text-space optimizer that trains a frozen agent's skill document like weights — rollouts, bounded edits, validation-gated updates — and ships a compact best_skill.md.

The curated verdict

Both evolve agent skills from real sessions; autoharness learns, merges and prunes continuously inside Claude Code with no eval, SkillOpt (incl. its nightly Sleep mode) accepts a skill edit only when a held-out validation score improves.

autoharnessSkillOpt
Stars11k18k
Forks6081.7k
LanguagePythonPython
LicenseMITMIT
Last activity2 days ago5 days ago
Topicsskills, codingskills, training
Curated connections54

autoharness — the curator's take

Install it if you live in Claude Code and want lessons captured without curating skills by hand — zero config, zero deps, and it never touches skills you wrote or installed. Skip it off Claude Code, or if you need measured gains: skills survive on adherence, not a held-out score. Needs Python 3.11+ as `python3` on PATH, or its hooks stay off (macOS's stock /usr/bin/python3 is 3.9).

SkillOpt — the curator's take

Use it when you have a task with a scorer and want a skill that measurably improves on it — an edit lands only if the held-out score rises, and deployment adds zero model calls. Not a passive learner: you need data, a benchmark and an optimizer-model budget; the nightly SkillOpt-Sleep mode is the bridge to everyday Claude Code/Codex sessions. No metric? An in-use skill learner fits better.