StackMap
Subscribe

agentic-context-engine vs SkillOpt

Learning loop for any agent: reflect on failures, distill strategies into a Skillbook, inject them next run — 2x consistency on Tau2, 49% token cuts. LiteLLM-based, 100+ providers. — versus — Microsoft's text-space optimizer that trains a frozen agent's skill document like weights — rollouts, bounded edits, validation-gated updates — and ships a compact best_skill.md.

The curated verdict

Both turn rollouts and reflection into a reusable text artifact injected next run; ACE grows a Skillbook from failures, SkillOpt accepts an edit only when held-out validation improves.

agentic-context-engineSkillOpt
Stars2.6k18k
Forks3121.7k
LanguagePythonPython
LicenseApache-2.0MIT
Last activity17 days ago5 days ago
Topicsmemory, agentsskills, training
Curated connections84

agentic-context-engine — the curator's take

The in-process answer to 'my agent repeats the same mistakes': wrap your agent, feed it corrections, and ACE extracts reusable strategies it injects on later runs — no fine-tuning, no reward signals, and the numbers are concrete (2x pass^4 on Tau2, ~$1.50 to learn its way through a 14k-line translation). Pick it over a memory *service* when you want the learning inside your Python process rather than behind an HTTP API. NOT magic memory: strategies come from explicit feedback loops you wire up, quality follows the judge model, and the open-source engine is the on-ramp to the hosted Kayba product — check where the managed line lands before betting infra on it.

SkillOpt — the curator's take

Use it when you have a task with a scorer and want a skill that measurably improves on it — an edit lands only if the held-out score rises, and deployment adds zero model calls. Not a passive learner: you need data, a benchmark and an optimizer-model budget; the nightly SkillOpt-Sleep mode is the bridge to everyday Claude Code/Codex sessions. No metric? An in-use skill learner fits better.