agentic-context-engine vs SkillOpt
Learning loop for any agent: reflect on failures, distill strategies into a Skillbook, inject them next run — 2x consistency on Tau2, 49% token cuts. LiteLLM-based, 100+ providers. — versus — Microsoft's text-space optimizer that trains a frozen agent's skill document like weights — rollouts, bounded edits, validation-gated updates — and ships a compact best_skill.md.
Both turn rollouts and reflection into a reusable text artifact injected next run; ACE grows a Skillbook from failures, SkillOpt accepts an edit only when held-out validation improves.
| agentic-context-engine | SkillOpt | |
|---|---|---|
| Stars | 2.6k | 18k |
| Forks | 312 | 1.7k |
| Language | Python | Python |
| License | Apache-2.0 | MIT |
| Last activity | 17 days ago | 5 days ago |
| Topics | memory, agents | skills, training |
| Curated connections | 8 | 4 |
agentic-context-engine — the curator's take
The in-process answer to 'my agent repeats the same mistakes': wrap your agent, feed it corrections, and ACE extracts reusable strategies it injects on later runs — no fine-tuning, no reward signals, and the numbers are concrete (2x pass^4 on Tau2, ~$1.50 to learn its way through a 14k-line translation). Pick it over a memory *service* when you want the learning inside your Python process rather than behind an HTTP API. NOT magic memory: strategies come from explicit feedback loops you wire up, quality follows the judge model, and the open-source engine is the on-ramp to the hosted Kayba product — check where the managed line lands before betting infra on it.
SkillOpt — the curator's take
Use it when you have a task with a scorer and want a skill that measurably improves on it — an edit lands only if the held-out score rises, and deployment adds zero model calls. Not a passive learner: you need data, a benchmark and an optimizer-model budget; the nightly SkillOpt-Sleep mode is the bridge to everyday Claude Code/Codex sessions. No metric? An in-use skill learner fits better.