Abide
Coding agents break your rules from the very first edit. Abide catches every one and makes your agent fix it
1 in 13 turns break a rule no linter can catch · abide does · 300 ms per check · a tenth of a cent per turn
Measured by replaying 93 real Claude Code sessions (1,256 edits, 147 turns) in two repos against their own AGENTS.md, for 22 cents. Jev flagged 39 edits and 15 turns; an independent reviewer confirmed 10 and 11. The turn-level catches (single-use abstractions, oversized files, duplicated logic) held up 11 times in 15. Method, per-rule table and what was wrong: benchmarks/replay.
npx @coldtea/abide login # pick a key type, paste it once
npx @coldtea/abide init # hooks into every agent on this machine
Then start claude, codex, opencode or pi as usual. That is the whole setup.
What it does
Your AGENTS.md, CLAUDE.md and the rest of your project instructions are full of rules no linter can check. "Never let a raw error reach a user." "Don't create premature abstractions." Nothing can script those, so nothing enforces them. In 93 real sessions, the agent broke one on 1 turn in 13, from the first edit on.
https://github.com/user-attachments/assets/2d45f6b0-c889-474c-ab4a-8d019fdc7140
Abide enforces exactly those rules. On every edit (or turn) it asks Jev, TypeSafe's decision model, one question per rule and gets a probability back. Jev sees the rule and the diff, never the conversation, so edit 200 is checked like edit 1. Break a rule and the agent is told which one and fixes it in the same turn.
- One call per edit, about 300 ms, a few thousandths of a cent.
- Rules a linter could check are handed to your linter instead.
- No built-in rules. No instruction files, nothing to enforce.
- Your key, your data. Nothing here talks to a server of ours.
Not previously possible
Checking every edit or turn against every rule was never worth doing (economically and latency-wise) with an ordinary LLM. A check is about 2,500 tokens. At typical model prices that is a cent or more, and a few seconds, per edit, and the answer comes back as prose you then have to parse and cannot fully trust. Two hundred edits a day made it a non-starter.
Jev changes the arithmetic. It is a decision model, so it answers a typed question with a calibrated probability and nothing else. There is no free text, so there is nothing to make up. It is up to 100x cheaper than a typical LLM and answers in about 300 ms. That is what makes it reasonable to check every edit, every time.
Three minutes to the first catch
- Get a TypeSafe API key at typesafe.ai, or use a Vercel AI Gateway key you already have.
- Run
npx @coldtea/abide login, pick which kind of key it is and where it lives, and paste it. It goes to~/.abide/.envfor every repo on the machine, or to.env.localin this repo, owner-only either way. A.envyou already have at the repo root works too. - Run
npx @coldtea/abide initin your repo. - Start your agent. Its first turn compiles your rules into
.abide/rubric.jsonand tells you what it found. - Ask for something your rules forbid. An AGENTS.md that says "use Yup, never validate by hand" produces this the moment the agent writes a manual guard: