fast-jev-compaction
Claude Code plugin that replaces the compaction summary with Jev decisions: every tool call and result is scored in one fast request, stale ones are dropped or truncated, everything kept stays verbatim. Also usable as an npm library.
What and why
Most context compaction asks an LLM to summarize old turns. A summary is lossy: a file path, exact error, constraint, or command can disappear even when it matters later. This library never rewrites anything. It only deletes tool calls and tool results Jev says are no longer needed, and it asks Jev while showing it the whole conversation. User and assistant text stays verbatim and in order.
The repository is both an npm package (src/) and a Claude Code plugin
(hooks/, .claude-plugin/) that uses the package to replace Claude Code's
built-in compaction summary with the original messages.
How it works
- Every
tool_useis paired with itstool_resultbytool_use_id. Calls in the first message or in the newestpreserveRecentMessagesmessages are pinned and never touched. - The state sent to Jev is the whole conversation so far, oldest first,
with every tool result replaced by a short note (
ok, 4213 chars (omitted)). Tool inputs are included, texts are included, nothing is summarized. - The state is fitted into
maxStateTokens(25k by default) in stages, each applied only if the previous one was not enough: tool inputs truncated to 1000, then 200, then 60 characters; long texts abridged to head + tail, oldest non-pinned messages first; old non-pinned messages collapsed to a[… N chars omitted …]note; old tool calls reduced to one line each (t12 Read file_path=src/a.ts → ok 480ch); old call-less messages left out; runs of old call-only messages folded into one entry. If it still does not fit, compaction throws. Tokens are estimated without a tokenizer (a word per six letters, half a token per digit, ~one per other symbol), calibrated to land a little above the counts Jev reports. - For every non-pinned call Jev gets two
noulquestions: should the call stay (knowing it was made, with its input, still matters), and should the result stay verbatim (its contents are still needed and re-running the tool would not do). - Questions are split into as many requests as needed so state plus questions
stays under
maxRequestTokens(30k by default, under Jev's 32k request limit). The same full state is resent with every request; requests run concurrently and their answers are merged. - Decisions per call, against
keepThreshold:keepResult ≥ threshold→ keep call and result;- else
keepCall ≥ threshold→ keep the call, truncate the result to its firsttruncateHeadCharscharacters plus a one-line note; - else → remove the call together with its result.
- The message list is rebuilt: a message that loses all its content is removed, untouched messages are returned as the same objects, and no result is ever left without its call.
Jev failures, malformed answers, a missing key, or a history that cannot be fitted throw; the caller (or the Claude Code hook) decides what to fall back to.
Install and usage
npm install fast-jev-compaction
export TYPESAFE_API_KEY=...
import { compactMessages, reductionRatio, type Message } from 'fast-jev-compaction';
const transcript: Message[] = [
{ role: 'user', text: 'Fix the failing test. Never edit src/generated.', toolUses: [] },
{
role: 'assistant',
text: '',
toolUses: [{ tool_use_id: 'toolu_1', tool: 'Read', input: { file_path: 'src/a.ts' } }],
},
{ role: 'user', text: '', toolUses: [], toolResults: [{ tool_use_id: 'toolu_1', text: '…file…' }] },
// …
];
const result = await compactMessages(transcript, { preserveRecentMessages: 4 });
console.log(result.messages, result.decisions, result.stats);
if (reductionRatio(result) < 0.25) {
// not worth it: keep the original