StackMap
Subscribe

CLM vs kev

Contrastive Language Models: CLM-8B, an open System-1 decision model scoring states against actions — typed choice/score/yes-no answers on a TypeSafe-compatible API, up to 9x faster than Jev. — versus — Open reproduction of TypeSafe's Jev: a LoRA + readout head on Qwen that answers many typed questions about one document in a single prefill pass, returning calibrated probabilities.

The curated verdict

Both are open stand-ins for TypeSafe's Jev; kev reproduces it as a LoRA + readout head on Qwen, CLM trains a contrastive state-action model and replays Jev requests unchanged.

CLMkev
Stars3.0k8.7k
Forks256571
LanguagePythonPython
LicenseApache-2.0Apache-2.0
Last activity6 days ago5 days ago
Topicsdecision-models, localdecision-models, training, local
Curated connections511

CLM — the curator's take

Use it when an agent must pick among many candidates fast — next action, tool, best-of-N verifier — and you want calibrated probabilities on your own GPU; cached state/action embeddings are why it beats Jev on latency. Not for anything generative: it ranks, never writes. Needs a Qwen3-8B encoder under vLLM plus the head, and states past 2,048 tokens are truncated unless you raise both limits.

kev — the curator's take

Reach for kev when the task is decisions, not prose: route this ticket, score this risk, answer 30 yes/no questions about one document — and you want a number you can threshold on rather than text you have to parse. The head is trained with cross-entropy on labelled outcomes, so the probabilities mean something (kev-8b: Brier 0.34, 8% confident errors out of domain), and a block-causal mask keeps questions from seeing each other, verified to 4e-6 against separate requests. It speaks TypeSafe's System One API, so their SDK works against localhost with a base_url change. Not a general chat or agent model - it never decodes, and it only answers the question types you define (noul/choice/score). Don't use it zero-shot on your own domain either: the value is in training the head on your labels, and the three 0.6B/4B/8B checkpoints are explicitly previews that failed the author's own release screen on held-out rule reasoning. Real Jev still wins out of domain (0.86 vs 0.77) - kev's pitch is that it runs on your laptop and you own the weights.