CLM vs kev
Contrastive Language Models: CLM-8B, an open System-1 decision model scoring states against actions — typed choice/score/yes-no answers on a TypeSafe-compatible API, up to 9x faster than Jev. — versus — Open reproduction of TypeSafe's Jev: a LoRA + readout head on Qwen that answers many typed questions about one document in a single prefill pass, returning calibrated probabilities.
Both are open stand-ins for TypeSafe's Jev; kev reproduces it as a LoRA + readout head on Qwen, CLM trains a contrastive state-action model and replays Jev requests unchanged.
| CLM | kev | |
|---|---|---|
| Stars | 3.0k | 8.7k |
| Forks | 256 | 571 |
| Language | Python | Python |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 6 days ago | 5 days ago |
| Topics | decision-models, local | decision-models, training, local |
| Curated connections | 5 | 11 |
CLM — the curator's take
Use it when an agent must pick among many candidates fast — next action, tool, best-of-N verifier — and you want calibrated probabilities on your own GPU; cached state/action embeddings are why it beats Jev on latency. Not for anything generative: it ranks, never writes. Needs a Qwen3-8B encoder under vLLM plus the head, and states past 2,048 tokens are truncated unless you raise both limits.
kev — the curator's take
Reach for kev when the task is decisions, not prose: route this ticket, score this risk, answer 30 yes/no questions about one document — and you want a number you can threshold on rather than text you have to parse. The head is trained with cross-entropy on labelled outcomes, so the probabilities mean something (kev-8b: Brier 0.34, 8% confident errors out of domain), and a block-causal mask keeps questions from seeing each other, verified to 4e-6 against separate requests. It speaks TypeSafe's System One API, so their SDK works against localhost with a base_url change. Not a general chat or agent model - it never decodes, and it only answers the question types you define (noul/choice/score). Don't use it zero-shot on your own domain either: the value is in training the head on your labels, and the three 0.6B/4B/8B checkpoints are explicitly previews that failed the author's own release screen on held-out rule reasoning. Real Jev still wins out of domain (0.86 vs 0.77) - kev's pitch is that it runs on your laptop and you own the weights.