AnyJev vs kev
Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM. — versus — Open reproduction of TypeSafe's Jev: a LoRA + readout head on Qwen that answers many typed questions about one document in a single prefill pass, returning calibrated probabilities.
Both turn an open LLM into a typed decision model read from one prefill; kev trains a LoRA and readout head, AnyJev needs no training at L0 and fits a closed-form head from 100 to 300 labels at L2.
| AnyJev | kev | |
|---|---|---|
| Stars | 998 | 7.6k |
| Forks | 130 | 466 |
| Language | Python | Python |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 3 days ago | 3 days ago |
| Topics | decision-models, local | decision-models, training, local |
| Curated connections | 4 | 10 |
AnyJev — the curator's take
The cleanest way to get decisions out of a model you already run: ask a typed question, read the next-token distribution from one prefill, nothing generated or parsed. Level L0 removes option-position bias with zero labels; L2 fits a tiny closed-form head from 100 to 300 labels and serves it from a vLLM pooling endpoint, and `anyjev.pipeline` measures accuracy, calibration and latency on your own box instead of asking you to trust their tables. One finding worth stealing: cutting Qwen2.5-7B to 18 of 28 blocks was faster with slightly better accuracy. It is a week-old research repo from two authors, so expect API churn, and it assumes you can run an open model; without a GPU, laya's small checkpoints are the lighter path.
kev — the curator's take
Reach for kev when the task is decisions, not prose: route this ticket, score this risk, answer 30 yes/no questions about one document — and you want a number you can threshold on rather than text you have to parse. The head is trained with cross-entropy on labelled outcomes, so the probabilities mean something (kev-8b: Brier 0.34, 8% confident errors out of domain), and a block-causal mask keeps questions from seeing each other, verified to 4e-6 against separate requests. It speaks TypeSafe's System One API, so their SDK works against localhost with a base_url change. Not a general chat or agent model - it never decodes, and it only answers the question types you define (noul/choice/score). Don't use it zero-shot on your own domain either: the value is in training the head on your labels, and the three 0.6B/4B/8B checkpoints are explicitly previews that failed the author's own release screen on held-out rule reasoning. Real Jev still wins out of domain (0.86 vs 0.77) - kev's pitch is that it runs on your laptop and you own the weights.