StackMap
Subscribe

AnyJev vs CLM

Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM. — versus — Contrastive Language Models: CLM-8B, an open System-1 decision model scoring states against actions — typed choice/score/yes-no answers on a TypeSafe-compatible API, up to 9x faster than Jev.

The curated verdict

Same goal — Jev-style typed decisions from open weights; anyjev converts any open LLM with one prefill, CLM ships a purpose-trained 8B model.

AnyJevCLM
Stars1.1k3.0k
Forks138256
LanguagePythonPython
LicenseApache-2.0Apache-2.0
Last activity4 days ago6 days ago
Topicsdecision-models, localdecision-models, local
Curated connections55

AnyJev — the curator's take

The cleanest way to get decisions out of a model you already run: ask a typed question, read the next-token distribution from one prefill, nothing generated or parsed. Level L0 removes option-position bias with zero labels; L2 fits a tiny closed-form head from 100 to 300 labels and serves it from a vLLM pooling endpoint, and `anyjev.pipeline` measures accuracy, calibration and latency on your own box instead of asking you to trust their tables. One finding worth stealing: cutting Qwen2.5-7B to 18 of 28 blocks was faster with slightly better accuracy. It is a week-old research repo from two authors, so expect API churn, and it assumes you can run an open model; without a GPU, laya's small checkpoints are the lighter path.

CLM — the curator's take

Use it when an agent must pick among many candidates fast — next action, tool, best-of-N verifier — and you want calibrated probabilities on your own GPU; cached state/action embeddings are why it beats Jev on latency. Not for anything generative: it ranks, never writes. Needs a Qwen3-8B encoder under vLLM plus the head, and states past 2,048 tokens are truncated unless you raise both limits.