AnyJev vs laya
Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM. — versus — Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint.
Both give you typed decisions with calibrated probabilities instead of generated text; laya ships its own small multilingual checkpoints that run without an LLM, AnyJev reads the decision out of an open LLM you already serve, with no training or a tiny fitted head.
| AnyJev | laya | |
|---|---|---|
| Stars | 998 | 30k |
| Forks | 130 | 2.6k |
| Language | Python | Python |
| License | Apache-2.0 | Apache-2.0 |
| Last activity | 3 days ago | today |
| Topics | decision-models, local | decision-models, local |
| Curated connections | 4 | 5 |
AnyJev — the curator's take
The cleanest way to get decisions out of a model you already run: ask a typed question, read the next-token distribution from one prefill, nothing generated or parsed. Level L0 removes option-position bias with zero labels; L2 fits a tiny closed-form head from 100 to 300 labels and serves it from a vLLM pooling endpoint, and `anyjev.pipeline` measures accuracy, calibration and latency on your own box instead of asking you to trust their tables. One finding worth stealing: cutting Qwen2.5-7B to 18 of 28 blocks was faster with slightly better accuracy. It is a week-old research repo from two authors, so expect API churn, and it assumes you can run an open model; without a GPU, laya's small checkpoints are the lighter path.
laya — the curator's take
Use laya where you now spend an LLM call on a classification: routing a ticket, scoring urgency, a yes/no guard. Typed questions in, calibrated probabilities out, in one pass of about 33 ms, in 100+ languages, with pip install and extras for LangGraph, LlamaIndex, CrewAI, MCP and an HTTP server. The probabilities are the point: you can threshold them. Mind the limits it publishes itself: the multilingual checkpoint cuts at 1,024 tokens unless you pass max_len=8192, and past about 4,000 tokens its own long-document bench drops to 8 to 17 correct of 20. It decides; it does not explain or extract, so free-form answers still need an LLM. Two weeks old despite the star count: pin the version.