StackMap
Subscribe

AnyJev vs laya

Nokia research: turn any open LLM into a Jev-style decision model. Typed choice, yes/no or score from one prefill, de-biased with no labels or a fitted head, served on vLLM. — versus — Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint.

The curated verdict

Both give you typed decisions with calibrated probabilities instead of generated text; laya ships its own small multilingual checkpoints that run without an LLM, AnyJev reads the decision out of an open LLM you already serve, with no training or a tiny fitted head.

AnyJevlaya
Stars99830k
Forks1302.6k
LanguagePythonPython
LicenseApache-2.0Apache-2.0
Last activity3 days agotoday
Topicsdecision-models, localdecision-models, local
Curated connections45

AnyJev — the curator's take

The cleanest way to get decisions out of a model you already run: ask a typed question, read the next-token distribution from one prefill, nothing generated or parsed. Level L0 removes option-position bias with zero labels; L2 fits a tiny closed-form head from 100 to 300 labels and serves it from a vLLM pooling endpoint, and `anyjev.pipeline` measures accuracy, calibration and latency on your own box instead of asking you to trust their tables. One finding worth stealing: cutting Qwen2.5-7B to 18 of 28 blocks was faster with slightly better accuracy. It is a week-old research repo from two authors, so expect API churn, and it assumes you can run an open model; without a GPU, laya's small checkpoints are the lighter path.

laya — the curator's take

Use laya where you now spend an LLM call on a classification: routing a ticket, scoring urgency, a yes/no guard. Typed questions in, calibrated probabilities out, in one pass of about 33 ms, in 100+ languages, with pip install and extras for LangGraph, LlamaIndex, CrewAI, MCP and an HTTP server. The probabilities are the point: you can threshold them. Mind the limits it publishes itself: the multilingual checkpoint cuts at 1,024 tokens unless you pass max_len=8192, and past about 4,000 tokens its own long-document bench drops to 8 to 17 correct of 20. It decides; it does not explain or extract, so free-form answers still need an LLM. Two weeks old despite the star count: pin the version.