StackMap
Subscribe

kev vs laya

Open reproduction of TypeSafe's Jev: a LoRA + readout head on Qwen that answers many typed questions about one document in a single prefill pass, returning calibrated probabilities. — versus — Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint.

The curated verdict

Both answer many typed questions about one text in a single pass with calibrated probabilities; kev puts a LoRA and readout head on a Qwen LLM, laya ships its own small non-autoregressive checkpoints with a multilingual router.

kevlaya
Stars7.6k30k
Forks4662.6k
LanguagePythonPython
LicenseApache-2.0Apache-2.0
Last activity3 days agotoday
Topicsdecision-models, training, localdecision-models, local
Curated connections105

kev — the curator's take

Reach for kev when the task is decisions, not prose: route this ticket, score this risk, answer 30 yes/no questions about one document — and you want a number you can threshold on rather than text you have to parse. The head is trained with cross-entropy on labelled outcomes, so the probabilities mean something (kev-8b: Brier 0.34, 8% confident errors out of domain), and a block-causal mask keeps questions from seeing each other, verified to 4e-6 against separate requests. It speaks TypeSafe's System One API, so their SDK works against localhost with a base_url change. Not a general chat or agent model - it never decodes, and it only answers the question types you define (noul/choice/score). Don't use it zero-shot on your own domain either: the value is in training the head on your labels, and the three 0.6B/4B/8B checkpoints are explicitly previews that failed the author's own release screen on held-out rule reasoning. Real Jev still wins out of domain (0.86 vs 0.77) - kev's pitch is that it runs on your laptop and you own the weights.

laya — the curator's take

Use laya where you now spend an LLM call on a classification: routing a ticket, scoring urgency, a yes/no guard. Typed questions in, calibrated probabilities out, in one pass of about 33 ms, in 100+ languages, with pip install and extras for LangGraph, LlamaIndex, CrewAI, MCP and an HTTP server. The probabilities are the point: you can threshold them. Mind the limits it publishes itself: the multilingual checkpoint cuts at 1,024 tokens unless you pass max_len=8192, and past about 4,000 tokens its own long-document bench drops to 8 to 17 correct of 20. It decides; it does not explain or extract, so free-form answers still need an LLM. Two weeks old despite the star count: pin the version.