StackMap
Subscribe

laya vs localjev

Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint. — versus — GitHub Next's local Jev bridge: a Bun/TypeScript POST /v1/systemone that translates typed decision questions into prompts for DiffusionGemma behind any OpenAI-compatible endpoint.

The curated verdict

Both serve Jev-style typed decisions; localjev translates questions into prompts for DiffusionGemma behind an OpenAI-compatible endpoint, laya runs its own checkpoints in-process with no LLM at all.

layalocaljev
Stars30k792
Forks2.6k53
LanguagePythonTypeScript
LicenseApache-2.0MIT
Last activitytoday13 days ago
Topicsdecision-models, localdecision-models, gateway, local
Curated connections57

laya — the curator's take

Use laya where you now spend an LLM call on a classification: routing a ticket, scoring urgency, a yes/no guard. Typed questions in, calibrated probabilities out, in one pass of about 33 ms, in 100+ languages, with pip install and extras for LangGraph, LlamaIndex, CrewAI, MCP and an HTTP server. The probabilities are the point: you can threshold them. Mind the limits it publishes itself: the multilingual checkpoint cuts at 1,024 tokens unless you pass max_len=8192, and past about 4,000 tokens its own long-document bench drops to 8 to 17 correct of 20. It decides; it does not explain or extract, so free-form answers still need an LLM. Two weeks old despite the star count: pin the version.

localjev — the curator's take

Use LocalJev when you want to point TypeSafe's SDK at your own hardware today, with a model you already serve: set TYPESAFE_BASE_URL and the quickstart runs unchanged, with chunking, admission control (max-inflight, HTTP 529 queue), and corrective retries on malformed JSON already handled. Read the honesty in its own README before trusting it: the probabilities are *generated* by the model as a JSON scalar/vector, not read from logits, so it is wire-compatible with Jev but not mathematically equivalent — OpenJev's structured read needs unmerged vLLM extensions that oMLX doesn't expose. That makes it fine for routing and triage, and the wrong tool for consequential decisions until you have run its own bake-off (AG News / BoolQ / SST-5, five models, two input lengths) on your workload. If you would rather own calibrated probabilities than borrow them, kev trains a readout head instead of prompting.