StackMap
Subscribe
Explore / laya
NandhaKishorM

laya

Non-autoregressive decision engine: typed choice, score and yes/no answers over text in one forward pass (~33 ms), 100+ languages, calibrated probabilities, a router picking the checkpoint.

29,781 2,581 Python Apache-2.0updated today
View on GitHubDispute this mapping →
Curator's take

Use laya where you now spend an LLM call on a classification: routing a ticket, scoring urgency, a yes/no guard. Typed questions in, calibrated probabilities out, in one pass of about 33 ms, in 100+ languages, with pip install and extras for LangGraph, LlamaIndex, CrewAI, MCP and an HTTP server. The probabilities are the point: you can threshold them. Mind the limits it publishes itself: the multilingual checkpoint cuts at 1,024 tokens unless you pass max_len=8192, and past about 4,000 tokens its own long-document bench drops to 8 to 17 correct of 20. It decides; it does not explain or extract, so free-form answers still need an LLM. Two weeks old despite the star count: pin the version.

Mapped by ShipWithAI editors · links verified

Continue your stack

What teams reach for next — and why each earns a place beside laya. Ranked by curator confidence.

pairs wellpairs wellalternativealternativealternativeLangGraphLlamaIndexkevAnyJevlocaljevlaya
pairs wellalternativebuilt withpick a node for the why · open it from the panel
Weekly digest
README.md1 min read

Laya

Multilingual, non-autoregressive System 1 decision engine. Typed decisions over 100+ languages in a single forward pass — 33 ms — trained with reinforcement learning against strictly proper scoring rules (RLCD), with a router that picks the right checkpoint per request.

Open In Colab PyPI version Docs Hugging Face Model Multilingual Hugging Face Space Dev.to Article Buy Me A Coffee License

Installation

python -m pip install laya

With uv, run uv add laya in a uv project or uv pip install laya in a virtual environment.

Python 3.10 or newer. Optional extras: laya[serve] (HTTP server), laya[mcp] (MCP server), laya[langchain] (LangChain and LangGraph), laya[llamaindex] (LlamaIndex selectors), laya[crewai] (CrewAI routing), laya[onnx] (ONNX Runtime), laya[fast] (TileLang GPU fast path). Step-by-step setup for each platform, CPU-only or GPU PyTorch builds, and troubleshooting are in Installation details.

For TypeScript / Node.js / browser, see laya-ts/. npm releases (npm install laya-ts) are published from this repository's laya-ts-v* release tags.

Long documents. laya-multilingual reads up to 8,192 tokens with max_len=8192. Measured accuracy and time by document length, reproducible with research/scripts/bench_long_context.py:

laya-multilingual with max_len=8192: 16 to 18 of 20 requests correct with up to about 4,000 tokens of text before them, more variable beyond

Quickstart

Long documents: laya-multilingual reads up to 8,192 tokens. It ships with a 1,024-token limit that cuts long documents off, so pass max_len=8192 for them:

result = router.predict(long_document, questions, model="multilingual", max_len=8192)

In the table above, 16 to 18 of 20 requests were answered correctly with up to about 4,000 tokens of text before them; beyond that results vary (8 to 17 of 20), so check long-document accuracy on your own data. Short inputs give identical answers with max_len=8192, and speed follows the input's real length, not the limit: short inputs are unchanged, and a 4,000-token input takes about 1.