Multilingual, non-autoregressive System 1 decision engine. Typed decisions over 100+ languages in a single forward pass — 33 ms — trained with reinforcement learning against strictly proper scoring rules (RLCD), with a router that picks the right checkpoint per request.
Installation
python -m pip install laya
With uv, run uv add laya in a uv project or uv pip install laya in a virtual environment.
Python 3.10 or newer. Optional extras: laya[serve] (HTTP server), laya[mcp] (MCP server), laya[langchain] (LangChain and LangGraph), laya[llamaindex] (LlamaIndex selectors), laya[crewai] (CrewAI routing), laya[onnx] (ONNX Runtime), laya[fast] (TileLang GPU fast path). Step-by-step setup for each platform, CPU-only or GPU PyTorch builds, and troubleshooting are in Installation details.
For TypeScript / Node.js / browser, see laya-ts/. npm releases (npm install laya-ts) are published from this repository's laya-ts-v* release tags.
Long documents. laya-multilingual reads up to 8,192 tokens with max_len=8192. Measured accuracy and time by document length, reproducible with research/scripts/bench_long_context.py:
Quickstart
Long documents:
laya-multilingualreads up to 8,192 tokens. It ships with a 1,024-token limit that cuts long documents off, so passmax_len=8192for them:result = router.predict(long_document, questions, model="multilingual", max_len=8192)In the table above, 16 to 18 of 20 requests were answered correctly with up to about 4,000 tokens of text before them; beyond that results vary (8 to 17 of 20), so check long-document accuracy on your own data. Short inputs give identical answers with
max_len=8192, and speed follows the input's real length, not the limit: short inputs are unchanged, and a 4,000-token input takes about 1.