Laya: Multilingual, non-autoregressive System 1 decision engine Laya released a multilingual, non-autoregressive "System 1" decision engine that returns typed decisions across 100+ languages in a single forward pass in 33 ms, installable via `python -m pip install laya` on Python 3.10 or newer. The library's `laya-multilingual` checkpoint reads up to 8,192 tokens with `max_len=8192`, and a router selects the appropriate checkpoint per request, sending non-English text to the multilingual model. Optional extras include `laya[serve]`, `laya[mcp]`, `laya[langchain]`, `laya[llamaindex]`, `laya[crewai]`, `laya[onnx]`, and `laya[fast]`, with a TypeScript/Node.js/browser port published to npm as `laya-ts`. Multilingual, non-autoregressive System 1 decision engine. Typed decisions over 100+ languages in a single forward pass — 33 ms — trained with reinforcement learning against strictly proper scoring rules RLCD , with a router that picks the right checkpoint per request. python -m pip install laya With uv https://docs.astral.sh/uv/ , run uv add laya in a uv project or uv pip install laya in a virtual environment. Python 3.10 or newer. Optional extras: laya serve HTTP server , laya mcp MCP server , laya langchain LangChain and LangGraph , laya llamaindex LlamaIndex selectors , laya crewai CrewAI routing , laya onnx ONNX Runtime , laya fast TileLang GPU fast path . Step-by-step setup for each platform, CPU-only or GPU PyTorch builds, and troubleshooting are in Installation details installation-details . For TypeScript / Node.js / browser, see laya-ts/ https://github.com/NandhaKishorM/laya/blob/main/laya-ts . npm releases npm install laya-ts are published from this repository's laya-ts-v release tags. Long documents. laya-multilingual reads up to 8,192 tokens with max len=8192 . Measured accuracy and time by document length, reproducible with research/scripts/bench long context.py https://github.com/NandhaKishorM/laya/blob/main/research/scripts/bench long context.py : Long documents: laya-multilingual reads up to 8,192 tokens. It ships with a 1,024-token limit that cuts long documents off, so pass max len=8192 for them: result = router.predict long document, questions, model="multilingual", max len=8192 In the table above, 16 to 18 of 20 requests were answered correctly with up to about 4,000 tokens of text before them; beyond that results vary 8 to 17 of 20 , so check long-document accuracy on your own data. Short inputs give identical answers with max len=8192 , and speed follows the input's real length, not the limit: short inputs are unchanged, and a 4,000-token input takes about 1.7 s on an Apple GPU. Name the checkpoint with model="multilingual" , since long mostly-English text would otherwise route to the English checkpoint. python from laya import Router router = Router downloads a checkpoint on first use; Router preload=True loads all three up front state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan." questions = { "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "other": "everything else"}}, "urgency": {"type": "score", "instructions": "How urgent is this?", "criteria": "not urgent", "soon", "blocking" }, "churn risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"}, } result = router.predict state, questions print result "answers" "department" "choice" billing print result "answers" "churn risk" "noul" probability the answer is yes print result "routing" "model" english The same call works in any of 100+ languages. The Router detects the script and language and sends non-English text to laya-multilingual : for text in "मुझसे मार्च में दो बार शुल्क लिया गया, कृपया डुप्लिकेट राशि वापस करें।", "La aplicación se cierra cada vez que abro la configuración." : r = router.predict text, {"department": questions "department" } print r "routing" "model" , r "answers" "department" "choice" multilingual billing multilingual technical From the command line, laya "My payment failed twice" --preset triage answers a ready-made question set. More in the full quickstart quickstart-route-mode-recommended and the docs https://nandhakishorm.github.io/laya/ . The shipped checkpoints work zero-shot, but fine-tuning on decisions from your own domain is where accuracy jumps. On the typed-decisions benchmark 2,000 decisions across four workflows , the fine-tuned laya-typed-decisions checkpoint scores 0.766 accuracy, against 0.362 for the base English checkpoint on the same decisions. Fine-tuning notebook https://github.com/NandhaKishorM/laya/blob/main/notebooks/laya finetune typed decisions 2xT4 kaggle.ipynb : runs the whole loop on Kaggle's free 2x T4 GPUs build the dataset, train, fit calibration temperatures, evaluate, and push the result to the Hub . Details in Fine-Tuning fine-tuning . nandhakishorm.github.io/laya https://nandhakishorm.github.io/laya/ : guides for prediction hooks https://nandhakishorm.github.io/laya/hooks/ , schema-driven decisions https://nandhakishorm.github.io/laya/structured/ , Docker https://nandhakishorm.github.io/laya/docker/ and LangChain and LangGraph https://nandhakishorm.github.io/laya/langchain/ , plus a full API reference https://nandhakishorm.github.io/laya/reference/ . - ONNX catches up with PyTorch. ONNXAgent gains predict batch with sort by length , predict long and decide batch , scripts/export onnx.py --quantize writes a per-channel INT8 copy for CPU, and laya-evals run --onnx scores an export with the same gates as the torch path. - Opt-in abstention. min confidence= on predict , predict batch , decide and decide batch flags answers below a threshold on answer confidence with low confidence: True , and decide returns None for them. - Batch everywhere. decide batch , Router.predict long , laya --batch FILE , the MCP laya predict batch / laya route batch / laya decide tools, and LangChain batch / abatch all run on shared forward passes. New LayaDecision LangChain , LlamaIndex selectors laya llamaindex and CrewAI routing laya crewai . - Per-request token budget. max len / head max len now reach every surface: laya-serve capped by LAYA MAX TOKEN BUDGET , Router.predict batch requests, the CLI --questions , --max-len , --head-max-len , MCP tools and LangChain nodes. - Operations. LAYA MAX LOADED , LAYA REVISION and per-checkpoint SHA-256 maps; /health reports the device a checkpoint really runs on and its CPU-fallback count; the 503 busy answer carries Retry-After ; compile=True no longer recompiles for every request shape. - Stricter inputs. A null or duplicate choice label, a short temperature list, a None state and non-dict questions are refused with a message that names them, and usage "options" says when the head budget left two options with the same tokens. Laya evaluates typed questions choice , score , noul over any state text, email, ticket or JSON document in a single forward pass — 33 ms for one question, 7.2 ms/question batched, measured on a T4. No text generation, so nothing to parse and nothing to hallucinate. Three checkpoints, and a Router that picks between them per request: | | encoder | params | context | use it for | |---|---|---|---|---| | laya https://huggingface.co/convaiinnovations/laya | ModernBERT-large | 421M | 512 | English | | laya-multilingual https://huggingface.co/convaiinnovations/laya-multilingual | mmBERT-base | 322M | 1024 up to 8,192 | 100+ languages, 2x faster | | laya-typed-decisions https://huggingface.co/convaiinnovations/laya-typed-decisions | ModernBERT-large | 421M | 1024 | the typed-decisions workflows | Python 3.10 or newer. The dependencies set that floor: huggingface hub 1.x, transformers 5.x and torch 2.14 all require 3.10. Optional PyTorch build selection: If you need a CPU-only or GPU-specific PyTorch build, follow PyTorch's installation guide https://pytorch.org/get-started/locally/ after creating your virtual environment and before installing Laya. Replace pip or pip3 in the selected command with the environment's Python executable followed by -m pip . If you already use a virtual environment, install the PyPI release with: python -m pip install laya For a new environment, choose the commands for your platform below. Run them from your project directory; the explicit Python paths keep installation and verification in the same environment. macOS / Linux with Python 3.10 or newer : On Debian/Ubuntu, the system Python may require sudo apt install python3-venv before creating a virtual environment. If venv reports that ensurepip is unavailable, install that package and retry. python3 -m venv .venv .venv/bin/python -m pip install laya .venv/bin/python -I -c "import laya; print laya. version " Windows PowerShell this example uses an installed Python 3.11 : py -3.11 -m venv .venv .\.venv\Scripts\python.exe -m pip install laya .\.venv\Scripts\python.exe -I -c "import laya; print laya. version " Both checks print the installed Laya version without loading a checkpoint. -I excludes the current directory from the import search path, so a local source copy cannot mask a missing installation. Keep using the same virtual environment's Python when running your application. Intel GPU XPU Install a supported Intel GPU driver first. For an XPU-enabled PyTorch build, install its wheel before Laya; the default PyPI wheel may be CPU-only. PyTorch's validated hardware and OS list is in the Intel GPU guide https://docs.pytorch.org/docs/2.14/notes/get start xpu.html . .\.venv\Scripts\python.exe -m pip install torch --index-url https://download.pytorch.org/whl/xpu .\.venv\Scripts\python.exe -m pip install laya .\.venv\Scripts\python.exe -c "import torch; print torch.xpu.is available " For a source checkout, replace pip install laya with pip install -e . . Laya automatically selects an available XPU when no device is specified; you can also request one explicitly with device="xpu" in laya.load or Router device="xpu" . Install from GitHub To use the development version instead of the PyPI release, create the virtual environment above and replace its installation command with the appropriate command below. Git must be installed. macOS / Linux .venv/bin/python -m pip install "git+https://github.com/NandhaKishorM/laya.git" Windows PowerShell .\.venv\Scripts\python.exe -m pip install "git+https://github.com/NandhaKishorM/laya.git" Run the same version check afterward. The GitHub version follows the repository's default branch and may differ from the published release. Install with uv uv https://docs.astral.sh/uv/ creates the virtual environment, downloads a matching Python if none is installed, and installs into it. The commands are the same on macOS, Linux and Windows PowerShell: uv venv --python 3.12 uv pip install laya uv pip install targets the .venv in the current directory without activating it, so run the version check for your platform above afterward. Extras and the GitHub version install the same way: uv pip install "laya serve " , uv pip install "git+https://github.com/NandhaKishorM/laya.git" . For a CPU-only or GPU-specific PyTorch build, add --torch-backend=auto to pick the build that matches the machine's GPU driver, or name one such as --torch-backend=cpu ; uv's PyTorch guide https://docs.astral.sh/uv/guides/integration/pytorch/ has the list. If your application is a uv project, add Laya as a dependency instead: python uv add laya uv run python -I -c "import laya; print laya. version " --torch-backend applies to uv pip only; in a uv project, uv's PyTorch guide shows how to set the PyTorch index in pyproject.toml . Model setup and troubleshooting Continue with the Router quickstart quickstart-route-mode-recommended to run inference. Loading a Hub checkpoint requires access to Hugging Face on its first download; the quickstart's Router preload=True loads all three configured checkpoints at construction. - ModuleNotFoundError: No module named 'laya' : run both installation and your script with the same virtual environment's Python executable shown above. In an editor, select that interpreter as well. - Missing rl agent config.json : this file ships with a Laya checkpoint alongside model.safetensors ; it is not a configuration file you need to create in the source repository. For a local model, pass the directory containing those checkpoint files. Installing the package also installs a laya command for quick local testing, no script needed: laya "I was charged twice, please refund" routing decision only; works offline, no download laya "Refactor this service" --predict full answers downloads the checkpoint on first use laya "Mein Konto wurde zweimal belastet" --lang de force a language instead of detecting it laya "My payment failed twice" --model ml pin a checkpoint: names, aliases and casing all resolve as the SDK resolves them laya "My payment failed twice" --preset triage answer a ready-made preset triage, email, guard, moderation, router laya --batch tickets.txt --predict score a file of requests, one per line, in one batch cat tickets.txt | laya --batch - --predict --json stdin; one JSON line of answers per request laya "Where is my card" --questions intents.json answer your own questions, written in a JSON file laya interactive mode Routing alone never downloads a checkpoint, so it returns in milliseconds. --predict loads the routed checkpoint, which needs network access to the Hugging Face hub the first time; if a checkpoint cannot be downloaded, the CLI says so instead of crashing. --batch with or without --predict sends the whole file through Router.predict batch in one process, so the requests share checkpoint loads and forward passes — measured 2.6x on 20 tickets vs looping predict one by one, with --batch-size N to bound the forward pass, --sort-by-length to group similarly sized requests inside it, and --json for JSONL output. Batch routing laya --batch FILE , no --predict likewise answers with route batch in one pass, still without loading anything. --sort-by-length reaches the length grouping Agent.predict batch has done since 294, which until now only the library call in Batch Mode batch-mode-score-many-states-in-one-forward-pass could ask for. A forward pass pads every state in it to the longest one, so a two-line request sharing a pass with a two-paragraph one spends most of its compute on padding; grouping by length first puts comparable sizes together. Measured on 128 real Yelp reviews of 69-2,293 characters, run as laya --batch FILE --predict --model laya --batch-size 8 with and without the flag, four interleaved rounds on an Apple-silicon Mac at the CLI's own device pick: 41.25 s median unsorted against 29.23 s sorted, a paired median of 1.42x 1.39-1.44x across the four pairs , with all 128 decisions identical and still in input order. research/ measures the same knob at 2.15x over 10,000 synthetic multilingual tickets. It needs --batch-size N with 1 < N < the number of requests — a single pass over the whole file has no group to reorder — and the CLI says so on stderr rather than printing a same-speed result and leaving you to notice. --questions takes the same question dict the SDK takes, as JSON: either the mapping itself, or {"state key": "body", "questions": {...}} when the question's instructions name a field other than request . It implies --predict , and a question set with many labels usually wants --head-max-len with it: on 58 MASSIVE-INTENT labels written as one choice question, the English checkpoint goes from 24/58 correct at its default 192-token option budget to 34/58 at --head-max-len 384 , for about 1.4x the per-request time on CPU. Widening it further costs the accuracy back, because max len then leaves fewer tokens for the request itself. Honest limits honest-limits describes the same budget ceiling for a 77-option question. examples/server.py is a self-contained FastAPI app for testing Laya without writing any code: a two-pane playground edit the request as a form or as JSON, run it with Ctrl+Enter, read each answer's full distribution and calibrated confidence, copy it as curl or Python , plus a plain JSON API /predict , /predict/batch for scripting against. pip install "laya serve " python examples/server.py http://127.0.0.1:8000 Open http://127.0.0.1:8000 in a browser for the playground, or hit it directly: curl -s localhost:8000/predict -H 'content-type: application/json' -d '{ "state": {"body": "We were billed twice for March. Please refund it today."}, "questions": { "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "other": "everything else"}}, "urgency": {"type": "score", "instructions": "How urgent is this?", "criteria": "not urgent", "soon", "critical" } } }' | python -m json.tool --no-preload loads checkpoints lazily instead of all three up front; --device cuda|cpu|mps pins the device. See python examples/server.py --help for the rest. A request body may carry model to pin a checkpoint instead of letting the router choose, and it takes exactly the spellings laya --model takes: names, aliases and casing all resolve through laya.router . GET /models lists the checkpoints and the aliases alongside them. /predict/batch accepts two optional body fields that control the shape of the forward passes without changing any answer: batch size states per pass; omit it and the whole batch is one pass and sort by length group similarly sized states so each pass pads to a shorter maximum — it needs a batch size below the number of states, since one pass has nothing to reorder . Measured on the running server: 64 real reviews of 91–2,293 characters, one choice question, batch size: 8 , three interleaved rounds on MPS — the default one-pass shape 5.35s median against 3.04s sorted, a paired median of 1.77x min 1.76x , and 64/64 labels identical in input order. batch size: 8 on its own was 4.95s, so the length grouping is what earns the second factor: 1.66x over the sized batch. laya-client https://github.com/NandhaKishorM/laya/blob/main/sdk/typescript/README.md supports JavaScript and TypeScript applications over HTTP to a self-hosted laya-serve /v1/systemone server. It includes inferred answer types, all five question presets, ESM/CommonJS exports, and cancellation. From this checkout, install and start the Python inference service: pip install -e '. serve ' LAYA HOST=127.0.0.1 LAYA MODELS=english laya-serve In another terminal: cd sdk/typescript npm ci npm run build node examples/triage.mjs The npm package is named laya-client . Use it when a JavaScript or TypeScript application talks over HTTP to self-hosted Python laya-serve . Use laya-ts when inference must run directly inside JavaScript through its local ONNX runtime, without a Python server. See installation, examples, and publishing https://github.com/NandhaKishorM/laya/blob/main/sdk/typescript/README.md and the repository analysis https://github.com/NandhaKishorM/laya/blob/main/docs/typescript-sdk.md for architecture and scope. To try the Python SDK in a CPU container, see the Docker Compose quickstart https://github.com/NandhaKishorM/laya/blob/main/docs/docker.md . It runs a sample request and keeps downloaded models between runs. Laya ships three checkpoints. The built-in Router is the recommended entry point: it evaluates any state in any language, automatically detects scripts and languages in sub-milliseconds, and dispatches to the optimal checkpoint in a single forward pass. python from laya import Router Preload checkpoints into memory for instant sub-35ms routing router = Router preload=True 1. State in any language or schema state = { "from": "user@acme.com", "subject": "Duplicate charge on invoice 4411", "body": "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan." } 2. Define your typed questions questions = { "department": { "type": "choice", "instructions": "Which department should handle this request?", "criteria": { "billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "sales": "pricing, new contracts", "other": "everything else" } }, "urgency": { "type": "score", "instructions": "How urgent is this request?", "criteria": "not urgent", "soon", "critical deadline or blocking issue" }, "churn risk": { "type": "noul", "instructions": "Does the user threaten to cancel or leave?" }, "refund requested": { "type": "noul", "instructions": "Does the user explicitly request a refund?" } } 3. English state - automatically routed to laya ModernBERT-large, 39.5 ms res en = router.predict state, questions print "Department :", res en "answers" "department" "choice" - billing confidence: 0.94 print "Routing :", res en "routing" "model" - english 4. Hindi state - automatically routed to laya-multilingual mmBERT-base, 32.8 ms res hi = router.predict {"body": "मुझसे दो बार शुल्क लिया गया, कृपया पैसे वापस करें।"}, questions print "Department :", res hi "answers" "department" "choice" - billing confidence: 0.86 print "Routing :", res hi "routing" "model" - multilingual 5. Explicit override when you want a specific checkpoint res td = router.predict state, questions, model="typed-decisions" Every result carries full routing metadata explaining why the choice was made: res hi "routing" { 'model': 'multilingual', 'repo': 'convaiinnovations/laya/multilingual', 'reason': 'non-Latin script devanagari, 100% of letters ; the English checkpoint cannot read it' } Inspect a routing decision without running any forward pass: router.route {"body": "Der Kunde wurde zweimal belastet"}, questions .reason "Latin script but language looks like 'de', not English" Very short Latin-script text often carries nothing that identifies its language "Quero cancelar" , "Esqueci minha senha" . Such text goes to default , which is "english" unless you change it. If most of your traffic is not English, set: router = Router default="multilingual" router.route {"body": "Esqueci minha senha"} .model - multilingual router.route {"body": "Please refund the duplicate charge"} .model - english If a lazy router receives an interleaved workload whose requests route to different checkpoints, calling predict in a loop can still cause unnecessary checkpoint churn when the required checkpoints exceed the resident cache, for example with max loaded=1 or when typed-decisions is also used. Router.predict batch routes the full workload first, groups requests by checkpoint, then groups requests with the same question schema within each checkpoint. Each compatible group is dispatched to Agent.predict batch so states can share forward passes, and results are restored to the original request order. requests = {"state": "Please refund invoice 1", "questions": questions}, {"state": "تم خصم المبلغ مرتين", "questions": questions}, {"state": "Please refund invoice 2", "questions": questions}, results = Router max loaded=1 .predict batch requests results stay in input order while compatible requests are batched by checkpoint Each item can independently set model , task , lang , lang guess , or the token budget max len , head max len . Use route batch requests when you only want the ordered routing decisions without loading any checkpoint. predict many is an alias for predict batch . Requests are validated before model loading. Different requests may use different question schemas; requests sharing a checkpoint, question schema and token budget are passed together to Agent.predict batch . Prediction hooks prediction-hooks installed on the Router run once per request, as they do for predict , so a redaction hook rewrites every state before the model sees it. Requests that share a checkpoint run all their start hooks before their shared forward pass; see docs/hooks/lifecycle.md https://github.com/NandhaKishorM/laya/blob/main/docs/hooks/lifecycle.md routerpredict batch . You can also bound the Agent-level forward-pass batch size, and forward the length grouping knob to every group see Batch Mode batch-mode-score-many-states-in-one-forward-pass : results = router.predict batch requests, batch size=8, sort by length=True On a shared benchmark 17,416 questions, one T4 GPU, identical questions per model : | Benchmark / Task | English laya | Multilingual laya-multilingual | Router Routed | |---|---|---|---| | MASSIVE intent, English | 0.783 | 0.657 | 0.783 | | MASSIVE intent, 13 other languages | 0.306 | 0.451 | 0.451 | | XNLI, English | 0.860 | 0.843 | 0.860 | | XNLI, 14 other languages | 0.521 | 0.731 | 0.731 | | Languages usable 3x random | 23 / 51 | 48 / 51 | 48 / 51 | | Latency, 1 question T4 GPU | 39.5 ms | 32.8 ms | 32.8 ms | | Latency, 10 questions batched | 158.6 ms | 72.3 ms | 72.3 ms | The English checkpoint collapses on non-Latin scripts Khmer scores 0.000 accuracy at 0.952 confidence , the raw-temperature figure; 0.705 as served after the clamp . Because the model stays confident while being wrong, confidence gating cannot save you. Router detects the script in <0.5 ms pure Python before the forward pass. A cold checkpoint build costs seconds; language detection costs microseconds. The lazy default keeps two checkpoints resident — english and multilingual , the only two automatic routing chooses between — so a language flip costs detection only once each has been built. max loaded=1 rebuilds the checkpoint it just evicted on every switch measured at a 7.4 s median reload on CPU and 10.3 s on T4 , and traffic that only ever sees one language never builds the second, so the default costs a single-language deployment nothing. For a server or production app, preload: Every checkpoint resident in memory; language flips cost detection only <1 ms router = Router preload=True router = Router preload=True, device="cuda" Or preload only the specific checkpoints you serve: router.preload "english", "multilingual" If your app already built an agent, attach it to avoid duplicate VRAM: router.attach "english", existing agent Manage resident memory default keeps two hot: english + multilingual, LRU eviction router = Router max loaded=3 keep all three hot, e.g. with auto task detection router = Router max loaded=1 memory-constrained host, reloads on every switch router.unload free memory | Deployment Mode | Per-Request Latency | Model Reloads | |---|---|---| | Router lazy, max loaded=2 | detection only <1 ms on a switch, after each language's first load | 1 the first time a language appears | | Router max loaded=1 | 7 to 10 s on every language switch | 1 per switch | | Router preload=True | 32.8 ms GPU / 193–464 ms CPU | none | A rebuild still re-reads the checkpoint, but each checkpoint's tokenizer is parsed once per process and reused by every Agent — including one the Router rebuilds after eviction. The multilingual tokenizer.json alone is 34 MB / 256k vocab, several times the cost of applying its weights. Preloading is still the right answer for a server: it removes the rebuild rather than making it cheaper. Routing asks one question: can the English checkpoint read this state? The built-in detector answers it from the script and a function-word heuristic, and is deliberately dependency-free. That heuristic is best-effort on Latin-script languages it holds no word list for, so a short request can carry no usable signal: python from laya.lang import analyse analyse "Care este ora in Tokyo?" {'script': 'latin', 'language': 'en', 'is english': True} - the English checkpoint If you already run a language-identification model, hand routing the answer instead of relying on the heuristic. lang guess takes a language code or a callable receiving the state, and is checked after an explicit lang= and before detection: A code you already know router.predict state, questions, lang guess="ro" A callable, e.g. wrapping fastText, CLD3 or a transformer LID router.predict state, questions, lang guess=lambda s: my lid s Or install one for every request on a server router = Router preload=True, lang guess=my lid The hint only decides English or not : a code whose primary subtag is en , eng or english routes to the English checkpoint, and every other code that names a language routes to the multilingual one. "en US" and "en US.UTF-8" are read as English, so $LANG can be passed straight through. Returning None , or a code that names no language, makes it abstain and the built-in detector decides as before — so a LID model that is unsure does not force a checkpoint. C , POSIX and C.UTF-8 abstain, which matters because C.UTF-8 is the default $LANG in the official Python image: passing it through no longer pins every request to the multilingual checkpoint, which is what it used to do. The ISO 639-2 special codes und , zxx and mul abstain for the same reason. An explicit model= , task= or lang= still wins, and the default path is unchanged. laya.serve exposes the Router over HTTP on the same POST /v1/systemone wire protocol as TypeSafe's hosted Jev API. Laya's answer payload is already schema-identical to what Jev returns choice / score / noul answers and a {input tokens, output tokens} usage block , so an existing Jev client — e.g. the hs-jev https://github.com/getmissionctrl/hs-jev Haskell client — just needs its baseUrl repointed; nothing else changes. pip install "laya serve " adds fastapi + uvicorn + python-multipart LAYA DEVICE=cuda LAYA PRELOAD=1 laya-serve binds 0.0.0.0:8000, preloads all 3 checkpoints curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{ "state": {"body": "billed twice, refund please or we cancel"}, "questions": {"dept": {"type": "choice", "instructions": "which team?", "criteria": {"billing": "refunds", "tech": "bugs"}}} }' To evaluate multiple states in a single call, send a states array to POST /v1/systemone/batch capped at 64 states : curl -s localhost:8000/v1/systemone/batch -H 'content-type: application/json' -d '{ "states": {"body": "billed twice, refund please"}, {"body": "cannot login, getting 500 error"} , "questions": {"dept": {"type": "choice", "instructions": "which team?", "criteria": {"billing": "refunds", "tech": "bugs"}}} }' The endpoint shares forward passes via Router.predict batch and returns an array of results in matching order along with aggregated total usage . Configuration is by environment variable: LAYA HOST , LAYA PORT , LAYA DEVICE , LAYA PRELOAD , LAYA MODELS comma list to preload , LAYA THREADS cap torch intra-op threads for CPU inference — keep at or below physical cores , LAYA AUTO TASK , LAYA MAX LOADED checkpoints resident at once, 2 by default; raise it to 3 when LAYA AUTO TASK makes a third one reachable on demand, or the server rebuilds one every time routing switches , and LAYA API KEY when set, clients must send Authorization: Bearer