Making Intent Routing Explainable A four-tier intent-routing cascade ordered by cost, with each tier allowed to abstain and every exit leaving a replayable log record, can replace a single frontier-model llm.classify(query) call for a large share of traffic, according to a design proposal that puts regex matching at tier one. The proposal argues that System One decision models (Jev, Laya, and contrastive dual-encoders such as CLM) return calibrated probabilities over defined options in a single forward pass without generating tokens, unlike LLM routers whose post-hoc explanations cannot be replayed or checked. The design keeps the frontier model for requests that need reasoning, and logs rule identifiers such as regex_txn_id_v1 so routing decisions remain auditable months later. The most expensive line in most intent-routing designs is the one that says llm.classify query . It looks harmless in a design doc. Every request goes to a frontier model, the model returns a label, a handler fires. The demo works on the first try. Then real traffic arrives and three things happen together: P99 latency climbs past what the product can tolerate, the inference bill grows linearly with a number the business is actively trying to grow, and two identical /cancel requests occasionally land on different handlers. Take that seriously and the question changes. You stop asking which model is smartest and start asking which is the cheapest engine that can make this decision correctly, with a confidence you can trust and a record you can replay. For a surprising share of traffic, that engine is a regex. The replay part matters more than it sounds. Ask an LLM router why it sent a refund request to the sales queue and it will write a fluent paragraph, generated after the fact by the same sampling process that made the decision. Nothing in that paragraph can be replayed or checked. System One decision models Jev, Laya, and contrastive dual-encoders such as CLM change this. They return calibrated probabilities over options you define, in a single forward pass, and never generate a token. What follows is the cascade I would put in front of any system routing a million intents a day: four tiers, ordered by cost, each allowed to abstain, every exit leaving a record that explains it, and the frontier model kept for the requests that need reasoning. For routing, an explanation is the answer to four questions you can answer from logs alone, months later: If any of the four is missing, you have a guess with a timestamp. An LLM router can answer the first two. Some LLM APIs also return token log-probabilities for the label, but those are per token, shift with the prompt, are not calibrated and are not exposed everywhere. System One models can answer all four by construction, provided you design the questions and the policy deliberately. Most of this article is about that design. Before a request pattern gets a model, it gets judged on four axes: The tiers line up in order of cost and latency: The arrows only point one way. A tier either answers with confidence above its threshold or passes the request down. Nothing climbs back up. If a pattern is invariant, match it. Do not predict it. Slash commands /cancel, /help, /reset , transaction IDs, tracking numbers, structured payload keys: anything where a false positive is unacceptable and the vocabulary never drifts belongs here. A request like "What is the status of TXN 99281X?" arriving with "transaction id": "TXN 99281X" already in the payload does not need a GPU to rediscover a string you already have. Explainability at this tier is free. The rule id is the whole explanation. Log it with a version regex txn id v1 and "why did this go there" becomes a grep. The failure mode is brittleness: typos, paraphrase and anything resembling natural language slip past. That is acceptable, because a miss here costs nothing. The request simply falls through. python import refrom typing import Optional, Tuple MAX INPUT CHARS = 2 000 bound regex work on untrusted input COMMANDS = {"/cancel": "handler cancel subscription", "/help": "handler display help"}TXN = re.compile r"\bTXN A-Za-z0-9 {5,10}\b" def tier1 query: str - Optional Tuple str, str : """Returns handler, rule id or None. The rule id is the whole explanation.""" query = query :MAX INPUT CHARS Match the whole first token, never a prefix of it, so "/cancellation policy?" does not fire the /cancel handler in the one tier that promises zero false positives tokens = query.strip .lower .split maxsplit=1 if tokens and tokens 0 in COMMANDS: return COMMANDS tokens 0 , "cmd trie exact" if TXN.search query : return "handler transaction status lookup", "regex txn id v1" return None --- Inference ---print tier1 "/cancel" 'handler cancel subscription', 'cmd trie exact' print tier1 "/cancellation policy?" None falls through to Tier 2 print tier1 "Check status for TXN 99281X" 'handler transaction status lookup', 'regex txn id v1' Regression guard for the prefix bugassert tier1 "/cancellation policy?" is Noneassert tier1 "/cancel please" 0 == "handler cancel subscription" Two details matter. The command lookup matches the whole first token, never a prefix. A naive trie walk returns as soon as it reaches a stored command, so “/cancellation policy?” would fire the cancel handler, in the one tier that promises zero false positives. For exact commands, a dictionary keyed on the whole token is a trie with that bug designed out, and the asserts at the bottom pin the behaviour so a later refactor cannot quietly bring it back. And every regex here runs on untrusted input, so the input is capped before matching; a catastrophic backtracking pattern turns your fastest tier into a CPU denial-of-service. Once phrasing starts to vary, rules break. Classical ML is the next cheapest thing that can generalise: CPU-only, no GPU, and explanations that come straight out of the model. A linear model over sparse n-gram features. It wins when specific tokens correlate strongly with a class. Support tickets that mention “invoice”, “receipt” or “charged” are billing tickets, and you do not need attention heads to know that. The coefficients tell you which n-grams pushed a ticket into billing. The model is rarely the slow part here: scikit-learn measures feature extraction at 100 to 500 times https://scikit-learn.org/stable/computing/computational performance.html the cost of the prediction itself, so tune the vectoriser before the classifier. python from sklearn.feature extraction.text import TfidfVectorizerfrom sklearn.linear model import LogisticRegressionfrom sklearn.pipeline import Pipeline TARGET = {"billing": "billing queue", "account access": "account access queue", "bug report": "technical queue"} X train = "where is my invoice receipt", "invoice receipt missing", "send me the invoice", "billing receipt for last month", "need a copy of my invoice", "charged twice on my card", "refund the extra charge", "my receipt is wrong", "cannot log into my account", "password reset link expired", "locked out of my account", "reset my password", "two factor code not arriving", "account access denied", "forgot my login email", "sign in fails", "app crashing on checkout", "getting 500 error on api call", "page will not load", "export button broken", "sync job fails every night", "webhook returns timeout", "dashboard shows blank chart", "upload keeps failing", y train = "billing" 8 + "account access" 8 + "bug report" 8 model = Pipeline "tfidf", TfidfVectorizer ngram range= 1, 2 , "clf", LogisticRegression C=20.0 model.fit X train, y train def tier2a query: str, threshold: float = 0.85 : probs = model.predict proba query 0 i = int probs.argmax label, p = model.classes i , float probs i return TARGET label if p = threshold else None , f"{label} p={p:.2f}" --- Inference ---print tier2a "where is my invoice receipt" 'billing queue', 'billing p=0.95' print tier2a "My screen flashed green and the app uninstalled itself" None, 'billing p=0.54' The None in the second call is deliberate, and it is the most important output in this section. The green-screen message shares almost no tokens with the training set, so the model’s best guess is billing, at 0.54. That guess is wrong, and the threshold is the only thing that stops a billing agent from opening a ticket about a crashed app. The router abstains and the request falls through; in Appendix B, Laya places it with the technical team at 0.92. A router that guesses on thin evidence is worse than one that says “not mine”. Gradient-boosted trees split continuous and categorical features along non-linear boundaries, and CatBoost takes raw text alongside tabular metadata natively. That matters because a lot of intent is not in the text at all. “I can’t log in” from a free-tier user with one failed attempt is a password reset. The same sentence from an Enterprise admin with five failed attempts in ten minutes is an escalation, possibly a security one. The text is identical and the right route differs. CatBoost reads failed login attempts 10m and user tier directly, and its own 2018 benchmark batch-scores rows at roughly 5 microseconds each https://catboost.ai/news/best-in-class-inference-and-a-ton-of-speedups on a single thread. That is batch throughput. A single request through Python and pandas spends far longer in overhead than in the trees, so measure the serving path end to end. CatBoost also gives you per-decision SHAP attributions for free, which is observability out of the box. python from typing import Any, Dict import pandas as pdfrom catboost import CatBoostClassifier, Pool TEXT, CATS = "query text" , "user tier" Toy-sized corpus: let every token into the text dictionary remove once you train on real volume TEXT PROCESSING = { "tokenizers": {"tokenizer id": "Space", "separator type": "ByDelimiter", "delimiter": " "} , "dictionaries": {"dictionary id": "Word", "occurrence lower bound": "1", "gram order": "1"} , "feature processing": {"default": {"dictionaries names": "Word" , "feature calcers": "BoW" , "tokenizers names": "Space" } },}train = pd.DataFrame { "query text": "can't log in", "need refund", "can't log in", "locked out of account", "refund my last charge", "forgot password" , "user tier": "Enterprise", "Free", "Free", "Enterprise", "Pro", "Pro" , "failed login attempts 10m": 5, 0, 1, 6, 0, 1 , "label": "escalate human", "billing queue", "account access queue", "escalate human", "billing queue", "account access queue" ,} model = CatBoostClassifier iterations=200, learning rate=0.1, depth=6, loss function="MultiClass", verbose=False, random seed=0, thread count=-1, text processing=TEXT PROCESSING, allow writing files=False model.fit train.drop columns= "label" , train "label" , text features=TEXT, cat features=CATS def tier2b query: str, meta: Dict str, Any , threshold: float = 0.85 : x = pd.DataFrame {"query text": query, "user tier": meta.get "user tier", "Free" , "failed login attempts 10m": meta.get "failed logins", 0 } probs = model.predict proba x 0 i = int probs.argmax Per-decision SHAP attributions for the predicted class, straight from CatBoost shap = model.get feature importance Pool x, text features=TEXT, cat features=CATS , type="ShapValues" 0 i :-1 target = model.classes i if probs i = threshold else None return target, round float probs i , 2 , {c: round float v , 3 for c, v in zip x.columns, shap } --- Inference ---print tier2b "can't log in", {"user tier": "Enterprise", "failed logins": 5} 'escalate human', 0.91, {'query text': 0.149, 'user tier': 0.177, 'failed login attempts 10m': 1.691} print tier2b "can't log in", {"user tier": "Free", "failed logins": 1} 'account access queue', 0.9, {'query text': 0.078, 'user tier': 0.451, 'failed login attempts 10m': 1.381} Look at the attribution. The sentence is identical in both calls, and the route flips. On the escalation, the query text contributed +0.149 and the login counter +1.691, more than ten times as much: the counter made the call. That is an explanation a support lead can act on, and it came out of the model itself. Classical ML fails when meaning diverges from vocabulary: indirect phrasing, negation, a message that reads like billing but is really a cancellation threat. This is where most teams jump straight to an LLM. There is now a middle layer. System One models such as Jev managed, from TypeSafe AI and Laya open source, Apache 2.0 read the input once and return typed answers without generating a token. You define the questions: a choice over named options, a score on an ordinal scale, or a noul, TypeSafe’s name for a calibrated yes/no. Each answer comes back as a probability distribution with a confidence. They are trained with Reinforcement Learning for Calibrated Decisions RLCD https://www.sanity.io/glossary/rlcd-reinforcement-learning-for-calibrated-decisions , which rewards probabilities that match reality rather than preferred text. Every System One sample below runs locally from one setup, with no API keys. The companion repository’s run.sh installs pinned versions of everything this article uses. Ollama serves the Jev-compatible model, the CLM encoder and the Tier 4 stand-in; Laya loads inside Python: One setup for every sample in this article. Needs Python 3.10 to 3.13 and Ollama 0.35 or later https://ollama.com/download on its default port: every sample talks to localhost:11434ollama serve & skip if Ollama is already runninggit clone https://github.com/indranildchandra/hybrid-intent-router.gitcd hybrid-intent-router Pinned: torch 2.14.0, laya 0.3.22, typesafe-sdk 0.7.2, catboost and scikit-learn; tev1:0.8b and qwen3:0.6b pulled into that Ollama; the Laya checkpoint; and, on machines with 16 GB of RAM and 10 GB of free disk, the CLM client without vLLM and its 8.25 GB encoder as clm-encoderHIR OLLAMA URL=http://localhost:11434 ./run.sh --setup-onlysource .venv-gpu/bin/activate .venv-cpu if run.sh found no GPU and fell back to CPU Do not ask one question “where does this go?” . Ask the questions a human triager would ask, separately: which team, how urgent, is the customer threatening to leave. Each comes back with its own calibrated distribution, and the route becomes a function of three auditable numbers instead of one opaque label. You do not need a TypeSafe account to run this. Ollama 0.35 serves open decision models behind the same /v1/systemone API as Jev https://ollama.com/blog/ollama-now-supports-jev-style-decision-models , so the official SDK works unchanged against a local base URL. The sample below uses the 0.8B tev1 model so it runs on a laptop CPU; Ollama also ships a 4B tev1 and the 9B Nimble, which it reports at 91 ms per decision on a MacBook Pro M5 Max. Moving to Jev in production is a change of base URL and model name, and the decision record should log whichever model tag actually answered. python from typesafe sdk import Choice, Noul, Score, TypeSafeClient Runs locally against tev1:0.8b, which the setup above pulls into Ollama. No TypeSafe account needed: Ollama serves open decision models behind the same /v1/systemone API as Jev.client = TypeSafeClient base url="http://localhost:11434", api key="ollama", the SDK requires a non-empty key; Ollama ignores it model="tev1:0.8b", pin the model: a decision you cannot replay is a decision you cannot explain timeout=120, the first call loads the model into memory In production on Jev: TypeSafeClient model="jev-1.13.0" with TYPESAFE API KEY set. state = { "message": "We were billed twice for invoice 99281. Refund it today or we cancel.", "user tier": "Enterprise", "open tickets": 2,} questions = { "department": Choice instructions="Which team should resolve this?", criteria={ "billing": "Duplicate charges, invoices, refunds", "technical": "Errors, outages, API failures", "sales": "Plan upgrades, seat expansions", }, , "urgency": Score instructions="How urgent is this for the customer?", criteria= "can wait", "today", "blocking" , , "churn risk": Noul instructions="Does the customer threaten to cancel or leave?" ,} resp = client.system one state=state, questions=questions dept = resp.answers "department" print dept.choice, {k: round v, 3 for k, v in dept.probabilities.items }, round dept.confidence, 3 print round resp.answers "urgency" .score, 2 , round resp.answers "churn risk" .noul, 2 billing {'billing': 0.967, 'technical': 0.033, 'sales': 0.0} 0.869 1.0 0.74 expect drift of about 0.01 across hardware Laya takes the same question shapes as plain dicts and runs on your own GPU or CPU, which matters when the state contains data you cannot send to a vendor: python from laya import Router do not name this file laya.py: it would shadow the package Self-hosted: same question shapes, plain dicts, runs on your own GPU or CPU. Pinned checkpoints for replayability laya.PINNED REVISIONS in laya 0.3.22 . The pins are per repository, so load each from its own repo: in the default bundle layout the multilingual pin does not apply and the first non-English request fails to load.router = Router revisions={ "english": "55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851", "multilingual": "e4e9ddf21a7b1903b7acffd8814ad4307bf63a67",}, standalone repos=True questions = { "department": { "type": "choice", "instructions": "Which team should resolve this?", "criteria": { "billing": "Duplicate charges, invoices, refunds", "technical": "Errors, outages, API failures", "sales": "Plan upgrades, seat expansions", }, }, "urgency": {"type": "score", "instructions": "How urgent is this for the customer?", "criteria": "can wait", "today", "blocking" }, "churn risk": {"type": "noul", "instructions": "Does the customer threaten to cancel or leave?"},} result = router.predict state, questions same state dict as the Jev exampledept = result "answers" "department" print dept "choice" , dept "probabilities" , dept "confidence" print result "answers" "churn risk" "noul" billing {'billing': 0.9688, 'technical': 0.0185, 'sales': 0.0126} 0.8546 0.5341 Asking more questions costs less than linearly. On a T4, Laya’s README measures 39.5 ms for one question and 158.6 ms for ten https://github.com/NandhaKishorM/laya in one call 72.3 ms for ten on the multilingual checkpoint , and Jev evaluates multiple questions in parallel. Decomposition stays cheap. Keep each question small, too. Laya’s own benchmarks show it at 0.425 accuracy on a 77-label intent set where Jev scores 0.870 on 72 labels, because the options share a fixed token budget. For large label sets, split the question, shortlist first, or use a dual-encoder below . The model’s job ends at probabilities. Turning probabilities into an action is a policy, and a policy belongs in versioned code that a reviewer can read and diff. The reason string below is rendered from the rules that fired. Nothing is generated. python import hashlibimport jsonfrom datetime import datetime, timezone POLICY VERSION = "routing-policy@v14" Each rule reads calibrated probabilities and returns fired, human-readable reason RULES = "churn priority", lambda a: a "churn risk" "noul" = 0.70 and a "department" "choice" == "billing", "billing with churn risk {churn:.2f} = 0.70 - retention billing queue" , "confident department", lambda a: a "department" "probabilities" a "department" "choice" = 0.90, "department={dept} at p={p dept:.2f} = 0.90" , "page on call", urgency scale 0 = can wait, 1 = today, 2 = blocking lambda a: a "urgency" "score" = 1.5, "urgency {urgency:.2f} = 1.50 near blocking - page on-call" , def decide redacted state: dict, answers: dict, model: str - dict: dept = answers "department" "choice" ctx = { "dept": dept, "p dept": answers "department" "probabilities" dept , "churn": answers "churn risk" "noul" , "urgency": answers "urgency" "score" , } fired = rid, tmpl.format ctx for rid, rule, tmpl in RULES if rule answers fired ids = {rid for rid, in fired} if "churn priority" in fired ids: target = "retention billing queue" elif "confident department" in fired ids: target = f"{dept} queue" else: target = "ABSTAIN" nothing cleared its threshold: fall through to the next tier return { "ts": datetime.now timezone.utc .isoformat timespec="seconds" , What the router saw: a fingerprint of the redacted state store the state itself separately "state sha256": hashlib.sha256 json.dumps redacted state, sort keys=True .encode .hexdigest :16 , "tier": "TIER 3B LAYA", "model": model, "policy version": POLICY VERSION, "target": target, "reason": "; ".join r for , r in fired or "no rule cleared its threshold", "answers": answers, the full distributions, not just the winner "rules fired": sorted fired ids , "page on call": "page on call" in fired ids, } Illustrative answers for the refund-or-cancel message above your model will return its own numbers redacted state = {"message": "We were billed twice for invoice