Tiny model. Big decisions. — How I built a 144M-parameter typed decision model that routes 82% of agent decisions off LLMs A developer built Phocinae-Largha-150M-v1, a 144.3M-parameter typed decision model that answers yes/no, pick-one and 1-10 questions in a single forward pass at 18.6 ms p50 on GPU, scoring 0.797 on English typed-decisions (400 cases / 2,000 decisions) versus published same-protocol scores of 0.766 for Laya, 0.768 for meraGPT and 0.727 for JEV-27B. With a τ=0.6 confidence gate that escalates uncertain cases to a larger LLM, the Chinese route cut LLM calls by 82% (100% to 18%) while combined accuracy rose from 0.789 to 0.7948. The Apache-2.0 model is served locally over a /v1/systemone protocol with calibrated probabilities, alongside an npm DeepSeek Harness bundle with a fail-closed PreToolUse approval gate. Or: why your agent's "should I run this command?" does not need a 70B chat model. Every agentic workflow I build hits the same wall: the interesting logic is 10 lines, and the other 90% is a language model deciding "allow or deny?", "which tool?", "pass or escalate?" — thousands of times a day, at chat-model prices and chat-model latency. So I built the opposite of a chatbot: Phocinae-Largha-150M-v1 , a 144.3M-parameter typed decision model. It cannot generate text. It takes a state plus a list of typed questions yes/no, pick-one, 1-10 score and returns, in one forward pass, a verdict per question with calibrated confidence. GPU: 18.6 ms p50 per decision. CPU-only: ~1.5 s, no GPU at all. Open-source, Apache-2.0. A decision is a typed output over a closed option set : state: "agent wants to run: rm -rf /var/log/app" question: { type: noul, qid: allow, options: false, true } answer: { allow: { label: false, prob: 0.96, confidence: 0.96 } } There is no sentence to generate, no chain-of-thought to emit, no format to parse back. So why rent a generative model for it? A 150M encoder mmBERT-small base, 256k vocab, Gemma tokenizer does one forward pass and outputs label logits per question — that is the entire inference. Deterministic: same input, same output. No sampling, no parsing failures. The evaluation protocol is typed-decisions the format used by Laya and others , which makes scores directly comparable across models on the same rows. English typed-decisions: 0.797 400 cases / 2,000 decisions . Chinese machine-translated eval set, no native zh training rows — disclosed : 0.789 . Same-protocol published scores: Laya 0.766 · JEV-27B 0.727 · meraGPT 0.768. The metrics most model cards skip, we publish: A small model does not need to be perfect — it needs to know when it is not. With a τ=0.6 confidence gate, confident decisions stay local and the rest escalate to a bigger model. Result on the zh route: LLM calls cut 82% 100% → 18% , while combined accuracy went 0.789 → 0.7948 — routing the hard 18% upward made the whole system slightly better , not just cheaper. That is the framing I want to leave with you: System 1 in BERT . The two-system picture for agents is not "small model vs big model" — it is typed, deterministic, milliseconds, free for the repetitive 82%, and generative, expensive only for the ambiguous 18%. pip install phocinae-server huggingface hub huggingface-cli download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./model PHOC MODEL DIR=./model python -m phocinae.main curl -s http://127.0.0.1:8155/v1/systemone \ -H 'Content-Type: application/json' \ -d '{"state":"The agent wants to run: rm -rf /var/log/app", "questions": {"type":"noul","qid":"allow", "question":"Allow this command?","options": "false","true" } }' → {"allow": {"label": "false", "prob": 0.96, "confidence": 0.96}} The server is local-only 127.0.0.1 , pure PyTorch at runtime, and the protocol /v1/systemone : noul / choice / score questions, calibrated probabilities is fully specified in the repo. There is also a DeepSeek Harness bundle dsh-phocinae, npm with a PreToolUse approval gate that fails closed. If you have a use case that is mostly repetitive decisions, I'd love to hear what breaks first.