{"slug": "tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that", "title": "Tiny model. Big decisions. — How I built a 144M-parameter typed decision model that routes 82% of agent decisions off LLMs", "summary": "A developer built Phocinae-Largha-150M-v1, a 144.3M-parameter typed decision model that answers yes/no, pick-one and 1-10 questions in a single forward pass at 18.6 ms p50 on GPU, scoring 0.797 on English typed-decisions (400 cases / 2,000 decisions) versus published same-protocol scores of 0.766 for Laya, 0.768 for meraGPT and 0.727 for JEV-27B. With a τ=0.6 confidence gate that escalates uncertain cases to a larger LLM, the Chinese route cut LLM calls by 82% (100% to 18%) while combined accuracy rose from 0.789 to 0.7948. The Apache-2.0 model is served locally over a /v1/systemone protocol with calibrated probabilities, alongside an npm DeepSeek Harness bundle with a fail-closed PreToolUse approval gate.", "body_md": "*Or: why your agent's \"should I run this command?\" does not need a 70B chat model.*\n\nEvery agentic workflow I build hits the same wall: the interesting logic is 10 lines, and the other 90% is a language model deciding *\"allow or deny?\", \"which tool?\", \"pass or escalate?\"* — thousands of times a day, at chat-model prices and chat-model latency.\n\nSo I built the opposite of a chatbot: **Phocinae-Largha-150M-v1**, a 144.3M-parameter typed decision model. It cannot generate text. It takes a state plus a list of typed questions (yes/no, pick-one, 1-10 score) and returns, in one forward pass, a verdict per question with calibrated confidence. GPU: **18.6 ms** p50 per decision. CPU-only: ~1.5 s, no GPU at all. Open-source, Apache-2.0.\n\nA decision is a **typed output over a closed option set**:\n\n```\nstate: \"agent wants to run: rm -rf /var/log/app\"\nquestion: { type: noul, qid: allow, options: [false, true] }\nanswer:  { allow: { label: false, prob: 0.96, confidence: 0.96 } }\n```\n\nThere is no sentence to generate, no chain-of-thought to emit, no format to parse back. So why rent a generative model for it? A 150M encoder (mmBERT-small base, 256k vocab, Gemma tokenizer) does one forward pass and outputs label logits per question — that is the entire inference. Deterministic: same input, same output. No sampling, no parsing failures.\n\nThe evaluation protocol is typed-decisions (the format used by Laya and others), which makes scores directly comparable across models on the same rows.\n\nEnglish typed-decisions: **0.797** (400 cases / 2,000 decisions). Chinese (machine-translated eval set, no native zh training rows — disclosed): **0.789**. Same-protocol published scores: Laya 0.766 · JEV-27B 0.727 · meraGPT 0.768.\n\nThe metrics most model cards skip, we publish:\n\nA small model does not need to be perfect — it needs to know when it is not. With a τ=0.6 confidence gate, confident decisions stay local and the rest escalate to a bigger model. Result on the zh route: LLM calls **cut 82%** (100% → 18%), while combined accuracy went **0.789 → 0.7948** — routing the hard 18% upward made the whole system slightly *better*, not just cheaper.\n\nThat is the framing I want to leave with you: **System 1 in BERT**. The two-system picture for agents is not \"small model vs big model\" — it is *typed, deterministic, milliseconds, free* for the repetitive 82%, and *generative, expensive* only for the ambiguous 18%.\n\n```\npip install phocinae-server huggingface_hub\nhuggingface-cli download Phocinae/Phocinae-Largha-150M-v1 --local-dir ./model\nPHOC_MODEL_DIR=./model python -m phocinae.main\ncurl -s http://127.0.0.1:8155/v1/systemone \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"state\":\"The agent wants to run: rm -rf /var/log/app\",\n       \"questions\":[{\"type\":\"noul\",\"qid\":\"allow\",\n                     \"question\":\"Allow this command?\",\"options\":[\"false\",\"true\"]}]}'\n# → {\"allow\": {\"label\": \"false\", \"prob\": 0.96, \"confidence\": 0.96}}\n```\n\nThe server is local-only (127.0.0.1), pure PyTorch at runtime, and the protocol (`/v1/systemone`: noul / choice / score questions, calibrated probabilities) is fully specified in the repo. There is also a DeepSeek Harness bundle (dsh-phocinae, npm) with a PreToolUse approval gate that fails closed.\n\nIf you have a use case that is mostly repetitive decisions, I'd love to hear what breaks first.", "url": "https://wpnews.pro/news/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that", "canonical_source": "https://dev.to/perrylink/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that-routes-82-of-23fi", "published_at": "2026-10-08 03:41:32+00:00", "updated_at": "2026-10-08 03:46:54.434596+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "machine-learning", "ai-infrastructure"], "entities": ["Phocinae-Largha-150M-v1", "Phocinae", "Laya", "JEV-27B", "meraGPT", "mmBERT-small", "DeepSeek Harness", "PyTorch"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that", "markdown": "https://wpnews.pro/news/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that.md", "text": "https://wpnews.pro/news/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that.txt", "jsonld": "https://wpnews.pro/news/tiny-model-big-decisions-how-i-built-a-144m-parameter-typed-decision-model-that.jsonld"}}