{"slug": "dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk", "title": "Dev-4B: quick calibrated decisions on a document, reason only when the router flags risk", "summary": "Suhaas Teja released Dev-4B, a 4B-parameter local model built on Qwen3-4B-Instruct-2507 plus roughly 133 MB of add-ons that answers typed document questions (choice, yes-no, score) with calibrated confidence and escalates to chain-of-thought only when a built-in router predicts the quick answer is wrong. The model ships with a Hugging Face Space demo, a Hugging Face model card, and an MLX 8-bit build, and is released under non-commercial CC BY-NC-SA weights with no Ollama or LM Studio support because of its LoRA adapter, decision head and router. Teja states the model is weaker on unseen task types than trained ones, is English-only, and was trained on documents of about 4,000 characters, and is seeking feedback on Mac performance and router escalation behavior.", "body_md": "I released Dev-4B — a small local model that answers typed questions about a document (choice / yes-no / score) with calibrated confidence, and only thinks step by step when a built-in router predicts the quick answer is likely wrong.\n\nSpace: [Dev-4B - a Hugging Face Space by suhaas-teja](https://huggingface.co/spaces/suhaas-teja/Dev-4B-demo)\n\nModel: [suhaas-teja/Dev-4B · Hugging Face](https://huggingface.co/suhaas-teja/Dev-4B)\n\nMLX 8-bit build is on the same profile as Dev-4B-MLX-8bit (search that name on Hugging Face).\n\nStack: Qwen3-4B-Instruct-2507 + ~133 MB of add-ons (LoRA adapter switched off while reading the document and on for the question, decision head, temperatures for calibration, tiny router). Base generation is unchanged when add-ons are off.\n\nResults on 7,100 frozen test questions:\n\nLimits to be clear about: weaker on unseen task types than on trained ones; reasoning is still a 4B model; English; docs ~4k chars in training; NC weights (CC BY-NC-SA); no Ollama/LM Studio (LoRA + decision head + router). Question format follows System One; not affiliated with TypeSafe.\n\nMost open System One clones are decision-only. Dev-4B’s wedge is local calibrated decisions plus a router that escalates to CoT when the quick pass looks wrong — for Mac indie / offline routing experimenters.\n\nWould love feedback from people running local decision models — especially Mac numbers and cases where the router should/shouldn’t escalate.", "url": "https://wpnews.pro/news/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk", "canonical_source": "https://discuss.huggingface.co/t/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-flags-risk/182924#post_1", "published_at": "2026-10-05 18:10:29+00:00", "updated_at": "2026-10-05 18:17:24.032455+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-tools", "ai-agents"], "entities": ["Dev-4B", "Suhaas Teja", "Qwen3-4B-Instruct-2507", "Hugging Face", "Dev-4B-MLX-8bit", "System One", "TypeSafe"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk", "markdown": "https://wpnews.pro/news/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk.md", "text": "https://wpnews.pro/news/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk.txt", "jsonld": "https://wpnews.pro/news/dev-4b-quick-calibrated-decisions-on-a-document-reason-only-when-the-router-risk.jsonld"}}