cd /news/large-language-models/dev-4b-quick-calibrated-decisions-on… · home › topics › large-language-models › article
[ARTICLE · art-145571] src=discuss.huggingface.co ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Dev-4B: quick calibrated decisions on a document, reason only when the router flags risk

Suhaas Teja released Dev-4B, a 4B-parameter local model built on Qwen3-4B-Instruct-2507 plus roughly 133 MB of add-ons that answers typed document questions (choice, yes-no, score) with calibrated confidence and escalates to chain-of-thought only when a built-in router predicts the quick answer is wrong. The model ships with a Hugging Face Space demo, a Hugging Face model card, and an MLX 8-bit build, and is released under non-commercial CC BY-NC-SA weights with no Ollama or LM Studio support because of its LoRA adapter, decision head and router. Teja states the model is weaker on unseen task types than trained ones, is English-only, and was trained on documents of about 4,000 characters, and is seeking feedback on Mac performance and router escalation behavior.

read1 min views2 publishedOct 5, 2026

I released Dev-4B — a small local model that answers typed questions about a document (choice / yes-no / score) with calibrated confidence, and only thinks step by step when a built-in router predicts the quick answer is likely wrong.

Space: [Dev-4B - a Hugging Face Space by suhaas-teja](https://huggingface.co/spaces/suhaas-teja/Dev-4B-demo)

Model: [suhaas-teja/Dev-4B · Hugging Face](https://huggingface.co/suhaas-teja/Dev-4B)

MLX 8-bit build is on the same profile as Dev-4B-MLX-8bit (search that name on Hugging Face).

Stack: Qwen3-4B-Instruct-2507 + ~133 MB of add-ons (LoRA adapter switched off while reading the document and on for the question, decision head, temperatures for calibration, tiny router). Base generation is unchanged when add-ons are off.

Results on 7,100 frozen test questions:

Limits to be clear about: weaker on unseen task types than on trained ones; reasoning is still a 4B model; English; docs ~4k chars in training; NC weights (CC BY-NC-SA); no Ollama/LM Studio (LoRA + decision head + router). Question format follows System One; not affiliated with TypeSafe.

Most open System One clones are decision-only. Dev-4B’s wedge is local calibrated decisions plus a router that escalates to CoT when the quick pass looks wrong — for Mac indie / offline routing experimenters.

Would love feedback from people running local decision models — especially Mac numbers and cases where the router should/shouldn’t escalate.

── more in #large-language-models 4 stories · sorted by recency
── more on @dev-4b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dev-4b-quick-calibra…] indexed:0 read:1min 2026-10-05 · —