I released Dev-4B — a small local model that answers typed questions about a document (choice / yes-no / score) with calibrated confidence, and only thinks step by step when a built-in router predicts the quick answer is likely wrong.
Space: [Dev-4B - a Hugging Face Space by suhaas-teja](https://huggingface.co/spaces/suhaas-teja/Dev-4B-demo)
Model: [suhaas-teja/Dev-4B · Hugging Face](https://huggingface.co/suhaas-teja/Dev-4B)
MLX 8-bit build is on the same profile as Dev-4B-MLX-8bit (search that name on Hugging Face).
Stack: Qwen3-4B-Instruct-2507 + ~133 MB of add-ons (LoRA adapter switched off while reading the document and on for the question, decision head, temperatures for calibration, tiny router). Base generation is unchanged when add-ons are off.
Results on 7,100 frozen test questions:
Limits to be clear about: weaker on unseen task types than on trained ones; reasoning is still a 4B model; English; docs ~4k chars in training; NC weights (CC BY-NC-SA); no Ollama/LM Studio (LoRA + decision head + router). Question format follows System One; not affiliated with TypeSafe.
Most open System One clones are decision-only. Dev-4B’s wedge is local calibrated decisions plus a router that escalates to CoT when the quick pass looks wrong — for Mac indie / offline routing experimenters.
Would love feedback from people running local decision models — especially Mac numbers and cases where the router should/shouldn’t escalate.