Small Decisions: Engineering a Leading Model Marc Brooker built Hobson, a roughly 2-billion-parameter calibrated classifier model, by replacing the LM head of a pre-trained Qwen3.5-2B torso with a pointer head of just over 1 million parameters and fine-tuning with a rank-16 LoRA adapter. Hobson reached a Brier score of 0.009 and 100% accuracy on the JevBench easy set, beating Qwen3.5-2B on both calibration and accuracy in its size range, though a newer decider-2b version beats it and has not yet been added to the leaderboard. The work followed TypeSafe AI's announcement of Jev on the 15th, a general-purpose calibrated classifier the author cites as a building block for workflow-oriented AI agents. My day job has been primarily in AI for three years now, but I’d be the first to admit that’s been almost entirely in one corner of AI: infrastructure, safety, and tools for AI agents. That work has brought me in contact with a lot of the AI science and I’d dabbled there over the previous decade , but I’m super far from the day-to-day of work like model building. I wanted to catch up a little after all you have to know what you’re talking about https://brooker.co.za/blog/2026/03/20/ic-leadership.html , and the last couple weeks provided a perfect opportunity. On the 15th of this month, the TypeSafe AI folks announced Jev https://typesafe.ai/blog/introducing-system-one-models-and-jev , a kind of general purpose calibrated classifier. You can see this as something exciting or not, but it sure has captured the world’s attention. And mine. I was particularly interested in the calibration, combined with low latency and the ability to answer questions in parallel, it’s a great building block for the more workflowy end of the spectrum of agents. In an effort to understand these things well, it was time to build my own model: Hobson. A small one, because I wanted to use the GPU I have at home, and because I wanted to see if I could push the bounds on accuracy and calibration at very low latency. I decided to limit myself to about 2 billion parameters. How have I done so far? Fairly well, I think. You can read that as a trajectory of how versions of my model have performed on accuracy on x and calibration on y as I’ve made improvements. The pareto optimal is the bottom right. On the jevbench public set https://benchmarkheaven.com/jev-models I’m at the top in my size range