{"slug": "stop-making-your-big-model-answer-multiple-choice", "title": "Stop making your big model answer multiple choice", "summary": "A developer has released Jeff, a 0.8B open model that answers multiple-choice decision calls — prompt-injection checks, ticket triage, tool choice, grounding checks and spam — by returning a probability per labeled option from a single forward pass instead of generating text. In the project's own benchmarks on an M4 Max with MLX, routing to Qwen3-27B only when Jeff is unsure lifted accuracy from 86.6% to 95.3% while cutting decision time from 8.1s to 0.25s and adding 1.96 GB of memory, with per-adapter confidence thresholds calibrated on held-out rows before test scoring. The project, discussed on Hacker News, ships nine ~41 MB LoRA adapters and a FastAPI server at /v1/systemone whose usage block always reports output_tokens: 0.", "body_md": "A support email lands in an agent's inbox.\n\nBefore the agent writes a single word back, it makes five small calls. Is this a prompt injection? How urgent is it? What does the customer actually want? Which tool runs first? Is the draft backed by the docs it pulled?\n\nEvery one of those is a multiple-choice question. In most stacks I have seen, every one goes to the biggest model on hand, which writes a paragraph that a regex then picks apart.\n\n[Jeff](https://github.com/firelex/jeff) is a 0.8B open model built for exactly those calls. You give it a situation and a list of labeled options. It gives back a probability for each option from one forward pass. It writes no text, so there is nothing to parse.\n\nThe author calls it a \"System 1\" model: the fast, reflexive half of a stack, with a big model as the slow, deliberate half. It was built at home as an open take on a hosted decision API, and [the HN thread](https://news.ycombinator.com/item?id=49883844) hit 574 points this week.\n\nThe base model handles any options you give it, zero-shot. On top sit nine adapters, small LoRA add-ons of about 41 MB each, one per job: a prompt-injection guard, ticket triage, tool choice, grounding checks, spam, and a few more. Per the README, each one trained in one epoch on one GPU, in half an hour to four hours.\n\nThe headline setup is the one I care about. Jeff answers first. Only when it is unsure does the question go on to Qwen3.8-27B. Across eight adapters, against the 27B deciding everything alone:\n\n|  | 27B decides alone | Jeff first, 27B when unsure | \n|---|---|---|\n| Accuracy | 86.6% | 95.3% | \n| Time per decision | 8.1 s | 0.25 s | \n| Extra memory | 28.6 GB | +1.96 GB | \n\nThose are the project's numbers, from an M4 Max with both models on MLX and the 27B's step-by-step reasoning turned off. The detail that earned my trust is how the \"unsure\" line was set. Each adapter's confidence threshold was picked on a separate set of calibration rows, and fixed before the test rows were scored. Plenty of model READMEs skip that step.\n\nThey also show the task where the small model does not win. On grounding (is this answer supported by its sources?), the 27B alone scores 96.7% and Jeff ends at 96.3%. That is one question in 300, at 20 times the speed. A results table with a near-loss in it reads like a measurement, not a pitch.\n\nThe server is a small FastAPI app. The route is `/v1/systemone`. Inside, the model scores every option and a softmax turns the scores into probabilities. The usage block it returns always says `output_tokens: 0`.\n\nOne flag I liked: `orders: 2`. It asks the same question a second time with the options reversed, then averages the two. Models lean toward options near the top of a list, and this cancels some of that out at the cost of a second pass.\n\nThe Python client, adapted from the README's own example (I did not run this):\n\n``` python\nfrom jeff import Client\nfrom jeff.client import choice_question\n\ntools = Client(\"http://localhost:8765\", model=\"tools\")  # the tool-choice adapter\nanswers = tools.ask(\"User: main is red again, can you see why?\", {\n    \"tool\": choice_question({\n        \"ci_logs\": \"Read the latest CI run logs\",\n        \"git_log\": \"List the recent commits\",\n        \"web_search\": \"Search the web\",\n    }, \"Which tool should the agent call first?\"),\n})\npick = answers.choice(\"tool\")\nif pick.confidence < THRESHOLD:\n    ...  # hand this one to the big model\n```\n\nEvery answer carries `key`, `probability` and `confidence`, where confidence runs from 0 (no better than a guess) to 1 (certain). That last field is the whole trick. The server also serves a playground page at its root URL, for trying questions by hand.\n\nHere is the claim I would build on, with or without Jeff. When your agent asks a big model \"which tool?\", you pay for a prompt, wait seconds, get prose back, and parse it. You also get no honest signal of how sure it was. A classifier hands you a calibrated probability, and the probability is the feature. It tells you exactly when to escalate.\n\nThat flips the usual design. The big model stops being the default for every small call and becomes the fallback for the hard ones. Your agent loop gets faster, your bill shrinks, and you gain something most stacks lack: a number that says \"I am not sure, ask someone smarter.\"\n\nThe guard adapter is where I would start. A prompt-injection check in front of every tool output, at about a tenth of a second per check in their table, is cheap enough to run on everything.\n\nThe edges are real, and the issue tracker is honest about them.\n\nThe maintainer does move fast. [Issue #1](https://github.com/firelex/jeff/issues/1) found that no option past the 26th was ever chosen. That was fixed and released as v1.1 within days, with 32,000 extra long-list questions in the training data.\n\nI read the README, the server and client code, three issues, and the HN thread. I did not run it. The install is a full Python ML environment plus a 1.7 GB model download, and that is more than I install for a post. Every number above is the project's own or comes from a named issue.\n\nIf you try it, start the server, open the playground, and throw your own five inbox questions at it.\n\nWhich decision in your agent would you hand to a 0.8B model first, and which one would you never let it make?", "url": "https://wpnews.pro/news/stop-making-your-big-model-answer-multiple-choice", "canonical_source": "https://dev.to/ianwieds/stop-making-your-big-model-answer-multiple-choice-4n9n", "published_at": "2026-10-03 18:00:00+00:00", "updated_at": "2026-10-03 18:07:47.754414+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "machine-learning", "ai-infrastructure"], "entities": ["Jeff", "Qwen3-27B", "MLX", "FastAPI", "Hacker News", "M4 Max"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/stop-making-your-big-model-answer-multiple-choice", "markdown": "https://wpnews.pro/news/stop-making-your-big-model-answer-multiple-choice.md", "text": "https://wpnews.pro/news/stop-making-your-big-model-answer-multiple-choice.txt", "jsonld": "https://wpnews.pro/news/stop-making-your-big-model-answer-multiple-choice.jsonld"}}