How I get a confidence score for every option without generating text, and why that score should only ever make a gate stricter.
A voice bot does not need a paragraph about what the caller wants. It needs a label, and a number that says how sure the model is. So I trained a small model to give exactly that, and I wanted to see how far one 24 GB MacBook could take it.
TypeSafe AI has described this style publicly as "decision models". This is my own version, built in the open-source forge fine-tuning tool. It is not their model, and I make no comparison to it.
The setup
The model is Gemma 4 E2B (4-bit) with a LoRA adapter, rank 16. The adapter is 52 MB. Training peaked at about 5.7 GB of memory and took about an hour. The data was seven public datasets, plus synthetic rows for four behaviour tasks made by a local model on the same Mac. No paid API and no cloud GPU.
Every example has the same shape: the conversation, a question, and lettered options.
Customer: I need to move my appointment, it's urgent.
Question: Which request is the customer making?
Options:
A) reschedule an appointment
B) cancel an appointment
C) ask about pricing
The model is trained to answer with one letter. At test time I do not let it write anything. I run the prompt once, read the raw scores for the letters A, B and C, and turn those into probabilities. A temperature fitted on a validation set makes them well calibrated. There is nothing to parse, and the answer can never be an invalid label.
Three details mattered:
What good calibration gives you
On about 68,000 test rows, the calibration error was between 0.006 and 0.04 on most tasks. Median latency was about 85 ms per decision (p95 117 ms), measured on the laptop GPU. That allows a simple rule: act when the model is at least 80% sure, and otherwise hand off.
| Test | Handled at 0.8 or higher | Accuracy on those |
|---|---|---|
| CLINC150 (held out) | 90% | 98.3% |
| MASSIVE intent | 79% | 95.4% |
| BANKING77 (held out) | 74% | 92.3% |
The part that is not about the model
A confident 0.92 says the label is probably right. It does not say the agent is allowed to act. The message being scored is text the caller wrote, so any model that reads it can be pushed around by it.
This is how I would combine the two in an agent gateway:
check identity and scoped permission first
if denied: stop
score = decision_model(message)
if score is below the threshold: send to a human
else: go ahead
The score can turn a "go ahead" into a "send to a human". It can never turn a "denied" into a "go ahead". Identity and scoped permission decide. The model only narrows. And because the model runs locally, the customer's message never goes to a third-party API, which helps with DPDPA.
What I am not claiming
The "held-out" intent tests pick from about ten options with random wrong answers. That is easier than the published 77-way and 151-way benchmarks.
The behaviour-task tests are small, and the test rows were made by the same kind of model that made the training rows.
I have not tested on real recorded calls. Everything above is on public data.
Where I am stuck
Training loss stays flat near ln(26), which is a uniform guess over the answer letters, for roughly the first 2,000 steps. Then it drops sharply. A lower learning rate moved the plateau but did not remove it. I do not know why.
Open question:
where should the hand-off threshold live? Per task, per data type, or per permission? I lean towards per data type. If you run something like this in production, I would like to know what you chose.