AI on Edge: One forward pass, one letter: a decision model that runs on a laptop A developer built a small local decision model on a 24 GB MacBook that returns a single lettered label plus a calibrated confidence score in one forward pass, using Gemma 4 E2B (4-bit) with a 52 MB rank-16 LoRA adapter trained in about an hour on seven public datasets plus synthetic rows. On roughly 68,000 test rows the calibration error was 0.006–0.04 on most tasks, with median latency around 85 ms (p95 117 ms), and a gate at 0.8 confidence handled 90% of held-out CLINC150 rows at 98.3% accuracy. The developer argues the score should only ever make a gate stricter — identity and scoped permission decide, the model only narrows — and asks where the hand-off threshold should live in production. How I get a confidence score for every option without generating text, and why that score should only ever make a gate stricter. A voice bot does not need a paragraph about what the caller wants. It needs a label, and a number that says how sure the model is. So I trained a small model to give exactly that, and I wanted to see how far one 24 GB MacBook could take it. TypeSafe AI has described this style publicly as "decision models". This is my own version, built in the open-source forge fine-tuning tool. It is not their model, and I make no comparison to it. The setup The model is Gemma 4 E2B 4-bit with a LoRA adapter, rank 16. The adapter is 52 MB. Training peaked at about 5.7 GB of memory and took about an hour. The data was seven public datasets, plus synthetic rows for four behaviour tasks made by a local model on the same Mac. No paid API and no cloud GPU. Every example has the same shape: the conversation, a question, and lettered options. Customer: I need to move my appointment, it's urgent. Question: Which request is the customer making? Options: A reschedule an appointment B cancel an appointment C ask about pricing The model is trained to answer with one letter. At test time I do not let it write anything. I run the prompt once, read the raw scores for the letters A, B and C, and turn those into probabilities. A temperature fitted on a validation set makes them well calibrated. There is nothing to parse, and the answer can never be an invalid label. Three details mattered: What good calibration gives you On about 68,000 test rows, the calibration error was between 0.006 and 0.04 on most tasks. Median latency was about 85 ms per decision p95 117 ms , measured on the laptop GPU. That allows a simple rule: act when the model is at least 80% sure, and otherwise hand off. | Test | Handled at 0.8 or higher | Accuracy on those | |---|---|---| | CLINC150 held out | 90% | 98.3% | | MASSIVE intent | 79% | 95.4% | | BANKING77 held out | 74% | 92.3% | The part that is not about the model A confident 0.92 says the label is probably right. It does not say the agent is allowed to act. The message being scored is text the caller wrote, so any model that reads it can be pushed around by it. This is how I would combine the two in an agent gateway: check identity and scoped permission first if denied: stop score = decision model message if score is below the threshold: send to a human else: go ahead The score can turn a "go ahead" into a "send to a human". It can never turn a "denied" into a "go ahead". Identity and scoped permission decide. The model only narrows. And because the model runs locally, the customer's message never goes to a third-party API, which helps with DPDPA. What I am not claiming The "held-out" intent tests pick from about ten options with random wrong answers. That is easier than the published 77-way and 151-way benchmarks. The behaviour-task tests are small, and the test rows were made by the same kind of model that made the training rows. I have not tested on real recorded calls. Everything above is on public data. Where I am stuck Training loss stays flat near ln 26 , which is a uniform guess over the answer letters, for roughly the first 2,000 steps. Then it drops sharply. A lower learning rate moved the plateau but did not remove it. I do not know why. Open question: where should the hand-off threshold live? Per task, per data type, or per permission? I lean towards per data type. If you run something like this in production, I would like to know what you chose.