Most AI reliability writing assumes a network. You call a bigger model to double-check, you log to a server, you push a hotfix when something goes wrong. A disaster app breaks every one of those assumptions. By the time someone opens it, the cell tower may be rubble and the phone may be in airplane mode for the next two days. There is no bigger model to call, no server to log to, and no hotfix that will arrive in time.
This post is about the engineering we did for HeliGO, an offline disaster and distress app built around an on-device 4B model. The runtime side (thread pinning and lazy on a phone) was covered earlier in this series. Here I want to focus on the harder problem: how do you make a small model's answers trustworthy when a wrong answer can get someone killed and nobody is online to catch it?
The obvious approach is a system prompt: "You are a disaster assistant. Never give dangerous advice. Say you are unsure when you are unsure." On a 4B model with no fallback, this is wishful thinking for two reasons.
First, a prompt is a suggestion, not a constraint. A small model under distribution shift (a panicked, ungrammatical question about a mushroom in the dark) will still occasionally produce a confident, wrong, fatal answer. There is no second model downstream to catch it.
Second, and this is the part that surprised us most, the model's own stated confidence is not just useless, it is inverted. When we asked the model to attach a confidence number to its answers, wrong answers tended to come with higher stated confidence than right ones. Asking a model "are you sure?" and trusting the reply is worse than a coin flip in exactly the cases that matter.
So the reliability layer cannot live inside the model's text output. It has to wrap the model.
Every answer the model produces passes through three stages before it reaches the screen. Think of it as a decision pipeline, not a filter.
model draft
|
v
[1] fatal-output block -- is this a category that can kill if wrong?
| (if yes and risky) -> suppress, escalate to stage 2
v
[2] official-text substitute -- is there a matching MOIS guideline?
| (if yes) -> answer with the official text, not the model's words
v
[3] refuse -- still not confident enough?
| (if yes) -> say "I am not sure", do not guess
v
answer shown to user
Stage 1, fatal-output block. Some answer categories are life-or-death: "yes, you can eat this", "yes, this water is safe", "yes, go that way". For these, the model's confidence is irrelevant. If the content falls in a fatal category and carries any risk signal, it is suppressed regardless of how sure the model sounds. This is the direct consequence of the inverted-confidence finding: in the categories where being wrong is fatal, we do not let the model's certainty vote at all.
Stage 2, official-text substitution. We packaged 75 official response guidelines from Korea's Ministry of the Interior and Safety (MOIS) on the device. When a question maps to one of them (earthquake, flood, wildfire, first aid, and so on), the app answers with the official text verbatim instead of the model's paraphrase. The model is still useful here: it does the matching and retrieval. But the words the user reads are the authoritative source, not a 4B paraphrase that might drift.
Stage 3, refuse. If an answer is neither safely substitutable nor confidently verifiable, the app declines. "I am not sure" is a valid and often correct output for a disaster tool. A refusal costs the user a little time. A confident wrong answer can cost a life. The gate is built so refusal is the default failure mode, never a guess.
The gate is deliberately not symmetric, and this is the design decision I am most sure about.
The app never says "this is safe." It always says "this is poisonous" when it has reason to.
The reasoning is a straightforward expected-cost argument. Consider the two ways the app can be wrong about something edible:
| Situation | App says | If app is wrong | Cost |
|---|---|---|---|
| Edible plant | "I cannot confirm this is safe" | User stays hungry | Low, recoverable |
| Poisonous plant | "This may be poisonous" | User avoids safe food | Low, recoverable |
| Poisonous plant | "This is safe to eat" | User is poisoned | Catastrophic |
Three of these outcomes are survivable. Only one is not. So the app is tuned to make the unsurvivable mistake nearly impossible, at the cost of being overly cautious in the harmless direction. The costs of silence are not equal on the two sides, so the gate is not equal either. We ship warnings for 12 venomous and poisonous species in the same spirit: a false alarm is cheap, a missed warning is not. Any edge safety system where errors have asymmetric cost should bias toward the cheap error.
Stages 1 and 3 both need a trustworthy confidence signal, and we already established that the model's spoken confidence is inverted. The fix is to stop asking.
TELL reads the model's internal state while it produces an answer and derives a reliability estimate from that, rather than from any number the model writes in its text. The intuition: the computation a transformer performs on its way to an answer carries signal about whether that answer is well-grounded, and that signal is more honest than the self-report the model emits at the end. Internal state does not have an incentive to sound confident; the output text, apparently, does.
I am not going to publish the thresholds or the exact internal signals we read, for the obvious reason that they are what makes the gate hard to game. The point for this article is architectural: the confidence check lives outside the token stream. We treat the model as an instrument to be read, not a witness to be interviewed.
The TELL technique was released alongside the model. In HeliGO it currently backs the gate's suppress-or-refuse decisions, and a future update will surface low-confidence filtering more directly to the user.
For an offline tool, "on-device" has to be literal. Everything the gate and the map need is packaged into the app so it works with the radio off:
None of this phones home. Location history in particular never leaves the device, which is both a privacy property and a reliability property: there is no network dependency to fail.
You move trust out of the model and into a deterministic wrapper. The model drafts and retrieves; a gate you control decides what reaches the user. Fatal categories are blocked on content, not on the model's confidence. Authoritative answers come from packaged official text. When neither applies, you refuse. The model never gets the final word on anything that can kill.
Because the phone is the budget. A 4B model is what runs in airplane mode on a consumer chip with room left for maps and rendering. The interesting claim of this project is not that a big model is safe, it is that a small model plus the right wrapper can be safe enough to ship for life-or-death use. The reliability came from architecture, not from parameter count.
Does HeliGO send anything to a server?
No. Maps, contours, altitude, compass, rescue coordinates, and the AI chat all run on-device in airplane mode. Location history does not leave the phone.
What is the model?
Edge-4B-TELL, built on Google's Gemma-4 E4B, published on Hugging Face. It handles offline chat, photo identification, and voice input on the phone's own chip.
Why does it never say something is safe to eat?
Because the cost of being wrong is asymmetric. A missed safe meal is hunger; a wrong "safe" is poisoning. The gate biases toward the recoverable error and only ever issues the poison warning side confidently.
Can I trust the AI's confidence when it answers?
The app does not trust the model's stated confidence, and neither should you in general. Measured on this model, wrong answers tended to be stated more confidently than right ones. HeliGO derives confidence from the model's internal state (TELL) instead, outside the text it writes.
Is this a replacement for calling emergency services?
No. HeliGO does not replace rescue services. If you have any signal, call emergency services first. It is a tool for when you have no connection.
Further reading (same account):