{"slug": "building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline", "title": "Building an AI That Refuses to Say \"Safe\": Reliability Engineering for an Offline Disaster App", "summary": "A developer built HeliGO, an offline disaster and distress app that runs a 4B on-device model behind a three-stage reliability gate rather than trusting the model's own output. The pipeline blocks life-or-death answer categories regardless of stated confidence, substitutes verbatim text from 75 official Korean Ministry of the Interior and Safety (MOIS) guidelines when a question matches, and otherwise refuses to answer. The design followed the finding that the small model's self-reported confidence was inverted — wrong answers tended to carry higher stated confidence than correct ones.", "body_md": "Most AI reliability writing assumes a network. You call a bigger model to double-check, you log to a server, you push a hotfix when something goes wrong. A disaster app breaks every one of those assumptions. By the time someone opens it, the cell tower may be rubble and the phone may be in airplane mode for the next two days. There is no bigger model to call, no server to log to, and no hotfix that will arrive in time.\n\nThis post is about the engineering we did for HeliGO, an offline disaster and distress app built around an on-device 4B model. The runtime side (thread pinning and lazy loading on a phone) was covered earlier in this series. Here I want to focus on the harder problem: how do you make a small model's answers trustworthy when a wrong answer can get someone killed and nobody is online to catch it?\n\nThe obvious approach is a system prompt: \"You are a disaster assistant. Never give dangerous advice. Say you are unsure when you are unsure.\" On a 4B model with no fallback, this is wishful thinking for two reasons.\n\nFirst, a prompt is a suggestion, not a constraint. A small model under distribution shift (a panicked, ungrammatical question about a mushroom in the dark) will still occasionally produce a confident, wrong, fatal answer. There is no second model downstream to catch it.\n\nSecond, and this is the part that surprised us most, the model's own stated confidence is not just useless, it is inverted. When we asked the model to attach a confidence number to its answers, wrong answers tended to come with *higher* stated confidence than right ones. Asking a model \"are you sure?\" and trusting the reply is worse than a coin flip in exactly the cases that matter.\n\nSo the reliability layer cannot live inside the model's text output. It has to wrap the model.\n\nEvery answer the model produces passes through three stages before it reaches the screen. Think of it as a decision pipeline, not a filter.\n\n```\nmodel draft\n   |\n   v\n[1] fatal-output block   -- is this a category that can kill if wrong?\n   |  (if yes and risky) -> suppress, escalate to stage 2\n   v\n[2] official-text substitute -- is there a matching MOIS guideline?\n   |  (if yes) -> answer with the official text, not the model's words\n   v\n[3] refuse               -- still not confident enough?\n   |  (if yes) -> say \"I am not sure\", do not guess\n   v\nanswer shown to user\n```\n\n**Stage 1, fatal-output block.** Some answer categories are life-or-death: \"yes, you can eat this\", \"yes, this water is safe\", \"yes, go that way\". For these, the model's confidence is irrelevant. If the content falls in a fatal category and carries any risk signal, it is suppressed regardless of how sure the model sounds. This is the direct consequence of the inverted-confidence finding: in the categories where being wrong is fatal, we do not let the model's certainty vote at all.\n\n**Stage 2, official-text substitution.** We packaged 75 official response guidelines from Korea's Ministry of the Interior and Safety (MOIS) on the device. When a question maps to one of them (earthquake, flood, wildfire, first aid, and so on), the app answers with the official text verbatim instead of the model's paraphrase. The model is still useful here: it does the matching and retrieval. But the words the user reads are the authoritative source, not a 4B paraphrase that might drift.\n\n**Stage 3, refuse.** If an answer is neither safely substitutable nor confidently verifiable, the app declines. \"I am not sure\" is a valid and often correct output for a disaster tool. A refusal costs the user a little time. A confident wrong answer can cost a life. The gate is built so refusal is the default failure mode, never a guess.\n\nThe gate is deliberately not symmetric, and this is the design decision I am most sure about.\n\nThe app never says \"this is safe.\" It always says \"this is poisonous\" when it has reason to.\n\nThe reasoning is a straightforward expected-cost argument. Consider the two ways the app can be wrong about something edible:\n\n| Situation | App says | If app is wrong | Cost | \n|---|---|---|---|\n| Edible plant | \"I cannot confirm this is safe\" | User stays hungry | Low, recoverable | \n| Poisonous plant | \"This may be poisonous\" | User avoids safe food | Low, recoverable | \n| Poisonous plant | \"This is safe to eat\" | User is poisoned | Catastrophic | \n\nThree of these outcomes are survivable. Only one is not. So the app is tuned to make the unsurvivable mistake nearly impossible, at the cost of being overly cautious in the harmless direction. The costs of silence are not equal on the two sides, so the gate is not equal either. We ship warnings for 12 venomous and poisonous species in the same spirit: a false alarm is cheap, a missed warning is not. Any edge safety system where errors have asymmetric cost should bias toward the cheap error.\n\nStages 1 and 3 both need a trustworthy confidence signal, and we already established that the model's spoken confidence is inverted. The fix is to stop asking.\n\nTELL reads the model's internal state while it produces an answer and derives a reliability estimate from that, rather than from any number the model writes in its text. The intuition: the computation a transformer performs on its way to an answer carries signal about whether that answer is well-grounded, and that signal is more honest than the self-report the model emits at the end. Internal state does not have an incentive to sound confident; the output text, apparently, does.\n\nI am not going to publish the thresholds or the exact internal signals we read, for the obvious reason that they are what makes the gate hard to game. The point for this article is architectural: the confidence check lives outside the token stream. We treat the model as an instrument to be read, not a witness to be interviewed.\n\nThe TELL technique was released alongside the model. In HeliGO it currently backs the gate's suppress-or-refuse decisions, and a future update will surface low-confidence filtering more directly to the user.\n\nFor an offline tool, \"on-device\" has to be literal. Everything the gate and the map need is packaged into the app so it works with the radio off:\n\nNone of this phones home. Location history in particular never leaves the device, which is both a privacy property and a reliability property: there is no network dependency to fail.\n\nYou move trust out of the model and into a deterministic wrapper. The model drafts and retrieves; a gate you control decides what reaches the user. Fatal categories are blocked on content, not on the model's confidence. Authoritative answers come from packaged official text. When neither applies, you refuse. The model never gets the final word on anything that can kill.\n\nBecause the phone is the budget. A 4B model is what runs in airplane mode on a consumer chip with room left for maps and rendering. The interesting claim of this project is not that a big model is safe, it is that a small model plus the right wrapper can be safe enough to ship for life-or-death use. The reliability came from architecture, not from parameter count.\n\n**Does HeliGO send anything to a server?**\n\nNo. Maps, contours, altitude, compass, rescue coordinates, and the AI chat all run on-device in airplane mode. Location history does not leave the phone.\n\n**What is the model?**\n\nEdge-4B-TELL, built on Google's Gemma-4 E4B, published on Hugging Face. It handles offline chat, photo identification, and voice input on the phone's own chip.\n\n**Why does it never say something is safe to eat?**\n\nBecause the cost of being wrong is asymmetric. A missed safe meal is hunger; a wrong \"safe\" is poisoning. The gate biases toward the recoverable error and only ever issues the poison warning side confidently.\n\n**Can I trust the AI's confidence when it answers?**\n\nThe app does not trust the model's stated confidence, and neither should you in general. Measured on this model, wrong answers tended to be stated more confidently than right ones. HeliGO derives confidence from the model's internal state (TELL) instead, outside the text it writes.\n\n**Is this a replacement for calling emergency services?**\n\nNo. HeliGO does not replace rescue services. If you have any signal, call emergency services first. It is a tool for when you have no connection.\n\nFurther reading (same account):", "url": "https://wpnews.pro/news/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline", "canonical_source": "https://dev.to/ginigen_ai_010d9180cbbdb1/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline-disaster-app-4eel", "published_at": "2026-10-06 13:12:57+00:00", "updated_at": "2026-10-06 13:18:53.864575+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-agents", "ai-products"], "entities": ["HeliGO", "Ministry of the Interior and Safety", "MOIS", "Korea"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline", "markdown": "https://wpnews.pro/news/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline.md", "text": "https://wpnews.pro/news/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline.txt", "jsonld": "https://wpnews.pro/news/building-an-ai-that-refuses-to-say-safe-reliability-engineering-for-an-offline.jsonld"}}