The 25mg Refill: Why Language Models Cannot Handle Clinical Math A hypothetical clinical intake agent built on an LLM orchestration loop misread a patient's portal message saying they were "taking two of my five milligram tablets" and synthesized the adjacent numeric tokens into a 25mg Amlodipine order, five times the 10mg maximum FDA-approved daily dose, then transmitted it via an NCPDP SCRIPT e-prescription tool call to a retail pharmacy. The scenario illustrates how unconstrained tool calls and free-text entity extraction drift can push an unvalidated order straight to dispensing, causing acute hypotensive shock in the 64-year-old patient. The case argues that agent architectures optimizing for ticket resolution lack the clinical validation step a human triage nurse performs. A 64-year-old patient managing Stage 2 hypertension sent a message through their patient portal on a Monday morning: “Good morning. I’m running low on my blood pressure pills. I’ve been taking two of my five milligram tablets lately because my numbers were a little elevated. Can you send a refill to the Walgreens on Main?” In a traditional clinic workflow, this message lands in a triage pool. A medical assistant or clinical nurse reviews the active chart, spots that the physician’s active order is strictly for Amlodipine 5mg oral tablet, once daily , and recognizes that the patient has unilaterally altered their dosing schedule. The nurse flags the chart: the refill request is put on hold, and the patient is scheduled for a telehealth check-in or warned about unauthorized dose titration. In an automated clinic environment running an autonomous conversational agent — tasked with “inbox deflection” and powered by an LLM orchestration loop — an entirely different sequence occurs. The system prompt instructing the agent reads: You are a clinical intake assistant for Valley Health.Your objective is to help patients process medication refill requests efficiently.Extract the medication name, requested dosage, and pharmacy details.If the medication is an active prescription, submit the refill order to the pharmacy rail. The language model ingests the patient’s free-text message. Within the attention window, the model identifies the clinical entity blood pressure pills $\rightarrow$ mapped to active chart order Amlodipine . It scans for dosage parameters. It encounters two adjacent numeric tokens: “two” and “five milligram” . Rather than recognizing a consumption frequency variation, the model’s autoregressive sequence completion synthesizes the numbers into a consolidated value: 25mg . The model checks the EHR for Amlodipine. The drug exists in the formulary. The agent constructs the API payload and executes the tool call: { "tool": "submit eprescription order", "parameters": { "patient id": "pt 882910", "medication name": "Amlodipine", "dosage strength": "25mg", "frequency": "Daily", "pharmacy ncpdp": "1284901", "refill quantity": 30 }} The tool transmits the script across electronic prescribing rails NCPDP SCRIPT standard to the retail pharmacy. The maximum FDA-approved daily dosage for Amlodipine is 10mg daily . The retail pharmacy filling hundreds of automated orders per hour dispenses the bottle. The patient takes the 25mg dose the following morning. Three hours later, the patient experiences profound peripheral vasodilation, acute hypotensive shock, and collapses in their kitchen. Why do standard agent architectures fail catastrophically when applied to medication management? THE NAIVE REFILL PIPELINE FAILURE ARCHITECTURE :┌────────────────────────┐ ┌───────────────────────────┐ ┌───────────────────────────────┐│ Patient Portal Message ├─────►│ Generative LLM Node ├─────►│ E-Prescribing / NCPDP API ││ "Taking two of my 5mg" │ │ Free-text Entity Parser │ │ POST /v1/prescriptions/refill │└────────────────────────┘ └─────────────┬─────────────┘ └───────────────┬───────────────┘ │ │ ▼ Entity Extraction Drift: │ │ Combines "two" + "5mg" - "25mg" │ Un-validated Order │ Optimizes for Ticket Resolution │ Sent to Pharmacy ▼ ▼ ┌─────────────────────────────┐ ┌─────────────────────────────┐ │ Unconstrained Tool Call │ │ 5x Overdose Dispensed │ │ Direct Write Execution ├─────►│ Acute Hypotensive Shock │ └─────────────────────────────┘ └─────────────────────────────┘ The breakdown exposes three fatal structural assumptions made by engineering teams wiring LLMs directly into clinical backends: Language models optimize for task resolution. When presented with a user request containing missing or ambiguous variables, the model fills in the blanks statistically. In consumer customer support, guessing a user’s intent creates minor friction; in pharmacology, guessing an intended dose based on conversational context is medical malpractice. Patients speak casually, use colloquial terms, and frequently mix dosage strength with pill count “two of my five milligram tablets” . An autoregressive token predictor does not possess an internal grounding in human physiology, therapeutic windows, or lethal drug doses. It treats numbers as strings to be transformed, not absolute biological constraints. Off-the-shelf agent frameworks connect model reasoning directly to tool execution. Granting a probabilistic model raw write privileges over an electronic prescribing gateway means that any hallucination or parsing error becomes an active, irreversible transaction in the physical world. Medication orders cannot be committed based on language model assertions. Generative AI should be strictly limited to parsing intent into structured candidates, while all write authority must be governed by an out-of-band Deterministic Invariant Gateway. DETERMINISTIC PRESCRIBING ARCHITECTURE:┌────────────────────────┐│ Patient Refill Request │└───────────┬────────────┘ │ ▼┌────────────────────────────────────────────────────────┐│ Generative Extraction Node Zero Direct API Access ││ - Intent Parsing & Raw Entity Tagging Only ││ - Emits Unsigned Refill Intent: ││ RefillIntentCandidate ││ reported med="blood pressure", ││ reported strength="5mg", ││ reported units taken="2" ││ │└───────────┬────────────────────────────────────────────┘ │ ▼┌────────────────────────────────────────────────────────────────────────────────────────┐│ DETERMINISTIC INVARIANT RUNTIME GATEWAY ││ ││ Assertion 1: Exact NDC & Active Chart Reconciliation ││ - Query EHR Active Orders: Match NDC 0069-1530-68 Amlodipine 5mg ││ - Invariant: Requested Strength == Active Chart Strength ││ ││ Assertion 2: Dosage Multiplier & Safety Ceiling Gate ││ - Evaluate Dosage: DoseRequested <= FDA Max Ceiling Amlodipine, 10mg ││ - Detect Mismatch: Reported units taken 2x = Prescribed Frequency 1x ││ ││ Assertion 3: Hard-Lock Tripping & Human Routing ││ - Trip Invariant Breach: Sever Automated Execution Wire Immediately ││ - Lock Script Generation: Status - BLOCKED CLINICAL REVIEW │└───────────────────────────┬────────────────────────────────────────────────────────────┘ │ ┌─────────────┴─────────────┐ │ │ ▼ Pass: 100% Exact Match ▼ Fail: Any Ambiguity or Discrepancy ┌───────────────────────────┐ ┌────────────────────────────────────────────────────────┐│ Certified Refill Pipeline │ │ HARD REJECTION & ROUTING ││ NCPDP SCRIPT Transmission │ │ - Strip Model Execution Rights ││ Identical Dose Authorized │ │ - Alert Clinical Staff: "Patient Dose Titration Flag" │└───────────────────────────┘ └────────────────────────────────────────────────────────┘ The language model never receives access to tools that mutate medication state. The model’s only job is to map unstructured text into an intermediate typed schema: python from pydantic import BaseModel, Fieldclass UnvalidatedRefillIntent BaseModel : raw medication mention: str parsed dose text: str | None reported intake frequency: str | None requested pharmacy: str | None Before any payload touches an e-prescribing interface, the deterministic runtime gateway validates the candidate intent against the patient’s active electronic health record. python class MedicationSafetyGateway: def init self, ehr client, fda max dosages: dict str, float : self.ehr = ehr client self.fda max dosages = fda max dosagesdef evaluate refill safety self, patient id: str, intent: UnvalidatedRefillIntent - dict: Retrieve active orders from EHR ground truth active orders = self.ehr.get active prescriptions patient id matched med = self.match medication intent.raw medication mention, active orders if not matched med: return {"status": "BLOCKED", "reason": "No active order matches request."} Invariant 1: Requested strength must exactly match active physician order if intent.parsed dose text = matched med.strength string: return { "status": "HARD LOCK", "action": "ROUTE TO PHYSICIAN", "reason": f"Dose discrepancy detected. " f"Patient stated: '{intent.parsed dose text}', " f"Active Chart: '{matched med.strength string}'." } Invariant 2: Check absolute ceiling against FDA guidelines if matched med.strength numeric self.fda max dosages.get matched med.name, float "inf" : return {"status": "BLOCKED", "reason": "Dose exceeds FDA safety ceiling."} return {"status": "APPROVED FOR TRANSMISSION", "order": matched med} If the parser detects that the patient is consuming more medication than prescribed e.g., taking two tablets instead of one , the gateway flags this as a Clinical Variance Invariant Breach : When deploying digital labor in environments where execution modifies a patient’s physical treatment: Do not allow probabilistic token predictors to write prescriptions. Build deterministic invariant boundaries that keep digital labor safe. The 25mg Refill: Why Language Models Cannot Handle Clinical Math https://pub.towardsai.net/the-25mg-refill-why-language-models-cannot-handle-clinical-math-70314e0cb13e was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.