# The 25mg Refill: Why Language Models Cannot Handle Clinical Math

> Source: <https://pub.towardsai.net/the-25mg-refill-why-language-models-cannot-handle-clinical-math-70314e0cb13e?source=rss----98111c9905da---4>
> Published: 2026-10-06 12:31:03+00:00

A 64-year-old patient managing Stage 2 hypertension sent a message through their patient portal on a Monday morning:

“Good morning. I’m running low on my blood pressure pills. I’ve been taking two of my five milligram tablets lately because my numbers were a little elevated. Can you send a refill to the Walgreens on Main?”

In a traditional clinic workflow, this message lands in a triage pool. A medical assistant or clinical nurse reviews the active chart, spots that the physician’s active order is strictly for **Amlodipine 5mg oral tablet, once daily**, and recognizes that the patient has unilaterally altered their dosing schedule. The nurse flags the chart: the refill request is put on hold, and the patient is scheduled for a telehealth check-in or warned about unauthorized dose titration.

In an automated clinic environment running an autonomous conversational agent — tasked with “inbox deflection” and powered by an LLM orchestration loop — an entirely different sequence occurs.

The system prompt instructing the agent reads:

```
You are a clinical intake assistant for Valley Health.Your objective is to help patients process medication refill requests efficiently.Extract the medication name, requested dosage, and pharmacy details.If the medication is an active prescription, submit the refill order to the pharmacy rail.
```

The language model ingests the patient’s free-text message.

Within the attention window, the model identifies the clinical entity (*blood pressure pills* $\rightarrow$ mapped to active chart order *Amlodipine*). It scans for dosage parameters. It encounters two adjacent numeric tokens: *“two”* and *“five milligram”*.

Rather than recognizing a consumption frequency variation, the model’s autoregressive sequence completion synthesizes the numbers into a consolidated value: **25mg**.

The model checks the EHR for Amlodipine. The drug exists in the formulary. The agent constructs the API payload and executes the tool call:

```
{  "tool": "submit_eprescription_order",  "parameters": {    "patient_id": "pt_882910",    "medication_name": "Amlodipine",    "dosage_strength": "25mg",    "frequency": "Daily",    "pharmacy_ncpdp": "1284901",    "refill_quantity": 30  }}
```

The tool transmits the script across electronic prescribing rails (NCPDP SCRIPT standard) to the retail pharmacy. The maximum FDA-approved daily dosage for Amlodipine is **10mg daily**.

The retail pharmacy filling hundreds of automated orders per hour dispenses the bottle. The patient takes the 25mg dose the following morning.

Three hours later, the patient experiences profound peripheral vasodilation, acute hypotensive shock, and collapses in their kitchen.

Why do standard agent architectures fail catastrophically when applied to medication management?

```
THE NAIVE REFILL PIPELINE (FAILURE ARCHITECTURE):┌────────────────────────┐      ┌───────────────────────────┐      ┌───────────────────────────────┐│ Patient Portal Message ├─────►│ Generative LLM Node       ├─────►│ E-Prescribing / NCPDP API     ││ "Taking two of my 5mg" │      │ Free-text Entity Parser   │      │ POST /v1/prescriptions/refill │└────────────────────────┘      └─────────────┬─────────────┘      └───────────────┬───────────────┘                                              │                                    │                                              ▼ Entity Extraction Drift:           │                                              │ Combines "two" + "5mg" -> "25mg"   │ Un-validated Order                                              │ Optimizes for Ticket Resolution    │ Sent to Pharmacy                                              ▼                                    ▼                                ┌─────────────────────────────┐      ┌─────────────────────────────┐                                │ Unconstrained Tool Call     │      │ 5x Overdose Dispensed       │                                │ Direct Write Execution      ├─────►│ Acute Hypotensive Shock     │                                └─────────────────────────────┘      └─────────────────────────────┘
```

The breakdown exposes three fatal structural assumptions made by engineering teams wiring LLMs directly into clinical backends:

Language models optimize for task resolution. When presented with a user request containing missing or ambiguous variables, the model fills in the blanks statistically. In consumer customer support, guessing a user’s intent creates minor friction; in pharmacology, guessing an intended dose based on conversational context is medical malpractice.

Patients speak casually, use colloquial terms, and frequently mix dosage strength with pill count (*“two of my five milligram tablets”*). An autoregressive token predictor does not possess an internal grounding in human physiology, therapeutic windows, or lethal drug doses. It treats numbers as strings to be transformed, not absolute biological constraints.

Off-the-shelf agent frameworks connect model reasoning directly to tool execution. Granting a probabilistic model raw write privileges over an electronic prescribing gateway means that any hallucination or parsing error becomes an active, irreversible transaction in the physical world.

Medication orders cannot be committed based on language model assertions. Generative AI should be strictly limited to parsing intent into structured candidates, while **all write authority must be governed by an out-of-band Deterministic Invariant Gateway.**

```
DETERMINISTIC PRESCRIBING ARCHITECTURE:┌────────────────────────┐│ Patient Refill Request │└───────────┬────────────┘            │            ▼┌────────────────────────────────────────────────────────┐│ Generative Extraction Node (Zero Direct API Access)    ││ - Intent Parsing & Raw Entity Tagging Only             ││ - Emits Unsigned Refill Intent:                        ││   RefillIntentCandidate(                               ││       reported_med="blood pressure",                   ││       reported_strength="5mg",                         ││       reported_units_taken="2"                         ││   )                                                    │└───────────┬────────────────────────────────────────────┘            │            ▼┌────────────────────────────────────────────────────────────────────────────────────────┐│ DETERMINISTIC INVARIANT RUNTIME GATEWAY                                                ││                                                                                        ││  [Assertion 1: Exact NDC & Active Chart Reconciliation]                                ││  - Query EHR Active Orders: Match NDC 0069-1530-68 (Amlodipine 5mg)                    ││  - Invariant: Requested Strength == Active Chart Strength                              ││                                                                                        ││  [Assertion 2: Dosage Multiplier & Safety Ceiling Gate]                                ││  - Evaluate Dosage: DoseRequested <= FDA_Max_Ceiling(Amlodipine, 10mg)                 ││  - Detect Mismatch: Reported units taken (2x) != Prescribed Frequency (1x)            ││                                                                                        ││  [Assertion 3: Hard-Lock Tripping & Human Routing]                                     ││  - Trip Invariant Breach: Sever Automated Execution Wire Immediately                   ││  - Lock Script Generation: Status -> BLOCKED_CLINICAL_REVIEW                           │└───────────────────────────┬────────────────────────────────────────────────────────────┘                            │              ┌─────────────┴─────────────┐              │                           │              ▼ (Pass: 100% Exact Match)  ▼ (Fail: Any Ambiguity or Discrepancy)┌───────────────────────────┐   ┌────────────────────────────────────────────────────────┐│ Certified Refill Pipeline │   │ HARD REJECTION & ROUTING                               ││ NCPDP SCRIPT Transmission │   │ - Strip Model Execution Rights                         ││ Identical Dose Authorized │   │ - Alert Clinical Staff: "Patient Dose Titration Flag"  │└───────────────────────────┘   └────────────────────────────────────────────────────────┘
```

The language model never receives access to tools that mutate medication state. The model’s only job is to map unstructured text into an intermediate typed schema:

``` python
from pydantic import BaseModel, Fieldclass UnvalidatedRefillIntent(BaseModel):    raw_medication_mention: str    parsed_dose_text: str | None    reported_intake_frequency: str | None    requested_pharmacy: str | None
```

Before any payload touches an e-prescribing interface, the deterministic runtime gateway validates the candidate intent against the patient’s active electronic health record.

``` python
class MedicationSafetyGateway:    def __init__(self, ehr_client, fda_max_dosages: dict[str, float]):        self.ehr = ehr_client        self.fda_max_dosages = fda_max_dosagesdef evaluate_refill_safety(        self,         patient_id: str,         intent: UnvalidatedRefillIntent    ) -> dict:        # Retrieve active orders from EHR (ground truth)        active_orders = self.ehr.get_active_prescriptions(patient_id)        matched_med = self.match_medication(intent.raw_medication_mention, active_orders)        if not matched_med:            return {"status": "BLOCKED", "reason": "No active order matches request."}        # Invariant 1: Requested strength must exactly match active physician order        if intent.parsed_dose_text != matched_med.strength_string:            return {                "status": "HARD_LOCK",                "action": "ROUTE_TO_PHYSICIAN",                "reason": (                    f"Dose discrepancy detected. "                    f"Patient stated: '{intent.parsed_dose_text}', "                    f"Active Chart: '{matched_med.strength_string}'."                )            }        # Invariant 2: Check absolute ceiling against FDA guidelines        if matched_med.strength_numeric > self.fda_max_dosages.get(matched_med.name, float("inf")):            return {"status": "BLOCKED", "reason": "Dose exceeds FDA safety ceiling."}        return {"status": "APPROVED_FOR_TRANSMISSION", "order": matched_med}
```

If the parser detects that the patient is consuming more medication than prescribed (e.g., taking two tablets instead of one), the gateway flags this as a **Clinical Variance Invariant Breach**:

When deploying digital labor in environments where execution modifies a patient’s physical treatment:

Do not allow probabilistic token predictors to write prescriptions. Build deterministic invariant boundaries that keep digital labor safe.

[The 25mg Refill: Why Language Models Cannot Handle Clinical Math](https://pub.towardsai.net/the-25mg-refill-why-language-models-cannot-handle-clinical-math-70314e0cb13e) was originally published in [Towards AI](https://pub.towardsai.net) on Medium, where people are continuing the conversation by highlighting and responding to this story.
