cd /news/large-language-models/shipping-an-llm-feature-without-brea… · home › topics › large-language-models › article
[ARTICLE · art-146021] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Shipping an LLM Feature Without Breaking Production

A developer published a production checklist for shipping large language model features, arguing that the gap between a working demo and a reliable product comes down to defensive engineering rather than model quality. The guidance covers wrapping every model call with timeouts and fallbacks, validating JSON output before use, maintaining a 30-to-50 example evaluation set, logging prompts and token counts with redaction, and rolling out behind a kill switch. The author also recommends telling users when text is machine-generated and assigning clear post-launch ownership of prompts and evaluations.

by read3 min views2 publishedOct 6, 2026

The demo worked. Everyone clapped. Then you deployed it, and within a week something strange happened: a user pasted a 40-page document, the model returned a response in a format your parser had never seen, and an error page appeared where a helpful answer should have been.

Sound familiar? Language model features behave differently from normal code, and the gap between "works on my laptop" and "works for real users" is wider than usual. Here's the checklist I wish more teams used.

Because it is one. Responses can be slow, rate-limited, or occasionally nonsense. So wrap every call with the basics:

The fallback matters most. If the summary feature fails, show the original text. If the classifier fails, route the item to a human queue. A boring fallback beats a broken page every time.

Even when you ask for JSON, you won't always get valid JSON. Validate before you use anything.

import json

REQUIRED = {"category", "confidence"}

def parse_reply(raw: str) -> dict | None:
    try:
        data = json.loads(raw)
    except json.JSONDecodeError:
        return None
    if not isinstance(data, dict) or not REQUIRED <= data.keys():
        return None
    return data

def classify(ticket_text: str) -> dict:
    raw = call_model(ticket_text)  # your model call, with timeout
    result = parse_reply(raw)
    if result is None:
        return {"category": "needs_review", "confidence": 0.0}
    return result

Notice the last branch. When parsing fails, the code doesn't crash. It degrades gracefully and flags the item for review.

You don't need a research lab. Collect 30 to 50 real examples with the answers you'd consider correct, and store them in a file. Run them every time you change the prompt, the model or the settings.

Why bother? Because prompt tweaks are sneaky. You fix one case and quietly break three others. Without a test set, you'll never notice until a user does.

Keep the set honest, too. Include the ugly inputs: typos, mixed languages, empty fields, absurdly long text.

When something goes wrong, you'll want to know what went in and what came out. Log the prompt version, model name, latency, token counts and a request ID. For the actual text, be careful. User content can contain names, phone numbers and other private details. Redact or hash what you can, and set a retention limit.

Debuggability and privacy pull in opposite directions. Decide the trade-off deliberately, not by accident.

Ship it dark. Turn it on for internal staff first, then 5% of users, then everyone. And keep a kill switch that disables the model path instantly without a redeploy.

You'll be glad of it the first time a provider has an outage or a prompt change goes sideways at 2 a.m.

Model calls cost money per request, and costs scale with traffic and input length. Add a few cheap guards:

It's far nicer to receive an alert than a surprise invoice.

Tell users when text is machine-generated, and give them an easy way to correct or reject it. A "this isn't right" button doubles as free evaluation data. Over a few weeks, those corrections become your best test cases.

Whether you're working solo or inside a software development company Lahore clients rely on, decide who owns the feature after launch. Prompts and models drift. Somebody needs to run the evaluation set monthly and review the logs. If nobody owns it, quality decays slowly and silently.

And if the feature involves heavier work, such as fine-tuning, retrieval pipelines or custom data handling, it may be worth bringing in an AI development company Lahore teams have worked with before, rather than learning every lesson the expensive way.

None of this is glamorous. All of it is what separates a demo from a product.

What would you add to the list? Drop it in the comments.

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/shipping-an-llm-feat…] indexed:0 read:3min 2026-10-06 · —