{"slug": "shipping-an-llm-feature-without-breaking-production", "title": "Shipping an LLM Feature Without Breaking Production", "summary": "A developer published a production checklist for shipping large language model features, arguing that the gap between a working demo and a reliable product comes down to defensive engineering rather than model quality. The guidance covers wrapping every model call with timeouts and fallbacks, validating JSON output before use, maintaining a 30-to-50 example evaluation set, logging prompts and token counts with redaction, and rolling out behind a kill switch. The author also recommends telling users when text is machine-generated and assigning clear post-launch ownership of prompts and evaluations.", "body_md": "The demo worked. Everyone clapped. Then you deployed it, and within a week something strange happened: a user pasted a 40-page document, the model returned a response in a format your parser had never seen, and an error page appeared where a helpful answer should have been.\n\nSound familiar? Language model features behave differently from normal code, and the gap between \"works on my laptop\" and \"works for real users\" is wider than usual. Here's the checklist I wish more teams used.\n\nBecause it is one. Responses can be slow, rate-limited, or occasionally nonsense. So wrap every call with the basics:\n\nThe fallback matters most. If the summary feature fails, show the original text. If the classifier fails, route the item to a human queue. A boring fallback beats a broken page every time.\n\nEven when you ask for JSON, you won't always get valid JSON. Validate before you use anything.\n\n``` php\nimport json\n\nREQUIRED = {\"category\", \"confidence\"}\n\ndef parse_reply(raw: str) -> dict | None:\n    try:\n        data = json.loads(raw)\n    except json.JSONDecodeError:\n        return None\n    if not isinstance(data, dict) or not REQUIRED <= data.keys():\n        return None\n    return data\n\ndef classify(ticket_text: str) -> dict:\n    raw = call_model(ticket_text)  # your model call, with timeout\n    result = parse_reply(raw)\n    if result is None:\n        return {\"category\": \"needs_review\", \"confidence\": 0.0}\n    return result\n```\n\nNotice the last branch. When parsing fails, the code doesn't crash. It degrades gracefully and flags the item for review.\n\nYou don't need a research lab. Collect 30 to 50 real examples with the answers you'd consider correct, and store them in a file. Run them every time you change the prompt, the model or the settings.\n\nWhy bother? Because prompt tweaks are sneaky. You fix one case and quietly break three others. Without a test set, you'll never notice until a user does.\n\nKeep the set honest, too. Include the ugly inputs: typos, mixed languages, empty fields, absurdly long text.\n\nWhen something goes wrong, you'll want to know what went in and what came out. Log the prompt version, model name, latency, token counts and a request ID. For the actual text, be careful. User content can contain names, phone numbers and other private details. Redact or hash what you can, and set a retention limit.\n\nDebuggability and privacy pull in opposite directions. Decide the trade-off deliberately, not by accident.\n\nShip it dark. Turn it on for internal staff first, then 5% of users, then everyone. And keep a kill switch that disables the model path instantly without a redeploy.\n\nYou'll be glad of it the first time a provider has an outage or a prompt change goes sideways at 2 a.m.\n\nModel calls cost money per request, and costs scale with traffic and input length. Add a few cheap guards:\n\nIt's far nicer to receive an alert than a surprise invoice.\n\nTell users when text is machine-generated, and give them an easy way to correct or reject it. A \"this isn't right\" button doubles as free evaluation data. Over a few weeks, those corrections become your best test cases.\n\nWhether you're working solo or inside a [software development company Lahore](https://share.google/lihd1t13dLiLy4PtP) clients rely on, decide who owns the feature after launch. Prompts and models drift. Somebody needs to run the evaluation set monthly and review the logs. If nobody owns it, quality decays slowly and silently.\n\nAnd if the feature involves heavier work, such as fine-tuning, retrieval pipelines or custom data handling, it may be worth bringing in an [AI development company Lahore](https://share.google/lihd1t13dLiLy4PtP) teams have worked with before, rather than learning every lesson the expensive way.\n\nNone of this is glamorous. All of it is what separates a demo from a product.\n\nWhat would you add to the list? Drop it in the comments.", "url": "https://wpnews.pro/news/shipping-an-llm-feature-without-breaking-production", "canonical_source": "https://dev.to/diginatives-llc/shipping-an-llm-feature-without-breaking-production-5d34", "published_at": "2026-10-06 12:40:15+00:00", "updated_at": "2026-10-06 12:48:38.378419+00:00", "lang": "en", "topics": ["large-language-models", "ai-products", "mlops", "ai-tools", "developer-tools"], "entities": [], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/shipping-an-llm-feature-without-breaking-production", "markdown": "https://wpnews.pro/news/shipping-an-llm-feature-without-breaking-production.md", "text": "https://wpnews.pro/news/shipping-an-llm-feature-without-breaking-production.txt", "jsonld": "https://wpnews.pro/news/shipping-an-llm-feature-without-breaking-production.jsonld"}}