Shipping an LLM Feature Without Breaking Production A developer published a production checklist for shipping large language model features, arguing that the gap between a working demo and a reliable product comes down to defensive engineering rather than model quality. The guidance covers wrapping every model call with timeouts and fallbacks, validating JSON output before use, maintaining a 30-to-50 example evaluation set, logging prompts and token counts with redaction, and rolling out behind a kill switch. The author also recommends telling users when text is machine-generated and assigning clear post-launch ownership of prompts and evaluations. The demo worked. Everyone clapped. Then you deployed it, and within a week something strange happened: a user pasted a 40-page document, the model returned a response in a format your parser had never seen, and an error page appeared where a helpful answer should have been. Sound familiar? Language model features behave differently from normal code, and the gap between "works on my laptop" and "works for real users" is wider than usual. Here's the checklist I wish more teams used. Because it is one. Responses can be slow, rate-limited, or occasionally nonsense. So wrap every call with the basics: The fallback matters most. If the summary feature fails, show the original text. If the classifier fails, route the item to a human queue. A boring fallback beats a broken page every time. Even when you ask for JSON, you won't always get valid JSON. Validate before you use anything. php import json REQUIRED = {"category", "confidence"} def parse reply raw: str - dict | None: try: data = json.loads raw except json.JSONDecodeError: return None if not isinstance data, dict or not REQUIRED <= data.keys : return None return data def classify ticket text: str - dict: raw = call model ticket text your model call, with timeout result = parse reply raw if result is None: return {"category": "needs review", "confidence": 0.0} return result Notice the last branch. When parsing fails, the code doesn't crash. It degrades gracefully and flags the item for review. You don't need a research lab. Collect 30 to 50 real examples with the answers you'd consider correct, and store them in a file. Run them every time you change the prompt, the model or the settings. Why bother? Because prompt tweaks are sneaky. You fix one case and quietly break three others. Without a test set, you'll never notice until a user does. Keep the set honest, too. Include the ugly inputs: typos, mixed languages, empty fields, absurdly long text. When something goes wrong, you'll want to know what went in and what came out. Log the prompt version, model name, latency, token counts and a request ID. For the actual text, be careful. User content can contain names, phone numbers and other private details. Redact or hash what you can, and set a retention limit. Debuggability and privacy pull in opposite directions. Decide the trade-off deliberately, not by accident. Ship it dark. Turn it on for internal staff first, then 5% of users, then everyone. And keep a kill switch that disables the model path instantly without a redeploy. You'll be glad of it the first time a provider has an outage or a prompt change goes sideways at 2 a.m. Model calls cost money per request, and costs scale with traffic and input length. Add a few cheap guards: It's far nicer to receive an alert than a surprise invoice. Tell users when text is machine-generated, and give them an easy way to correct or reject it. A "this isn't right" button doubles as free evaluation data. Over a few weeks, those corrections become your best test cases. Whether you're working solo or inside a software development company Lahore https://share.google/lihd1t13dLiLy4PtP clients rely on, decide who owns the feature after launch. Prompts and models drift. Somebody needs to run the evaluation set monthly and review the logs. If nobody owns it, quality decays slowly and silently. And if the feature involves heavier work, such as fine-tuning, retrieval pipelines or custom data handling, it may be worth bringing in an AI development company Lahore https://share.google/lihd1t13dLiLy4PtP teams have worked with before, rather than learning every lesson the expensive way. None of this is glamorous. All of it is what separates a demo from a product. What would you add to the list? Drop it in the comments.