How to Audit an ML Design Before It Ships Karmendra Pandey, a Practice Architect in AI & ML at TEKsystems, has published a five-gate audit for reviewing machine learning designs before they ship, covering data provenance and leakage, decision-cost alignment, latency and cost budgets, security and prompt-injection exposure, and accountability with rollback plans. Pandey, who builds production agentic AI systems on AWS, argues that most ML failures are design failures visible before launch, and recommends teams complete a single-page checklist at design review where any blank answer blocks the ship. He notes the security gate matters roughly ten times more for agentic systems, where every unnecessary tool call burns tokens and inflates downstream context windows. Most ML failures are not model failures. They are design failures — and they are almost always visible before launch to anyone who knows where to look. Training-serving skew nobody checked. An evaluation set that doesn't match production. A feedback loop that will quietly poison the next retraining run. Cost and latency budgets nobody wrote down. I've watched these kill systems that had perfectly good models. After 19+ years of shipping software and reviewing AI research, I now run every ML design through the same five-gate audit before it ships. Here it is. Ask three questions. Where did every training example come from, and can you prove it? Could any information from the future — or from the test set — have leaked into training? And what happens when the real world drifts away from your training distribution? The classic killer here is eval leakage: the "99% accurate" classifier that fails on its first real day because the eval set was accidentally drawn from training data. It happens more than anyone admits. Demand a written data provenance statement and a leakage check as merge requirements, not nice-to-haves. A model can ace its metric and still fail its job. Offline accuracy means nothing if the production decision has different costs for different errors — a fraud model optimized for accuracy will happily approve everything when fraud is rare. For every model, write down: what decision does this model actually drive, what does a mistake cost in each direction, and does our eval set look like production traffic? If the answers are vague, the design isn't ready. Nobody writes down the latency budget until the first user complaint. Nobody prices the prediction until the first cloud bill. For each model, record: p95 latency budget, cost per 1,000 predictions at expected volume, and what happens when the model is down or too slow. "The page breaks" is not a fallback strategy — cache the last good prediction, degrade to a simpler model, or fail open explicitly and loudly. For agentic systems this gate matters 10x more: every unnecessary tool call burns tokens and inflates every downstream context window. Cost discipline starts with call discipline. Who can feed inputs to this model, and what can they make it do? Cover the basics: input validation, rate limiting, access controls on the model endpoint and its training data. For LLM-backed systems, add prompt-injection review — treat every tool the agent can call as a privilege, and scope it like one. I model agents as IAM principals: unique identity, least-privilege roles, session-scoped credentials. When the model is wrong — and it will be — who is accountable, and what is the remediation path? Every production ML system needs a named owner, a rollback plan, and a written policy for the failure modes you know exist. "The model decided" is not an accountability structure. Turn these gates into a single page your team fills out at design review: Any blank answer blocks the ship. It takes an hour, and it catches the failures that postmortems always describe as "obvious in hindsight." Karmendra Pandey is a Practice Architect in AI & ML at TEKsystems. He builds production agentic AI systems on AWS and peer-reviews AI research on PREreview.