Build Agents That Keep Their Rules: What to Test Before You Ship a Vertical Agent Tej Pandya, founder of GrowEasy.ai, argues that vertical AI agents succeed by addressing three failure points general agents leave to users: multi-turn instruction drift, underspecified prompts, and a blank-box UX. He cites the ICLR 2026 paper "LLMs Get Lost in Multi-Turn Conversation," which reports an average 39% drop across six generation tasks when instructions arrive step by step, and an ACL Findings 2026 paper finding underspecified prompts roughly twice as likely to regress across model or prompt changes, and recommends regression checks that replay agent rules after many runs and after every model or prompt change. Subtitle: Multi-turn drift, underspecified prompts and a blank-box UX are the three failure points to design around. By Tej Pandya, founder of GrowEasy.ai I think vertical agents will do well because they fix three problems a general agent leaves to the user. If you build one, these are the things to test. In my own work, agreed rules slip a few deliveries later. That is my observation, not a measured rate. The ICLR 2026 paper "LLMs Get Lost in Multi-Turn Conversation" reports an average 39% drop across six generation tasks when instructions arrive step by step 15 models, 200,000+ simulated conversations . The ACL Findings 2026 paper "What Prompts Don't Say" reports underspecified prompts are about 2x as likely to regress across model or prompt changes. Both are lab tests of underspecified instructions. Build a regression check that replays your rules after many runs and after every model or prompt change. DETAIL Kim, Dec 2025; 30 tasks, GPT-4 and o3-mini found specificity improved accuracy, most for smaller models and procedural tasks. Weak evidence, but it matches practice. State the goal, steps, limits and quality bar, and ship them with the agent, so users do not write them. A CMU study 31 participants, Operator and Manus found usability barriers with general agents, including capabilities that did not match user expectations. It does not say people lack ideas. My opinion: a blank box asks users to invent the job, so ship an agent for one job. Boom forecasts come from investors and analysts. Thin wrappers can be absorbed by model vendors. Narrow agents still need rule checks. For legal, health and finance work, keep a human sign-off. Sources: ICLR 2026 https://proceedings.iclr.cc/paper files/paper/2026/file/59f6421e64707225fdf5b28840679a07-Paper-Conference.pdf , ACL Findings 2026 https://aclanthology.org/2026.findings-acl.441.pdf , CMU study https://arxiv.org/abs/2509.14528 , DETAIL https://arxiv.org/html/2512.02246v1 . Video version: https://youtu.be/bGr6M1ddN U https://youtu.be/bGr6M1ddN U