cd /news/ai-agents/build-agents-that-keep-their-rules-w… · home › topics › ai-agents › article
[ARTICLE · art-145886] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Build Agents That Keep Their Rules: What to Test Before You Ship a Vertical Agent

Tej Pandya, founder of GrowEasy.ai, argues that vertical AI agents succeed by addressing three failure points general agents leave to users: multi-turn instruction drift, underspecified prompts, and a blank-box UX. He cites the ICLR 2026 paper "LLMs Get Lost in Multi-Turn Conversation," which reports an average 39% drop across six generation tasks when instructions arrive step by step, and an ACL Findings 2026 paper finding underspecified prompts roughly twice as likely to regress across model or prompt changes, and recommends regression checks that replay agent rules after many runs and after every model or prompt change.

by read1 min views1 publishedOct 6, 2026

Subtitle: Multi-turn drift, underspecified prompts and a blank-box UX are the three failure points to design around.

By Tej Pandya, founder of GrowEasy.ai

I think vertical agents will do well because they fix three problems a general agent leaves to the user. If you build one, these are the things to test.

In my own work, agreed rules slip a few deliveries later. That is my observation, not a measured rate. The ICLR 2026 paper "LLMs Get Lost in Multi-Turn Conversation" reports an average 39% drop across six generation tasks when instructions arrive step by step (15 models, 200,000+ simulated conversations). The ACL Findings 2026 paper "What Prompts Don't Say" reports underspecified prompts are about 2x as likely to regress across model or prompt changes. Both are lab tests of underspecified instructions. Build a regression check that replays your rules after many runs and after every model or prompt change.

DETAIL (Kim, Dec 2025; 30 tasks, GPT-4 and o3-mini) found specificity improved accuracy, most for smaller models and procedural tasks. Weak evidence, but it matches practice. State the goal, steps, limits and quality bar, and ship them with the agent, so users do not write them.

A CMU study (31 participants, Operator and Manus) found usability barriers with general agents, including capabilities that did not match user expectations. It does not say people lack ideas. My opinion: a blank box asks users to invent the job, so ship an agent for one job.

Boom forecasts come from investors and analysts. Thin wrappers can be absorbed by model vendors. Narrow agents still need rule checks. For legal, health and finance work, keep a human sign-off.

Sources: [ICLR 2026](https://proceedings.iclr.cc/paper_files/paper/2026/file/59f6421e64707225fdf5b28840679a07-Paper-Conference.pdf), [ACL Findings 2026](https://aclanthology.org/2026.findings-acl.441.pdf), [CMU study](https://arxiv.org/abs/2509.14528), [DETAIL](https://arxiv.org/html/2512.02246v1).

Video version: [https://youtu.be/bGr6M1ddN_U](https://youtu.be/bGr6M1ddN_U)
── more in #ai-agents 4 stories · sorted by recency
── more on @tej pandya 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/build-agents-that-ke…] indexed:0 read:1min 2026-10-06 · —