cd /news/ai-agents/tests-enforce-journeys-agents-verify… · home › topics › ai-agents › article
[ARTICLE · art-141717] src=mlnotes.substack.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Tests Enforce Journeys. Agents Verify Goals: How Slack Uses AI for End-to-End Testing

Slack's engineering team published findings from over 200 automated agentic test runs comparing three architectures: an agent driving Playwright via Model Context Protocol, an agent issuing Playwright CLI commands, and LLM-generated deterministic Playwright tests. The team found that agent-driven tests reached their intended goal on every run while varying their exact step sequence, but flagged cost and runtime concerns of roughly $15 to $30 and seven minutes per execution.

by read8 min views2 publishedSep 29, 2026
Tests Enforce Journeys. Agents Verify Goals: How Slack Uses AI for End-to-End Testing
Image: Mlnotes (auto-discovered)

Every software engineer and product manager has a love-hate relationship with end-to-end (E2E) testing.

You write an automated test suite with Cypress or Playwright. You script every individual click, type, and assertion across thirty sequential steps: click(button#checkout) -> type(input#promo, "SAVE20") -> wait(500) -> assert(text.contains("Discount Applied"))

Everything passes locally. Then a frontend engineer redesigns the checkout modal, renames a CSS class, or introduces a minor asynchronous animation delay.

Your next pull request turns red. The entire CI/CD pipeline grinds to a halt. A senior developer spends three hours investigating, only to realize the application was working perfectly: the test script was simply too brittle to survive a minor layout change.

In the AI era, the obvious question emerged: why not replace brittle test scripts with autonomous AI agents?

Instead of programming a rigid sequence of DOM selectors, you hand an AI agent a high-level goal: “Log in as a trial user, send a message in the #general channel, and verify that the message appears in the thread view.”

If a button moves two inches to the right, the agent simply looks at the screen, adapts, and clicks it anyway.

It sounds like developer paradise. But can an autonomous agent that costs $15 to $30 per execution and takes seven minutes to run actually fit into modern development workflows?

A few days ago, the engineering team at Slack published findings from over 200 automated agentic test runs across real test workspaces.

Their findings reveal a foundational truth that every engineering team needs to understand before plugging LLMs into their testing stack.

1. The Core Paradigm: Journeys vs. Goals #

The team at Slack distilled the difference between scripted tests and AI testing into a single architectural distinction:

Tests enforce journeys. Agents verify goals.

Traditional deterministic tests validate a specific, rigid journey through the user interface:

Step 1 (Click) -> Step 2 (Type) -> Step 3 (Select) -> Step 4 (Assert)

If anything along that exact path deviates by a single pixel or millisecond, the test breaks.

Agent-driven tests, by contrast, operate on declarative intent:

Goal Declaration -> Autonomous Perception -> Adaptive Execution -> Outcome Verification

Across 200+ runs inside Slack workspaces, the agent reached the intended goal every time, but its exact sequence of steps varied widely:

  • Alternative Input Patterns: In one run, the agent clicked a search suggestion dropdown; in the next run, it typed the query and pressed Enter.
  • Dynamic Navigation: Sometimes the agent navigated back via breadcrumbs; other times it used global keyboard shortcuts.
  • Self-Healing Actions: When a modal unexpectedly popped up, the agent closed the modal and resumed its task without throwing an unhandled exception.

This flexibility makes agent testing extraordinary for exploratory quality assurance. But it also introduces a massive economic reality check.

2. Slack’s 200-Run Experiment: MCP vs. CLI vs. Generated Code #

To evaluate whether agents could realistically replace standard testing, Slack benchmarked three distinct architectures across simple workflows (creating a channel and replying to a thread) and complex workflows (multi-step search, filter, and message discovery):

  1. Agent + Playwright MCP (Model Context Protocol): The agent interacts with a live browser instance via persistent protocol connection, inspecting real-time DOM snapshots.
  2. Agent + Playwright CLI: The agent executes headless Playwright commands via terminal shell commands, rebuilding state snapshot-by-snapshot.
  3. AI-Generated Playwright Tests: An LLM generates deterministic Playwright code upfront, runs the script, and iteratively edits the code until it passes.

Here is what the empirical data revealed:

The Two Major Lessons from the Data:

Lesson A: Direct Protocol (MCP) Crushes Shell CLI

The Playwright MCP connection achieved a 0% failure rate on standard flows and remained stable on complex workflows.

Why did CLI commands fail twice as often? Because firing individual shell commands forces the agent to reconstruct browser state from scratch on every turn. Authentication cookies drop, DOM focus is lost, and timing inconsistencies accumulate.

Direct protocol connections (MCP) maintain a live, stateful websocket connection to the browser, allowing the agent to reuse in-session context effortlessly.

Lesson B: Generating Static Code Fails on Complex Workflows

Many developers assume the best use of AI is asking an LLM to generate standard Playwright scripts to run in CI/CD.

Slack’s data proved the opposite: on complex workflows, generated scripts failed 48% of the time. The generated scripts fell right back into the brittle selector trap, breaking on asynchronous assertions and dynamic UI state.

3. The 3-Tier Architecture: Where Agents Actually Belong #

The biggest mistake an engineering team can make is trying to replace their existing pull-request CI/CD tests with AI agents.

If you run an agentic test on every git push:

  • A developer pushing three commits an hour waits 10 minutes per commit.
  • Your testing bill skyrockets to $25 per PR.
  • Flaky non-deterministic model runs delay production deployments.

Production teams do not replace deterministic tests with agents: they structure their testing into a 3-Tier Testing Pyramid:

Tier 1: Unit & Component Tests (Sub-Second CI Gate)

  • Execution: Deterministic, mocked, runs in milliseconds on every local save and commit.
  • Role: Catches logic errors and regression bugs before code ever leaves the developer’s laptop.

Tier 2: Happy-Path Scripted E2E Tests (PR Merge Gate)

  • Execution: Deterministic Playwright or Cypress scripts covering the top 5% of critical user journeys (e.g. signup, billing checkout, core creation flow).
  • Runtime: 2 to 3 minutes. Runs on every pull request.
  • Role: Enforces non-negotiable user journeys with zero tolerance for nondeterminism.

Tier 3: Autonomous Agentic Explorers (Nightly & Staging Sweeps)

  • Execution: Asynchronous AI agents connected via Playwright MCP, running against staging environments every night or after major feature releases.
  • Runtime: 5 to 15 minutes per goal.
  • Role: Acting as an automated exploratory tester. The agent is given loose user goals (“Try to break the workspace search filter with edge-case characters” ) to uncover subtle regression bugs and unhandled exceptions that scripted tests never anticipated.

4. The 3 Production Scars: Why Naive Agent Testing Breaks #

If you attempt to implement agentic E2E testing on your application, watch out for these three enterprise traps:

Scar #1: The Financial Runaway in CI Pipelines

A standard CI suite runs hundreds of times a day across an active engineering organization. If you trigger an agentic E2E workflow on every branch, you can easily rack up tens of thousands of dollars in LLM API bills within weeks.

  • The Guardrail: Restrict agentic testing toscheduled cron jobs (nightly runs) or trigger them only when high-risk architectural changes touch core shared libraries.

Scar #2: Non-Deterministic Flakiness Masks Real Bugs

When an agent takes a slightly different navigation path on run #14 and encounters an unhandled exception, did the application fail, or did the model make an erratic hallucinated decision?

  • The Guardrail: Every agentic test must record a full session video alongside a synchronized DOM action log. An alert should only fire if the agent can reproduce the exact failure across two consecutive independent seeds.

Scar #3: Dangerous Mutation in Staging Databases

Giving an autonomous browser agent open credentials on a shared staging environment is risky: an exploratory agent clicking around admin panels can easily wipe test tenant data or trigger thousands of automated test emails to real users.

  • The Guardrail: Run agentic tests exclusively against isolated, ephemeral test workspaces seeded with synthetic data, with outbound mail and webhook gateways strictly mocked.

5. How to Prototype This Weekend (Playwright MCP Blueprint) #

You can spin up an agentic tester locally in under 30 lines of configuration using the open-source Playwright MCP:

// claude_desktop_config.json or MCP Client Settings
{
  "mcpServers": {
    "playwright": {
      "command": "npx",
      "args": ["@modelcontextprotocol/server-playwright"]
    }
  }
}

Once connected, hand the agent a high-level goal prompt:

Goal: Verify Slack Workspace Search Resilience
Target URL: https://staging.app.company.internal/login

Instructions:
1. Log in using test credentials (USER_ENV_KEY).
2. Navigate to the main workspace search bar.
3. Search for the term "invoice_test_412".
4. If an empty state appears, clear the search and filter by "Messages only".
5. Verify that at least one search result renders successfully.
6. If an unhandled error modal appears, take a screenshot and log the exact DOM path.

The agent uses native browser primitives (DOM inspection, clicks, keyboard events) to fulfill the goal adaptively, logging errors without requiring you to maintain fragile CSS selectors.

6. The Strategic Bottom Line #

For engineering leaders and product directors, the takeaway from Slack’s research is crystal clear:

Do not use AI agents to replace your deterministic CI/CD test suite.

Deterministic tests are fast, cheap, and essential for verifying that specific code changes didn’t break known contracts.

The true value of agentic testing is exploratory coverage. It replaces the tedious, manual smoke-testing that QA engineers and developers dread doing before every major production release.

By letting deterministic scripts enforce the journeys and autonomous agents verify the goals, you get the best of both worlds: lightning-fast pull request gating, and a relentless digital QA engineer hunting down edge cases while your team sleeps.

Further Reading & Resources

If you enjoyed this breakdown, subscribe to MLnotes for weekly, bite-sized systems engineering and AI architecture deep-dives. If your QA or dev team is wrestling with flaky test suites, share this article with them.

── more in #ai-agents 4 stories · sorted by recency
── more on @slack 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tests-enforce-journe…] indexed:0 read:8min 2026-09-29 · —