{"slug": "agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the", "title": "Agent Replay Caching for E2E Tests: How TesterArmy Records Once, Replays Until the App Changes", "summary": "TesterArmy's E2E testing framework records the action sequences an agent takes on a first run and replays them on subsequent runs without calling the model, cutting token costs and eliminating non-determinism until the app changes. The cache stores action sequences and their locators rather than DOM snapshots, invalidating and re-running the agent step when a locator no longer resolves to the same element, with each agent.act() call acting as a separate replay boundary. The framework relies on assertions rather than fuzzy matching to catch stale caches, and requires stable Playwright locators such as roles, labels and test IDs.", "body_md": "TesterArmy's E2E framework solves the two hard problems of agent-driven testing: cost and flakiness. You write a test that describes a goal in natural language. An agent drives the browser or mobile app to reach it. On the first run, the framework records every action the agent takes. On subsequent runs, it replays those actions without calling the model until the app changes enough to invalidate the recording.\n\nThis is not snapshot testing. It is not record-and-playback in the Selenium sense. It is a cache layer for agent decisions, and it changes the economics of agentic QA.\n\nAgent-driven tests are expensive. Every test run that calls a model costs tokens. A suite of 50 tests might burn through tens of thousands of tokens per run. If you run that suite on every pull request, you are paying for the same decisions over and over.\n\nAgent-driven tests are also non-deterministic. The same prompt can produce different action sequences. A test that passes today might fail tomorrow because the model chose a different button to click, even if the app did not change.\n\nTesterArmy's replay cache solves both problems. After the first run, the framework stores the action sequence. On the next run, it replays those actions without a model call. If the app has not changed, the test is deterministic and free. If the app has changed, the framework detects the mismatch and re-runs the agent step.\n\nThe framework does not store DOM snapshots. It stores action sequences and the locators that triggered them. When a test runs, the framework checks whether each locator still resolves to the same element. If it does, the cached action is safe to replay. If it does not, the framework invalidates the cache and calls the agent again.\n\nThis is a heuristic. It assumes that if the locator still works, the app has not changed in a way that matters. That assumption breaks in two cases:\n\nThe framework does not try to detect these cases automatically. Instead, it relies on assertions. If an assertion fails after a replayed action, the test fails, and you know the cache is stale.\n\nThe replay cache stores three things:\n\nEach agent step is a replay boundary. If you write:\n\n```\nawait agent.act('log in as a member');\nawait agent.act('upgrade to Pro');\n```\n\nThe framework treats these as two separate cached sequences. If the login flow has not changed but the upgrade flow has, only the second step re-runs with the model.\n\nTests without agent steps never call the model. You can mix agent-driven exploration with traditional locator-based assertions:\n\n```\nawait agent.act('navigate to the billing page');\nawait expect(screen.getByRole('heading')).toContainText('Billing');\nawait screen.getByRole('button', { name: 'Upgrade' }).click();\n```\n\nThe first line uses the agent and the cache. The rest is Playwright.\n\nDynamic content breaks naive replay. If your app shows a timestamp, a random ID, or a live data feed, the cached locator will not match on the next run.\n\nTesterArmy does not solve this with fuzzy matching. It solves it by making you choose stable locators. The framework is built on Playwright, so you use Playwright's locator strategies: roles, labels, test IDs. If you rely on text content that changes, the cache invalidates.\n\nThis is a feature, not a bug. It forces you to write tests that depend on semantic structure, not incidental content. If your test breaks because a timestamp changed, your locator was too brittle.\n\nFor truly non-deterministic elements (a random coupon code, a generated UUID), you skip the agent step and write the assertion manually:\n\n``` js\nawait agent.act('apply a coupon code');\nconst code = await screen.getByTestId('coupon-code').textContent();\nexpect(code).toMatch(/^[A-Z0-9]{8}$/);\n```\n\nThe agent gets you to the state. The assertion verifies the shape, not the exact value.\n\nIf an assertion fails mid-replay, the framework does not automatically re-run the agent step. It fails the test. This is deliberate. A failed assertion means either:\n\nYou decide which. If the app changed, you update the test. If the cache is stale, you delete the cache file and re-run. The framework stores cache files in `.e2e/cache` by default. You can commit them to version control or add them to `.gitignore`.\n\nCommitting cache files makes tests deterministic across machines. Ignoring them makes tests adapt to local changes faster. Most teams commit them.\n\nThe framework is a Playwright wrapper with two additions:\n\nThe orchestration layer uses a 426-token prompt system. The prompt includes:\n\n`agent.act()`.\nThe agent returns a sequence like:\n\n```\n[\n  { \"action\": \"click\", \"locator\": \"role=button[name='Upgrade']\" },\n  { \"action\": \"wait\", \"locator\": \"role=dialog\" },\n  { \"action\": \"click\", \"locator\": \"role=button[name='Confirm']\" }\n]\n```\n\nThe framework executes this sequence and stores it. On the next run, it checks whether each locator still resolves. If yes, it replays. If no, it calls the agent again.\n\nThe cache store is a JSON file per test. The file maps agent step descriptions to action sequences:\n\n```\n{\n  \"upgrade the workspace to the Pro plan\": {\n    \"actions\": [],\n    \"locators\": [],\n    \"timestamp\": \"2026-10-05T18:32:11Z\"\n  }\n}\n```\n\nThe timestamp is metadata. The framework does not use it for invalidation.\n\n| Dimension | Replay Cache | Always-On Agent | Traditional E2E | \n|---|---|---|---|\n| **Cost per run** | Free after first run | High (model call every run) | Free | \n| **Determinism** | High (until app changes) | Low (model variance) | High | \n| **Maintenance** | Medium (cache invalidation) | Low (agent adapts) | High (brittle locators) | \n| **Failure signal** | Clear (assertion or cache miss) | Noisy (agent flake) | Clear | \n| **Setup complexity** | Medium (model + Playwright) | Medium (model + Playwright) | Low (Playwright only) | \n\nThe replay cache trades maintenance cost for runtime cost. You pay once to record, then replay for free. But you must manage cache invalidation. If you change the app and forget to invalidate the cache, tests pass when they should fail.\n\nThe failure modes:\n\n``` js\nimport { test, expect } from 'e2e';\n\ntest('member upgrades and sees prorated invoice', async ({ app, agent, screen }) => {\n  // Agent-driven navigation (cached after first run)\n  await app.open('/settings/billing');\n  await agent.act('upgrade the workspace to the Pro plan');\n\n  // Agent-driven assertion (cached)\n  await agent.assert('the invoice preview shows a prorated amount');\n\n  // Traditional locator assertion (no model call, no cache)\n  await expect(screen.getByRole('status')).toContainText('Pro');\n\n  // Traditional interaction (no model call, no cache)\n  const invoiceTotal = await screen.getByTestId('invoice-total').textContent();\n  expect(parseFloat(invoiceTotal)).toBeGreaterThan(0);\n});\n```\n\nThe agent handles the upgrade flow. The locators verify the result. The cache stores the upgrade sequence. On the next run, the agent steps replay without a model call. The locator steps run as normal.\n\nUse agent steps when:\n\nUse locators when:\n\nUse agent assertions when:\n\nUse locator assertions when:\n\nThe framework runs in any environment that supports Playwright. For CI:\n\n`npm install e2e playwright`.` npx e2e test`.\nThe cache files live in `.e2e/cache`. If you commit them, tests replay in CI without model calls. If you ignore them, CI records on the first run and replays after that.\n\nFor cost control, you can set a cache TTL. The framework invalidates caches older than N days, forcing a re-record. This is useful if you want to periodically verify that the agent still makes good decisions.\n\nThe framework logs:\n\nLogs go to stdout by default. You can send them to a structured log store (Datadog, Honeycomb) by configuring a custom reporter.\n\nThe most useful signal is cache hit rate. If your cache hit rate drops below 80%, your app is changing too fast for replay to be useful, or your locators are too brittle.\n\nUse TesterArmy's E2E framework when you have complex, stable user flows where the cost of repeated model calls outweighs cache maintenance overhead. It shines for onboarding sequences, checkout flows, and admin workflows that change monthly, not daily. The replay cache cuts testing costs by 90% in these scenarios, and the Playwright foundation means you can drop down to traditional locators whenever the agent is overkill.\n\nAvoid it when your UI changes multiple times per day, making cache invalidation a constant chore. Skip it if your app is dominated by live data feeds or random content that breaks locator stability. And if you need sub-100ms test execution, the locator resolution overhead during replay (even without model calls) will frustrate you. In those cases, stick with pure Playwright or accept the cost of always-on agent testing.\n\nThe framework works best when you treat agent steps as expensive, reusable building blocks and locator steps as cheap, precise verification. If you find yourself caching every interaction, you are using the wrong tool. If you find yourself writing locators for complex flows, you are missing the point.", "url": "https://wpnews.pro/news/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the", "canonical_source": "https://dev.to/mech_app_ai/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the-app-changes-5hi", "published_at": "2026-10-05 20:07:15+00:00", "updated_at": "2026-10-05 20:18:08.422327+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-tools", "mlops"], "entities": ["TesterArmy", "Playwright"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the", "markdown": "https://wpnews.pro/news/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the.md", "text": "https://wpnews.pro/news/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the.txt", "jsonld": "https://wpnews.pro/news/agent-replay-caching-for-e2e-tests-how-testerarmy-records-once-replays-until-the.jsonld"}}