TesterArmy's E2E framework solves the two hard problems of agent-driven testing: cost and flakiness. You write a test that describes a goal in natural language. An agent drives the browser or mobile app to reach it. On the first run, the framework records every action the agent takes. On subsequent runs, it replays those actions without calling the model until the app changes enough to invalidate the recording.
This is not snapshot testing. It is not record-and-playback in the Selenium sense. It is a cache layer for agent decisions, and it changes the economics of agentic QA.
Agent-driven tests are expensive. Every test run that calls a model costs tokens. A suite of 50 tests might burn through tens of thousands of tokens per run. If you run that suite on every pull request, you are paying for the same decisions over and over.
Agent-driven tests are also non-deterministic. The same prompt can produce different action sequences. A test that passes today might fail tomorrow because the model chose a different button to click, even if the app did not change.
TesterArmy's replay cache solves both problems. After the first run, the framework stores the action sequence. On the next run, it replays those actions without a model call. If the app has not changed, the test is deterministic and free. If the app has changed, the framework detects the mismatch and re-runs the agent step.
The framework does not store DOM snapshots. It stores action sequences and the locators that triggered them. When a test runs, the framework checks whether each locator still resolves to the same element. If it does, the cached action is safe to replay. If it does not, the framework invalidates the cache and calls the agent again.
This is a heuristic. It assumes that if the locator still works, the app has not changed in a way that matters. That assumption breaks in two cases:
The framework does not try to detect these cases automatically. Instead, it relies on assertions. If an assertion fails after a replayed action, the test fails, and you know the cache is stale.
The replay cache stores three things:
Each agent step is a replay boundary. If you write:
await agent.act('log in as a member');
await agent.act('upgrade to Pro');
The framework treats these as two separate cached sequences. If the login flow has not changed but the upgrade flow has, only the second step re-runs with the model.
Tests without agent steps never call the model. You can mix agent-driven exploration with traditional locator-based assertions:
await agent.act('navigate to the billing page');
await expect(screen.getByRole('heading')).toContainText('Billing');
await screen.getByRole('button', { name: 'Upgrade' }).click();
The first line uses the agent and the cache. The rest is Playwright.
Dynamic content breaks naive replay. If your app shows a timestamp, a random ID, or a live data feed, the cached locator will not match on the next run.
TesterArmy does not solve this with fuzzy matching. It solves it by making you choose stable locators. The framework is built on Playwright, so you use Playwright's locator strategies: roles, labels, test IDs. If you rely on text content that changes, the cache invalidates.
This is a feature, not a bug. It forces you to write tests that depend on semantic structure, not incidental content. If your test breaks because a timestamp changed, your locator was too brittle.
For truly non-deterministic elements (a random coupon code, a generated UUID), you skip the agent step and write the assertion manually:
await agent.act('apply a coupon code');
const code = await screen.getByTestId('coupon-code').textContent();
expect(code).toMatch(/^[A-Z0-9]{8}$/);
The agent gets you to the state. The assertion verifies the shape, not the exact value.
If an assertion fails mid-replay, the framework does not automatically re-run the agent step. It fails the test. This is deliberate. A failed assertion means either:
You decide which. If the app changed, you update the test. If the cache is stale, you delete the cache file and re-run. The framework stores cache files in .e2e/cache by default. You can commit them to version control or add them to .gitignore.
Committing cache files makes tests deterministic across machines. Ignoring them makes tests adapt to local changes faster. Most teams commit them.
The framework is a Playwright wrapper with two additions:
The orchestration layer uses a 426-token prompt system. The prompt includes:
agent.act().
The agent returns a sequence like:
[
{ "action": "click", "locator": "role=button[name='Upgrade']" },
{ "action": "wait", "locator": "role=dialog" },
{ "action": "click", "locator": "role=button[name='Confirm']" }
]
The framework executes this sequence and stores it. On the next run, it checks whether each locator still resolves. If yes, it replays. If no, it calls the agent again.
The cache store is a JSON file per test. The file maps agent step descriptions to action sequences:
{
"upgrade the workspace to the Pro plan": {
"actions": [],
"locators": [],
"timestamp": "2026-10-05T18:32:11Z"
}
}
The timestamp is metadata. The framework does not use it for invalidation.
| Dimension | Replay Cache | Always-On Agent | Traditional E2E |
|---|---|---|---|
| Cost per run | Free after first run | High (model call every run) | Free |
| Determinism | High (until app changes) | Low (model variance) | High |
| Maintenance | Medium (cache invalidation) | Low (agent adapts) | High (brittle locators) |
| Failure signal | Clear (assertion or cache miss) | Noisy (agent flake) | Clear |
| Setup complexity | Medium (model + Playwright) | Medium (model + Playwright) | Low (Playwright only) |
The replay cache trades maintenance cost for runtime cost. You pay once to record, then replay for free. But you must manage cache invalidation. If you change the app and forget to invalidate the cache, tests pass when they should fail.
The failure modes:
import { test, expect } from 'e2e';
test('member upgrades and sees prorated invoice', async ({ app, agent, screen }) => {
// Agent-driven navigation (cached after first run)
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
// Agent-driven assertion (cached)
await agent.assert('the invoice preview shows a prorated amount');
// Traditional locator assertion (no model call, no cache)
await expect(screen.getByRole('status')).toContainText('Pro');
// Traditional interaction (no model call, no cache)
const invoiceTotal = await screen.getByTestId('invoice-total').textContent();
expect(parseFloat(invoiceTotal)).toBeGreaterThan(0);
});
The agent handles the upgrade flow. The locators verify the result. The cache stores the upgrade sequence. On the next run, the agent steps replay without a model call. The locator steps run as normal.
Use agent steps when:
Use locators when:
Use agent assertions when:
Use locator assertions when:
The framework runs in any environment that supports Playwright. For CI:
npm install e2e playwright. npx e2e test.
The cache files live in .e2e/cache. If you commit them, tests replay in CI without model calls. If you ignore them, CI records on the first run and replays after that.
For cost control, you can set a cache TTL. The framework invalidates caches older than N days, forcing a re-record. This is useful if you want to periodically verify that the agent still makes good decisions.
The framework logs:
Logs go to stdout by default. You can send them to a structured log store (Datadog, Honeycomb) by configuring a custom reporter.
The most useful signal is cache hit rate. If your cache hit rate drops below 80%, your app is changing too fast for replay to be useful, or your locators are too brittle.
Use TesterArmy's E2E framework when you have complex, stable user flows where the cost of repeated model calls outweighs cache maintenance overhead. It shines for onboarding sequences, checkout flows, and admin workflows that change monthly, not daily. The replay cache cuts testing costs by 90% in these scenarios, and the Playwright foundation means you can drop down to traditional locators whenever the agent is overkill.
Avoid it when your UI changes multiple times per day, making cache invalidation a constant chore. Skip it if your app is dominated by live data feeds or random content that breaks locator stability. And if you need sub-100ms test execution, the locator resolution overhead during replay (even without model calls) will frustrate you. In those cases, stick with pure Playwright or accept the cost of always-on agent testing.
The framework works best when you treat agent steps as expensive, reusable building blocks and locator steps as cheap, precise verification. If you find yourself caching every interaction, you are using the wrong tool. If you find yourself writing locators for complex flows, you are missing the point.