Introducing e2e: open source agentic testing for web, iOS, and Android TesterArmy CTO Oskar Kwasniewski has open sourced e2e, a TypeScript end-to-end testing framework that lets agent-driven steps and exact assertions coexist in the same test across web, iOS, and Android through one API. The framework adds agent.act(), agent.assert(), and agent.extract() alongside Playwright-style locators and auto-retrying assertions, and caches successful agent actions to a trace cache so subsequent runs replay without model calls. "In practice, this means your token spend follows how often your UI changes rather than how often your tests run," Kwasniewski wrote. Written by Oskar Kwasniewski, CTO at TesterArmy. Originally published on the TesterArmy blog https://tester.army/blog/introducing-e2e . We're open sourcing e2e, a TypeScript testing framework that lets agent steps and exact assertions live in the same test. It runs on web, iOS, and Android with one API, it works with the AI model or subscription you already pay for, and the setup wizard takes most projects from install to a first passing test in about a minute: npx e2e init e2e comes out of the work we do every day at TesterArmy, where our testing agent tests other teams' apps. We built it based on the knowledge and feedback we got from our customers, and it will be the open foundation that powers TesterArmy in the future. Here's what you can expect from this first release: agent.act , agent.assert , and agent.extract where describing a goal is easier than scripting it. Because both live in one test, the report always says exactly what was verified. // Detect dark theme var iframe = document.getElementById 'tweet-2105675143464763540-763' ; if document.body.className.includes 'dark-theme' { iframe.src = "https://platform.twitter.com/embed/Tweet.html?id=2105675143464763540&theme=dark" } Testing so many different apps keeps surfacing the same trade-off. An agent can complete a goal like "buy the cheapest item on this list" without a single selector, but when the test passes, it's hard to say exactly what was checked. A scripted end-to-end test tells you precisely what it verified, and you pay for that precision by rewriting a selector every time someone renames a button or the flow in your app changes. Teams usually pick one approach for the whole suite and live with its costs. e2e lets you make that choice per step instead of per suite. My favorite thing about e2e is that you can gradually adopt agentic APIs where it makes sense. Migration from frameworks like Playwright to e2e is super simple: you port your tests using the same familiar APIs, then add agent steps where they help. On top of that, e2e can run exploration bug bashes via e2e explore before you open a pull request, which makes a great verification step in your software factory. More on that is coming soon. If you've written Playwright tests, most of e2e will feel familiar: role-based locators, auto-retrying assertions, and the same test API. The three agent steps cover the rest. agent.act carries out a goal you describe, agent.assert checks a condition you describe in plain language, and agent.extract reads data off the screen so you can use it later in the test. js import { test, expect } from "e2e"; test "a member upgrades to Pro", async { app, agent, screen } = { await app.open "/settings/billing" ; await agent.act "upgrade the workspace to the Pro plan" ; await expect screen.getByRole "status" .toContainText "Pro" ; await agent.assert "the invoice preview shows the Pro price" ; } ; In this test, the agent works out the upgrade flow on its own, so a reworked billing page is far less likely to break it. The expect line then pins down the result with an exact locator, which means a passing run tells you the status really reads "Pro", whatever path the agent took to get there. Agent steps are slower and more expensive than scripted ones on their first run, because each action needs a model call. We designed e2e so that you pay this cost once per step rather than on every run. When an agent step passes with a recorded check, e2e saves the actions it took to a trace cache. On the next run, it replays those actions directly, without calling the model, so the step runs at the speed of a scripted test and uses no tokens. If the UI changes enough that the replay fails, the runner hands the step back to the live agent to find a new path, and that run costs model calls again. In practice, this means your token spend follows how often your UI changes rather than how often your tests run. e2e drives browsers through Playwright https://playwright.dev , and iOS simulators and Android emulators through agent-device https://github.com/callstack/agent-device . Your tests use the same fixtures, locators, and assertions on every platform, so your web app and mobile app can share one suite and one config: python import type { E2EConfig } from "e2e"; import { web } from "@e2e-dev/web"; import { mobile } from "@e2e-dev/mobile"; import { gateway } from "ai"; export default { targets: { name: "web", engine: web , app: { url: "http://127.0.0.1:3000" }, }, { name: "ios", engine: mobile { platform: "ios" } , app: { bundleId: "com.example.app" }, }, , agents: { default: { model: gateway "openai/gpt-6-luna-fast" }, }, } satisfies E2EConfig; If you'd rather your CI run on infrastructure someone else maintains, e2e has first-class support for hosted browsers from Kernel https://kernel.sh and mobile simulators from Expo https://expo.dev . Your tests stay the same either way, and only the config changes. Agent steps run on any AI SDK https://ai-sdk.dev model. You can bring your own key through Vercel AI Gateway https://vercel.com/ai-gateway or OpenRouter https://openrouter.ai , point e2e at a local model server, or sign in with the ChatGPT, GitHub Copilot, or SuperGrok subscription you already have. You pay your provider's price for tokens, with no added markup. Test credentials stay in environment variables, outside the model's context. The agent can type a password into a login form without the password ever appearing in a prompt, which matters when the model runs on a third-party API. e2e ships an agent skill and an MCP server Model Context Protocol, the standard most coding agents use to call external tools . With them, Claude Code, Cursor, or Codex can open your app, explore it, write tests with locators taken from the real UI, and read the failure report when something breaks: npx skills add tester-army/e2e This is how tests get written in our own repo: the coding agent that built a feature also explores it and adds the regression test in the same change. The two-minute explainer below shows how agent goals and deterministic checks fit together in one test, then sets up e2e from scratch on a Next.js app. It's a quick way to see the whole wizard before you run it on your own project. e2e is available today on npm as e2e , under the Apache 2.0 license. It's still pre-1.0: the core API is the one we use every day, and we expect parts of it to change as more teams run it on their own apps. Web, iOS, and Android are supported now. Desktop apps and other platforms aren't covered yet. Engines are pluggable, so you can write your own and keep the test API unchanged, and we're working on more platforms ourselves. To get started, run this in your project: npx e2e init The docs https://e2e.tester.army/docs cover writing tests, choosing a model, and running in CI, and tester.army/e2e https://tester.army/e2e has the full overview. The code lives at github.com/tester-army/e2e https://github.com/tester-army/e2e . If e2e saves you from rewriting a selector or two, a star helps more people find it, and issues and PRs are very welcome.