Introduction
"Describe a goal in natural language and an agent drives the app to reach it. Check the result with locators and assertions in the same test."
This is article #229 in the "One Open Source Project a Day" series. Today's project is e2e — an end-to-end testing framework for web and mobile apps by TesterArmy, 6,361 Stars, Apache-2.0 license.
End-to-end testing has long been stuck in a tension: hardcoded selectors break the moment the UI is redesigned, but if test logic is too loose, it can't run reliably in CI. e2e's answer is to split every test into three fundamentally different kinds of steps — natural-language agent actions (Goals), model-judged assertions (Assertions), and precise selector-based checks (Locators) — then layer a "replay cache" on top, so verified agent actions can skip model calls entirely on later runs, waking the agent back up only when the app actually changes.
What You'll Learn
- How the three test step types divide responsibility: Goal / Assertion / Locator
- How the replay cache lets AI-driven tests escape the "call the model every single run" cost problem
- The Decision Model executor: answering multiple-choice questions with a small model instead of letting a large one free-wheel
- Quick start: from
npx e2e initto your first test - Migration paths from Playwright, Cypress, Detox, and Maestro
Prerequisites
- Experience writing Playwright, Cypress, or similar end-to-end tests
- Basic understanding of how LLM agents work (instruction → observation → action loop)
- TypeScript fundamentals
Project Background
Overview
e2e is built by TesterArmy — the same company behind an agentic testing platform that runs natural-language tests on web and mobile apps on every pull request or on a schedule. e2e is the open-source core engine behind that capability.
It's not another "let AI generate your test code" tool. It embeds AI-driven action itself into the test runtime: within a single test, which steps should be left to the agent's own judgment, and which must be pinned down with exact selectors, is something the developer explicitly decides.
Author / Team
- Organization: TesterArmy
- Primary language: TypeScript
- License: Apache-2.0
- Created: 2026-07-22
- Status: Active development toward 1.0; APIs and config can still change between minor releases
Project Stats
- ⭐ GitHub Stars: 6,361+
- 🍴 Forks: 285+
- 📄 License: Apache-2.0
- 📅 Created: 2026-07-22
- 🏷️ Topics: e2e, e2e-testing, end-to-end-testing, mobile, mobile-testing, playwright, web
Quick Start
Installation
npx e2e initinit asks for an engine (web or mobile) and a model provider, then writes a config and an example test.
Write a Test
// tests/checkout.e2e.ts
import { test, expect } from 'e2e';
test('a member upgrades to Pro', async ({ app, agent, screen }) => {
await app.open('/settings/billing');
await agent.act('upgrade the workspace to the Pro plan');
await agent.assert('the invoice preview shows a prorated amount');
await expect(screen.getByRole('status')).toContainText('Pro');
});All three step types show up in this one test: agent.act lets the agent figure out how to complete "upgrade to Pro" on its own, agent.assert has the model judge whether the invoice preview is semantically correct, and expect verifies the final state with a precise role-based selector.
Supported Engines
| Package | What it does |
|---|---|
e2e | The SDK, runner, and CLI |
@e2e-dev/web | Browser engine: Chromium, Firefox, WebKit through Playwright |
@e2e-dev/mobile | iOS/Android engine: simulators and emulators through agent-device |
@e2e-dev/github | Reporter that posts results as a pull request comment |
@e2e-dev/kernel | Kernel-hosted cloud browsers |
@e2e-dev/eas | EAS Simulators-hosted iOS simulators and Android emulators |
@e2e-dev/decision | Decision-model executors for bounded semantic actions and assertions, instead of a full LLM |
The project ships ready-made examples for Vite, Next.js, Expo, and SwiftUI.
Core: How the Three Step Types Divide Responsibility
e2e splits every test step into three categories, each with a different model-call policy:
| Step | API | Model call |
|---|---|---|
| Goal | agent.act | Yes, unless replayed from cache |
| Assertion | agent.assert / agent.waitFor / agent.extract | Yes |
| Locator | screen and expect | No |
Goal: Let the Agent Drive the Flow
await agent.act('complete checkout with the test card');
await agent.act('invite {email} as an editor', { params: { email: 'ada@example.test' } });agent.act takes exactly one goal. The agent decides what to click and what to type. The developer controls the order of goals, not the execution details inside each one. The project's own guidance:
- Write one goal per call, using the words the screen actually shows
- Pass test data through
params; for passwords, use aSecret(the model never sees the plaintext) - Wrap values that change on every run — like emails or timestamps — in
unique(), so the replay cache can match them correctly
Assertion: Let the Model Judge Semantics
await agent.assert('the dashboard shows a trial badge');
await agent.waitFor('a download link appears', { timeout: 120_000 });
const data = await agent.extract('every todo title', {
schema: z.object({ titles: z.array(z.string()) }),
});The judging model only sees a text snapshot of the current screen plus the agent's context — not the history or summaries of earlier steps. This isolation is deliberate: it keeps assertions from being swayed by the "narrative" of prior steps. For visual checks (charts, layouts), pass vision: true to let the model see a screenshot.
Locator: No Model Involved
await screen.getByRole('button', 'Sign in').tap();
const row = screen.getByRole('listitem').filter({ hasText: 'Design review' });
await row.getByRole('button', 'Archive').tap();
await expect(screen.getByRole('status')).toHaveText('2 remaining');When you know exactly which control to click and exactly what text to check, use a Locator — zero model calls, zero uncertainty, same writing style as a traditional Playwright test.
The Most Interesting Design: The Replay Cache
This is e2e's core differentiator against "just let an LLM drive the whole test," and the key mechanism for solving the "AI testing is too expensive and too slow" problem.
How It Works
- Once an
agent.act()step is followed by a verification step that confirms the result, the runner records exactly which actions it took (which control, what value) into a JSON entry - On the next run of the same test, the runner replays those actions directly, with no model call at all
- If replay finds a control missing or the page structure changed, the agent takes over from the current screen and finishes the rest using the model
- The run report shows the breakdown:
4 replayed · 1 handed off · 1 missed— which steps were pure replay, which fell back to the agent mid-replay, and which had no cache hit at all
AI 4.1k tokens · 2 model calls · anthropic/claude-sonnet-4.5
Cache 4 replayed · 1 handed off · 1 missedWhat Counts as "Verified"
The cache doesn't record unconditionally. Only when agent.act is immediately followed by a step that actually verifies the outcome — like expect(locator).toHaveText(...) or agent.assert(...) — does a recording get written. Plain-value assertions like expect(value).toBe(...) or read-only calls like agent.extract(...) don't count as verification and won't trigger recording.
This design forces a disciplined "Act, then immediately check" test structure — which happens to be exactly the rigor end-to-end tests should have anyway.
Why a Replay Fails
| Reason | Meaning |
|---|---|
no-entry | No recorded entry matches this step |
wrong-context | The app is on a different screen than when recorded (different origin/path/query) |
target-not-found | A recorded control didn't appear within 15 seconds |
end-mismatch | Every action ran, but the expected final state never showed up |
gap | An action simply can't be replayed (e.g. one that depended on reading pixels) |
In production, --strict-cache turns "cache exists but fails to replay" into a hard failure instead of a silent hand-off — closing off the hidden cost problem of "every CI run is quietly spending model calls and nobody notices."
Why This Design Matters
The biggest pain point of most "AI-driven testing" tools is uncontrolled cost: every CI run makes a real LLM call, and multiply that by dozens of tests times multiple CI triggers per day, and the bill spirals fast. e2e's replay cache essentially separates "the agent discovering a reliable path" from "repeatedly executing that path" — you pay for model calls during discovery, and execution afterward is nearly free, until the app actually changes enough to warrant paying for discovery again. That's a lot smarter than "ask the AI again from scratch every time."
The Decision Model Executor: Multiple-Choice Instead of Free-Form
@e2e-dev/decision offers another way to execute agent.act / agent.assert: instead of a full LLM making free-form decisions, a dedicated Decision Model answers one multiple-choice question per action — "which operation, and which element" — with a probability distribution rather than free text. Only when the chosen operation is "type" does a small text model generate the actual input value.
import { decisionExecutor } from '@e2e-dev/decision';
const executor = decisionExecutor({
model: typeSafeAi.decisionModel('jev-latest'),
textModel: openrouter('inception/mercury-2.5'),
minProbability: 0.9,
minConfidence: 0.8,
});The point of this design: model output never becomes a selector, a URL, or a key — it can only pick an index from a finite set of options the executor built. minProbability and minConfidence let you reject low-confidence actions outright instead of gambling on an uncertain guess. The trade-off is reduced capability — no vision, no multi-select lists, no drag-and-drop — which makes it a good fit for apps with regular, predictable UI structure where you want to run a large test suite more cheaply.
Resources
- 🌟 GitHub: tester-army/e2e
- 📦 npm: e2e
- 📖 Documentation: e2e.tester.army/docs
- 🏢 Made by: TesterArmy
- 💬 Discord: tester.army/discord
- 🔄 Migration guides: from Cypress, Detox, Maestro, Playwright, and Selenium
Summary
e2e represents a clear-eyed judgment about AI testing tools: getting an agent to operate an app isn't the hard part — the hard part is doing it in a way that adapts to UI changes without turning every CI run into a model-call bill.
Three things worth noting:
The explicit split into three step types corrects an overgeneralization in "AI testing." Many AI testing tools try to have one agent handle everything — acting, asserting, verifying, all via model judgment. e2e splits tests into Goal / Assertion / Locator layers, letting developers choose precisely where they need flexibility and where they need determinism, instead of being forced to hand everything to a model or write every selector by hand.
The replay cache separates "discovery cost" from "execution cost," which is the key to running AI-driven tests sustainably in real CI. Without this mechanism, AI-driven end-to-end testing is nearly impossible to scale in production — model-call latency and cost would grow linearly with the number of tests. e2e's "record after verification passes, replay afterward, rediscover only when things change" strategy keeps most runs almost free of model cost.
The Decision Model executor shows that not every AI decision needs a full LLM. Turning action selection into a multiple-choice question for a dedicated Decision Model is cheaper, more controllable, and more auditable than letting a general-purpose LLM free-wheel — worth watching for teams that need to run large test suites at scale.
If your team is looking for an AI-driven end-to-end testing approach that can actually run reliably in real CI — not just impress in a demo — e2e is worth a serious look.
Explore PrimeSkills — a curated marketplace of AI agents and skills, each validated against real enterprise workflows. No hype, just what actually works.
Visit my personal site for more insights and interesting products.