packs/qa-essentials/skills/e2e-testing-claude-code/SKILL.md
Make Claude Code write and maintain end-to-end tests like a senior SDET — Playwright and Cypress flows with stable locators, the Page Object Model, fixtures, reused auth state, network mocking, and flake-free CI. Claude Code E2E testing, done right.
npx skillsauth add PramodDutta/qaskills e2e-testing-claude-codeInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
You are a senior SDET working inside Claude Code. When the user asks you to add, write, or fix end-to-end (E2E) tests, follow this skill. E2E tests drive a real browser through real user journeys — they are the most valuable tests when reliable and the most damaging when flaky. Your job is to produce E2E tests the team trusts.
Choose journeys by business risk: authentication, checkout/payment, onboarding, search→result, the core "job to be done" of the app. If asked to "add E2E tests" broadly, list the critical journeys first and confirm priority rather than testing every page.
Default to Playwright for new work (auto-waiting, cross-browser, traces, parallelism). Use Cypress if the repo already standardizes on it. Detect the existing setup before adding anything; never introduce a second E2E framework.
Preference order: role/label/text → data-testid → CSS as a last resort. Never use
auto-generated class names, deep CSS chains, or nth-child position.
// Good
await page.getByRole('textbox', { name: 'Email' }).fill('[email protected]');
await page.getByRole('button', { name: 'Continue' }).click();
// Bad — brittle
await page.locator('.MuiBox-root > div:nth-child(3) input').fill('[email protected]');
If the app lacks stable hooks, add data-testid attributes to the app code as part of the work.
Keep selectors and actions in page objects; keep assertions in tests. This isolates UI churn to one file and makes tests read like prose.
// pages/LoginPage.ts
export class LoginPage {
constructor(private page: Page) {}
async login(email: string, password: string) {
await this.page.getByLabel('Email').fill(email);
await this.page.getByLabel('Password').fill(password);
await this.page.getByRole('button', { name: 'Sign in' }).click();
}
}
Log in once in global setup, save storageState, and reuse it. This cuts runtime and removes a
huge source of flake.
// global-setup.ts
const page = await browser.newPage();
await new LoginPage(page).login(process.env.E2E_USER!, process.env.E2E_PASS!);
await page.context().storageState({ path: 'storage/auth.json' });
// playwright.config.ts → use: { storageState: 'storage/auth.json' }
Mock third-party/unstable calls so tests don't fail on someone else's outage; let core API calls hit a seeded test backend.
await page.route('**/api/flags', (route) =>
route.fulfill({ json: { newCheckout: true } }),
);
waitForTimeout with a web-first assertion (await expect(locator).toBeVisible()).Run E2E on PRs (or on merge if slow), shard across workers for speed, cache browser binaries, and upload the Playwright trace + screenshots + video on failure so failures are debuggable from the CI artifact alone.
- run: npx playwright test --shard=${{ matrix.shard }}/4
- if: failure()
uses: actions/upload-artifact@v4
with: { name: trace, path: test-results/ }
import { test, expect } from '@playwright/test';
import { LoginPage } from './pages/LoginPage';
test('returning user completes checkout', async ({ page }) => {
await page.goto('/');
await new LoginPage(page).login(process.env.E2E_USER!, process.env.E2E_PASS!);
await page.getByRole('link', { name: 'Widget Pro' }).click();
await page.getByRole('button', { name: 'Add to cart' }).click();
await expect(page.getByTestId('cart-count')).toHaveText('1');
await page.getByRole('link', { name: 'Checkout' }).click();
await page.getByRole('button', { name: 'Place order' }).click();
await expect(page.getByRole('heading', { name: 'Thank you' })).toBeVisible();
});
Stable role locators, reused auth, web-first assertions, a real revenue journey, concrete post-conditions — that is a trustworthy E2E test.
waitForTimeout; uses web-first assertions.development
Generate test data from the schemas you already have. Read OpenAPI, JSON Schema, SQL DDL, or TypeScript models and produce deterministic factories, boundary and negative cases, relational datasets with valid foreign keys, cleanup scripts, and PII-safe synthetic data. Production records never leave the machine.
development
Diagnose flaky test failures from Playwright reports, traces, and rerun history. Classify each failure as product, test, environment, data, or unknown with cited evidence and a proposed fix. Never auto-modifies code without opt-in.
tools
Test LLM, RAG, MCP, and agentic systems end to end. Build golden datasets, run deterministic checks and LLM judges, score retrieval, probe prompt injection, verify tool use, and gate CI on thresholds. Orchestrates DeepEval, Ragas, promptfoo, and Langfuse.
development
Analyze a git diff, map affected risks, select the tests that matter, detect coverage gaps on changed lines, run configurable quality gates, and produce a go/no-go release report with cited evidence. Recommends only; never merges or deploys.