Playwright Tutorial: End-to-End Tests That Aren't Flaky
Playwright drives real browsers to test your app the way users use it. Install it, write your first test, use role-based locators and auto-waiting assertions, log in once and reuse the session, run against a dev server, debug with UI mode and traces, and run it in CI.
Unit tests tell you functions work. End-to-end tests tell you the app works: a real browser loads your pages, clicks your buttons and checks what appears. (Unit vs integration vs E2E tests) Playwright is the most popular tool for this today — fast, reliable, and it drives Chromium, Firefox and WebKit (Safari's engine).
Install
In your project:
npm init playwright@latest
It asks a few questions, creates playwright.config.ts, an example test in tests/, and downloads browsers. Run the example:
npx playwright test
npx playwright show-report
Your first real test
// tests/signup.spec.ts
import { test, expect } from '@playwright/test'
test('visitor can sign up and reach the dashboard', async ({ page }) => {
await page.goto('/signup')
await page.getByLabel('Email').fill(`test+${Date.now()}@example.com`)
await page.getByLabel('Password').fill('a-long-test-password')
await page.getByRole('button', { name: 'Create account' }).click()
await expect(page).toHaveURL(/\/dashboard/)
await expect(page.getByRole('heading', { name: 'Welcome' })).toBeVisible()
})
Read it like a user story. That's the goal.
Locators: find things the way users do
Prefer locators based on what a user sees and what assistive technology reads:
| Locator | Finds |
|---|---|
getByRole('button', { name: 'Save' }) |
A button labelled Save — the best default |
getByLabel('Email') |
The input with that label |
getByText('Order confirmed') |
Visible text |
getByPlaceholder('Search') |
Input by placeholder |
getByTestId('cart-total') |
data-testid="cart-total" — when nothing else fits |
Avoid CSS selectors like .btn-primary > span:nth-child(2) — they break whenever the markup changes. Role-based locators also nudge your app towards being accessible. (Web accessibility basics)
Auto-waiting: why Playwright isn't flaky (if you let it)
Actions like click() wait for the element to be visible, enabled and stable. expect(...) assertions retry until they pass or time out:
await expect(page.getByText('Saved')).toBeVisible() // waits for it to appear
So:
- Never use
page.waitForTimeout(3000). Fixed sleeps are the main cause of flaky tests — too short on a slow CI machine, wasted time everywhere else. - Use web-first assertions (
await expect(locator).toHaveText(...)) rather than reading a value and comparing it yourself.
Run against your dev server automatically
In playwright.config.ts:
export default defineConfig({
use: { baseURL: 'http://localhost:3000', trace: 'on-first-retry' },
webServer: {
command: 'npm run dev',
url: 'http://localhost:3000',
reuseExistingServer: !process.env.CI,
},
})
Playwright starts the app before tests and stops it after.
Log in once, reuse it
Logging in through the UI in every test is slow. Do it once in a setup project and save the session:
// tests/auth.setup.ts
import { test as setup } from '@playwright/test'
setup('authenticate', async ({ page }) => {
await page.goto('/login')
await page.getByLabel('Email').fill(process.env.E2E_USER!)
await page.getByLabel('Password').fill(process.env.E2E_PASSWORD!)
await page.getByRole('button', { name: 'Log in' }).click()
await page.waitForURL('/dashboard')
await page.context().storageState({ path: 'playwright/.auth/user.json' })
})
// playwright.config.ts
projects: [
{ name: 'setup', testMatch: /.*\.setup\.ts/ },
{
name: 'chromium',
use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
dependencies: ['setup'],
},
]
Add playwright/.auth to .gitignore. Use a dedicated test account against a test database, never real users. (Database seeding)
Debugging failing tests
- UI mode:
npx playwright test --ui— watch tests run step by step, time-travel through each action, see the DOM at every point. The best way to understand a failure. - Trace viewer: with
trace: 'on-first-retry', failed CI runs produce a trace with screenshots, network requests and console logs. Open it withnpx playwright show-trace. - Codegen:
npx playwright codegen localhost:3000records your clicks as test code — a good starting point; clean up the locators afterwards. await page.pause()stops a test and opens the inspector.
Keeping tests reliable
- Independent tests. Each test sets up its own data; never rely on another test having run first.
- Unique data. Use timestamps or random suffixes for emails and names so parallel runs don't collide.
- Control the network when needed.
page.route()can stub third-party APIs (payments, AI calls) so tests don't depend on them. - Test what matters. A handful of critical flows — sign up, log in, the core action, checkout — beats hundreds of fragile UI checks.
In CI
- run: npm ci
- run: npx playwright install --with-deps chromium
- run: npx playwright test
- uses: actions/upload-artifact@v7
if: ${{ !cancelled() }}
with:
name: playwright-report
path: playwright-report/
Upload the report so you can open traces from failed runs. (CI with GitHub Actions)
Playwright and AI agents
E2E tests are a superb "definition of done" for coding agents: "the feature is done when npx playwright test tests/checkout.spec.ts passes." Agents can also use Playwright directly to look at the app they're building. Ask them to write tests with role-based locators and no fixed waits. (Getting AI to write tests that catch bugs)
The summary
npm init playwright@latest, then write tests that read like user stories.- Use
getByRole/getByLabellocators and web-firstexpectassertions — no sleeps. - Start the dev server via
webServer; log in once withstorageState. - Debug with UI mode and traces; run in CI and upload the report.
EasySpawn gives Claude Code a persistent server with your app, database and a real browser available, so it can run Playwright suites against the live app on every change. See how it works or join the waitlist.
Related: Unit vs Integration vs E2E Tests · How to Test Your App Before Launch · Getting AI to Write Tests · Set Up CI With GitHub Actions
Keep reading
How to Test Webhooks Locally: Stripe CLI, Tunnels, and Replays
Webhook providers can't reach localhost, so local testing needs a forwarder or a tunnel. How to use provider CLIs like stripe listen, general tunnels like ngrok and Cloudflare Tunnel, request-capture tools, and fixture replays in automated tests — plus the signature-verification gotchas.
ESM vs CommonJS: import vs require and the Errors Between Them
JavaScript has two module systems. CommonJS uses require and module.exports; ES modules use import and export. How Node decides which a file is, and how to fix 'Cannot use import statement outside a module', 'require is not defined', ERR_REQUIRE_ESM and __dirname is not defined.