FR EN
Claude Code writes a Playwright test from scratch
Claude Code writes a Playwright test from scratch | AutomationDataCamp
October 7, 2026 ADC Team 6 min read

Watch Claude Code write a Playwright test from scratch (and fix a baseURL trap)

Everyone says AI agents can write end-to-end tests. We wanted to see one do it, start to finish, without editing the run. We gave Claude Code an empty Playwright project and the public TodoMVC demo. It explored the app, wrote three tests, failed all three on the first run, found why, fixed it, then checked for flakiness. Here is the session, step by step, with our review.

Key takeaways
  • It explored before writing: an ARIA snapshot and a list of data-testids, then locators based on them.
  • The first run failed: a baseURL without a trailing slash sent page.goto('/') to a 404 page.
  • It debugged like a tester: it read Playwright’s error-context.md, saw the 404 and fixed the config.
  • It checked flakiness unprompted: --repeat-each=5, 15 runs, 15 passes.
  • Your job is still the test design: it covered exactly what the prompt listed, nothing more.

What did we ask the agent?

We started from an empty Playwright project (Playwright 1.63) with one config line: baseURL: 'https://demo.playwright.dev/todomvc'. Then we gave Claude Code (version 2.1.291, model Claude Opus 5.5) a single prompt, the kind a tester would write:

Explore the app to find stable locators (prefer getByRole, getByPlaceholder or getByTestId over CSS), then write tests/todo.spec.ts covering: adding two todos, completing one and checking the “items left” counter, and filtering with Active / Completed. Run the tests and fix any failure. Do not use waitForTimeout. At the end, explain which locators you chose and why.

We did not edit the run or retry it. The full session has 18 messages and took about two and a half minutes.

Step 1: it explored the app before writing a test

The agent did not guess selectors. It wrote a short Node script that opens the page, adds two todos, then prints the page’s ARIA snapshot and every data-testid it can find. The output told it that the input is a textbox named “What needs to be done?”, that each todo has a checkbox named “Toggle Todo”, and that the app exposes todo-item, todo-title and todo-count test ids.

This is the step to copy, with or without AI: look at what the page actually exposes before choosing a locator.

Step 2: the first run failed, all 3 tests

The first npx playwright test ended with 3 failures, each a 30-second timeout:

Error: locator.fill: Test timeout of 30000ms exceeded.
Call log:
  - waiting for getByPlaceholder('What needs to be done?')

A timeout on the very first locator usually means you are not on the page you think. The agent guessed the cause, then checked it: it opened the error-context.md file that Playwright writes next to a failed test and found a page snapshot with a heading “404” and links to GitHub’s help pages. The tests had landed on a GitHub Pages 404.

Why did page.goto('/') leave the app?

Because Playwright builds the URL with the standard URL() constructor, not by gluing strings. The Playwright docs give the exact case: with baseURL: 'http://localhost:3000/foo' (no trailing slash), navigating to ./bar.html gives http://localhost:3000/bar.html. The last path segment is dropped. And an absolute path like / always goes back to the domain root.

So baseURL: 'https://demo.playwright.dev/todomvc' plus page.goto('/') opened https://demo.playwright.dev/, which does not exist. The fix the agent made:

  • add the trailing slash: baseURL: 'https://demo.playwright.dev/todomvc/';
  • navigate with a relative path: page.goto('./').

It also said plainly that it had changed a file we described as “already set up”. That is what you want from an agent: the change, and a clear note that it touched something outside the task.

Step 3: green, then a flakiness check nobody asked for

After the fix, the 3 tests passed in 2.7 seconds. The agent then ran npx playwright test --repeat-each=5: 15 runs, 15 passes. One clean run proves little; repeating each test is a cheap first check against flaky tests.

The test it wrote

import { test, expect, type Page } from '@playwright/test';

const TODOS = ['Buy milk', 'Walk the dog'];

async function addTodos(page: Page, titles: string[]) {
  const newTodo = page.getByPlaceholder('What needs to be done?');
  for (const title of titles) {
    await newTodo.fill(title);
    await newTodo.press('Enter');
  }
}

test.beforeEach(async ({ page }) => {
  // Relative path: resolves against baseURL's /todomvc/ (an absolute '/' would hit the domain root).
  await page.goto('./');
});

test('adds two todos', async ({ page }) => {
  await addTodos(page, TODOS);

  await expect(page.getByTestId('todo-title')).toHaveText(TODOS);
  await expect(page.getByPlaceholder('What needs to be done?')).toBeEmpty();
  await expect(page.getByTestId('todo-count')).toHaveText('2 items left');
});

test('completing a todo updates the items left counter', async ({ page }) => {
  await addTodos(page, TODOS);

  const first = page.getByTestId('todo-item').filter({ hasText: TODOS[0] });
  await first.getByRole('checkbox', { name: 'Toggle Todo' }).check();

  await expect(first).toHaveClass(/completed/);
  await expect(page.getByTestId('todo-count')).toHaveText('1 item left');
});

test('filters todos with Active and Completed', async ({ page }) => {
  await addTodos(page, TODOS);
  await page
    .getByTestId('todo-item')
    .filter({ hasText: TODOS[0] })
    .getByRole('checkbox', { name: 'Toggle Todo' })
    .check();

  await page.getByRole('link', { name: 'Active' }).click();
  await expect(page).toHaveURL(/#\/active$/);
  await expect(page.getByTestId('todo-title')).toHaveText([TODOS[1]]);

  await page.getByRole('link', { name: 'Completed' }).click();
  await expect(page).toHaveURL(/#\/completed$/);
  await expect(page.getByTestId('todo-title')).toHaveText([TODOS[0]]);

  await page.getByRole('link', { name: 'All' }).click();
  await expect(page.getByTestId('todo-title')).toHaveText(TODOS);
});

Are its locator choices good?

Mostly yes, and they follow the order the Playwright docs recommend: user-facing attributes and explicit contracts first. A few remarks from our review:

  • getByRole('link', { name: 'Completed' }) is a good choice: matching by role avoids the “Clear completed” button, which also contains the word.
  • Picking the todo by its text (filter({ hasText })) before clicking its checkbox makes the test independent of the order of the list.
  • toHaveClass(/completed/) is the only check tied to styling. Asserting the checkbox with toBeChecked() would check the same state without depending on a CSS class name.
  • Missing cases: the prompt did not ask for them, so the agent did not write them: empty input, editing a todo (double-click), “Clear completed”, persistence after reload. An agent covers what you ask; the test design is still yours.

What this session shows about AI-written tests

  • The prompt is the test plan. The agent covered exactly the three cases we listed, no more.
  • Constraints work. “No waitForTimeout” and “prefer getByRole” were respected everywhere.
  • Debugging is where it shines. Reading error-context.md and spotting the 404 took two tool calls. A beginner can lose an hour on that timeout.
  • You still review. Read the diff, including files you did not expect it to touch.

Frequently asked questions

Can Claude Code write Playwright tests on its own?

Yes, in this session it explored the app, wrote three passing tests and fixed a configuration bug without help. It only covered the cases listed in the prompt, so a tester still has to design the test plan and review the code.

Why does page.goto('/') ignore the path in my baseURL?

Playwright resolves URLs with the standard URL() constructor. An absolute path such as '/' always goes back to the domain root, and without a trailing slash the last segment of baseURL is dropped. Use a trailing slash in baseURL and relative paths such as './'.

How do I check that a new Playwright test is not flaky?

A cheap first check is to run it several times in a row, for example npx playwright test --repeat-each=5, and look for any failure.

Recorded on 6 October 2026 with Claude Code 2.1.291 (Claude Opus 5.5) and Playwright 1.63; the session was not edited, only local paths were shortened.

Learn AI-assisted test automation

Our online programme covers Playwright, API testing, CI and AI-assisted testing, with ISTQB exam preparation. Courses are taught in French.

View our courses

Related articles