FR EN
Flaky Playwright tests: 8 common causes and their fixes
Flaky Playwright tests: 8 common causes and their fixes | AutomationDataCamp
October 8, 2026 ADC Team 7 min read

Flaky Playwright tests: 8 common causes and how to fix each one

A flaky test fails, then passes on the next run with no code change. Playwright has a name for it: a test that “failed on the first run, but passed when retried” is reported as flaky. Retries hide the problem, they do not fix it. Below are eight common causes, how to spot each one, and the fix the Playwright documentation recommends, including two features from the latest releases: test locks (1.63) and --shuffle (1.64).

Key takeaways
  • Never wait for time: Playwright’s own docs say tests that use waitForTimeout are “inherently flaky”.
  • Await the assertion, not the value: await expect(locator).toBeVisible() retries, expect(await locator.isVisible()) does not.
  • Order-dependent tests can now be found with npx playwright test --shuffle (Playwright 1.64).
  • Tests that share one resource can declare a named lock so they never run at the same time (Playwright 1.63).
  • Measure before you fix: --repeat-each to reproduce, traces on the first retry to understand, --fail-on-flaky-tests to stop new ones.

1. Hard waits

Symptom: the test passes on your laptop and fails on a slower CI runner, or the reverse.

Cause: page.waitForTimeout(2000) guesses how long the app needs. The Playwright API reference marks it as discouraged: “Never wait for timeout in production. Tests that wait for time are inherently flaky.”

Fix: wait for a state, not a duration. Locator actions already wait: before a click, Playwright checks that the element is visible, stable (“the same bounding box for at least two consecutive animation frames”), receives events and is enabled. Then assert the result with a web-first assertion such as await expect(page.getByRole('status')).toHaveText('Saved').

The same goes for waitForLoadState('networkidle'), which the docs also mark as discouraged: “Don't use this method for testing, rely on web assertions to assess readiness instead.”

2. Assertions that do not retry

Symptom: an assertion fails, but the screenshot shows the expected element.

Cause: the value was read once, before the page was ready. The Playwright best practices page gives this exact anti-pattern: expect(await page.getByText('welcome').isVisible()).toBe(true).

Fix: await expect(page.getByText('welcome')).toBeVisible(). Web-first assertions wait “until the expected condition is met”. A quick search for expect(await in your test folder often finds several of these.

3. Locators tied to the DOM structure

Symptom: the test breaks after a harmless markup change, or clicks the wrong one of two matching elements.

Cause: long CSS or XPath selectors. The docs warn: “Your DOM can easily change so having your tests depend on your DOM structure can lead to failing tests.”

Fix: prefer user-facing locators (getByRole, getByLabel, getByTestId). When a hidden copy of an element exists (a mobile menu, a template), Playwright 1.63 added locator.visible(), which “returns a locator that matches only visible elements”, for example page.locator('button').visible().

4. Tests that depend on each other

Symptom: a test passes in the full suite but fails when run alone, or the reverse.

Cause: one test creates data or state that another one silently needs. Playwright already gives each test a fresh browser context (“Playwright creates a context for each test”), so cookies and storage are not shared. But data in your backend is.

Fix: each test creates what it needs, for example through an API call in a fixture. To find the hidden dependencies, Playwright 1.64 added npx playwright test --shuffle, which “schedules tests in a random order, which helps to find tests that accidentally depend on each other”. The run prints a seed: pass it back (--shuffle 271828182 in the docs) to replay the same order while you debug.

5. Parallel tests fighting over one resource

Symptom: failures only appear with several workers, and never on the same test twice.

Cause: by default, Playwright runs test files in parallel in separate worker processes. Two tests that edit the same account setting or the same record overwrite each other.

Fix: first, give each test its own data (unique names, one user per worker). When the resource really is shared, Playwright 1.63 added test locks: “Tests that share a lock name never run concurrently, across files, workers and projects, while everything else keeps running in parallel.” You declare it with test('update user settings', { lock: 'user-settings' }, async ({ page }) => { ... }), or for a whole file with the lock option of test.describe.configure(). This is more precise than switching the whole file to serial mode.

6. Calls to services you do not control

Symptom: failures come in waves, on every branch at once.

Cause: a third-party API, payment sandbox or analytics script is slow or down. Your test is measuring someone else’s uptime.

Fix: the best practices page says to use “the Playwright Network API and guarantee the response needed”: page.route() with route.fulfill() returns a fixed response. Keep a few separate, clearly labelled tests that hit the real service.

7. Time, time zone and locale

Symptom: the test fails around midnight, at the end of the month, or only on the CI machine.

Cause: the app formats “today” or a price with the machine’s clock, time zone and language, and those differ between your laptop and CI.

Fix: set them. In the config, use: { locale: 'en-GB', timezoneId: 'Europe/Paris' } (the example from the emulation docs). For the date itself, the Clock API controls time: page.clock.setFixedTime() fixes Date.now() and new Date(), and page.clock.install() lets you pause or fast-forward timers.

8. A CI machine under too much load

Symptom: unrelated tests time out together, close to the timeout limit, and all pass when re-run alone.

Cause: too many workers for the machine, combined with a tight timeout. When workers is not set, Playwright uses 50% of the logical CPU cores. We hit this on our own website’s test suite: 5 tests failed in the full run, each running for 11 to 18 seconds with a 15-second test timeout. With --workers=1, all 89 tests passed. Nothing was wrong with the site.

Fix: set workers explicitly for CI, and do not lower timeouts to make the suite look faster. Before reporting a bug from a full run, re-run the failing test alone.

How to measure flakiness before fixing it

  • Reproduce: npx playwright test my.spec.ts --repeat-each=20 runs each test 20 times. A test that passes 20 times in a row is not proof, but one failure is.
  • Record evidence: the default config records a trace on-first-retry. For CI failures, the docs recommend “the Playwright trace viewer instead of videos and screenshots”: it shows the DOM, network and console at each step.
  • Read the report: with retries enabled, the report sorts tests into passed, flaky and failed. Track the flaky list like a bug list.
  • Stop new ones: --fail-on-flaky-tests makes the run “fail if any test is flagged as flaky”. Useful on a pull request, once the existing list is under control.

Can AI help?

AI agents can read a trace and suggest a fix, and Playwright now ships a healer agent for broken tests. But most of the causes above are design decisions: what data a test owns, which service is mocked, which resource is shared. An agent can propose a waitForTimeout just as easily as a web-first assertion, so review its fix with the same checklist.

Frequently asked questions

What is a flaky test in Playwright?

Playwright reports a test as flaky when it failed on the first run but passed when retried. With retries set to zero, a flaky test simply shows up as an intermittent failure.

Should I just turn on retries?

Retries keep the pipeline green while you investigate, and they let Playwright record a trace on the first retry. They do not fix the cause. Track the tests reported as flaky and fix them one by one.

How do I find tests that depend on each other?

Since Playwright 1.64, run npx playwright test --shuffle. It runs tests in a random order and prints a seed so you can replay the same order. Also run the failing test alone: if it passes alone and fails in the suite, it depends on another test.

Is waitForTimeout ever acceptable?

Only while debugging locally. The Playwright API reference marks it as discouraged and says tests that wait for time are inherently flaky. Use locator actions and web-first assertions instead.

Written with AI assistance and reviewed by a test automation engineer; every Playwright option and quote was checked against the official documentation on 8 October 2026 (Playwright 1.64).

Learn AI-assisted test automation

Our online programme covers Playwright, API testing, CI and AI-assisted testing, with ISTQB exam preparation. Courses are taught in French.

View our courses

Related articles