TL;DR: Flaky tests are not a minor annoyance anymore, they are a scaling problem. Bitrise's analysis of more than 10 million mobile CI builds found the share of teams reporting real test flakiness climbed from 10 percent in 2022 to 26 percent in 2025, and a widely cited developer survey puts the day to day number even higher: 58 percent of engineers say they deal with a flaky test at least once a month. The fix is not more retries. It is finding which of five mechanical causes is actually behind the flake and closing it. New to the term? Start with what a flaky test actually is before this one.
Quick answers
What actually counts as a flaky test versus a real bug?
A test is flaky when it produces different results, pass one run, fail the next, against the exact same code and the exact same test. A real bug is deterministic: given the same inputs, it fails every time. If re-running the identical test on the identical commit sometimes passes, that is flakiness, not a regression, and the fix belongs in the test's timing or environment, not in application code.
Should a flaky test be fixed immediately or quarantined first?
Quarantine first, fix on a deadline second. Pulling a confirmed-flaky test out of the release-blocking suite the moment it is flagged stops it from costing the whole team time on every build after that. Leaving it in the blocking suite until someone gets around to a proper fix is how teams end up re-running a build three times and calling that normal.
Can AI self-healing locators eliminate flaky tests entirely?
No, and treating them as a full fix is a mistake. Self-healing locators solve exactly one of the five causes below: a selector breaking because the DOM changed. They do nothing for race conditions, network timing, or state bleed between tests. Self-healing closes one door. The other four still need real engineering.
The 5 Root Causes of Flaky Tests
Flaky tests almost always trace back to one of five mechanical causes, and the failure message rarely tells you which one you are looking at. A timeout looks identical whether the real cause is a slow network call or a race condition three steps earlier. These are the five, in the order teams usually run into them.
- DOM race conditions: the test clicks or reads an element before the page has actually finished rendering it. This is the most common cause in UI automation, because the DOM sits in an intermediate state for a few hundred milliseconds, and a test running slightly faster than that render catches it mid-update.
- Network latency: an assertion runs before an API response actually lands, especially over an unstable CI network or a rate-limited third-party endpoint. It passes locally on a fast connection and fails in CI on a slower one, which is exactly the kind of environment-dependent failure that makes flakiness hard to reproduce on a laptop.
- State bleed: one test leaves data behind, a created user, an open session, a row in a shared table, that changes how the next test behaves. These are the hardest to debug, because the test that fails is innocent and the test that ran before it is the actual cause.
- Hardcoded sleep: a fixed page.waitForTimeout(3000) works fine until the CI runner is under load or the page loads slower than usual, at which point 3 seconds was never long enough. Hardcoded sleeps do not fail loudly, they fail intermittently, which makes them one of the least obvious root causes even though they are one of the easiest to fix.
- Dynamic IDs: selectors built on an auto-generated id or a hashed class name, common in frameworks that hash class names for CSS-in-JS, break the moment that hash regenerates on the next build. The test was never actually flaky in the classic sense, the selector was just never stable to begin with.

Replacing Arbitrary waitForTimeout() with Explicit DOM Assertions and Polling
Hardcoded sleep is the easiest of the five causes to fix, and also the thing most engineers reach for the moment a test gets flaky, which is worth covering first. The instinct when a test fails intermittently is to add a wait. That treats the symptom. The actual fix is replacing the wait with a condition.
Playwright's own web-first assertions already retry against the page automatically until the condition is true or a timeout is hit, which is a different thing from waiting a fixed amount of time and hoping. expect(locator).toBeVisible() polls the DOM every few milliseconds until the element is actually visible, then moves on immediately, often faster than a fixed sleep would have allowed, and always more reliable.
- Find every waitForTimeout() in the suite first. Grep the codebase for it before touching anything, so the migration has a real scope instead of a vague sense of somewhere in there.
- Replace each one with the specific condition the sleep was actually waiting for: element visible, network response resolved, text present, item count changed. The sleep duration was always a guess at how long that condition would take. The assertion removes the guess.
- For conditions with no built-in assertion, waiting on a custom application state for example, use expect.poll() or a small retrying helper instead of a fixed wait, so the test still moves on the moment the real condition is met rather than at a fixed clock time.
- Run the migrated tests against a deliberately throttled network a few times before trusting them. A wait-based test that only passed because CI happened to be fast that day is still flaky, just less often.
How AI Self-Healing Locators Dynamically Resolve Broken Selectors at Runtime
Dynamic IDs and shifting selectors are a different problem from timing, and no amount of better waiting fixes a selector that no longer matches anything on the page. This is where self-healing locators earn their place. Instead of one brittle CSS or XPath string, a self-healing locator records several identifying signals for an element at once, its text, its role, its position relative to stable landmarks, its attributes, and scores which combination still uniquely identifies that element when the page changes.
When a build ships and one signal breaks, a moved attribute, a renamed class, the locator does not fail outright. It falls back to the next strongest signal, resolves the same element, and logs that a healing event happened, so a human can confirm it later instead of finding out from a failed pipeline.
For a deeper look at how this actually works under the hood, see our guide to self-healing test automation. It is worth being precise about scope here: this closes exactly one of the five root causes above, dynamic selectors, not the other four. A self-healing locator will happily and correctly resolve an element that a race condition is about to click before it is actually ready. Treat it as one layer, not the whole defense.
Every fix above still takes real engineering time: isolating state, migrating waits, tuning selectors. ContextQA's agentic AI testing platform pairs self-healing locators with the same evaluation layer that runs your functional suite, so a broken selector heals itself before the next run instead of turning into another flaky ticket. See it on a 15-minute demo.
Implementing Automated Flaky Quarantine and Root Cause Telemetry
Root cause fixes take real time, and no team can leave a release blocked on a test someone is actively fixing. This is what quarantine solves, and it is a real practice, not a euphemism for ignoring the problem. Google described its own version of this on its testing blog: track flaky tests separately from the main suite so they cannot fail a build, while still surfacing them for someone to actually fix.
- Detection: track every test's pass or fail history across the last 20 to 50 runs, keyed by commit, not just by test name. A test that fails on the same commit it passed on a run earlier is flaky by definition, no manual judgment needed.
- Auto-quarantine threshold: move a test out of the release-blocking suite automatically once it crosses a flakiness threshold you set, for example three inconsistent results across the last 20 runs, and file a ticket the moment it happens instead of waiting for someone to notice.
- Root cause tagging: tag every quarantined test with which of the five causes above triggered it, timeout, network, state, selector, race condition. Without this tag, quarantine becomes a queue nobody works through. With it, an engineer can batch-fix every hardcoded-sleep test in one pass instead of debugging each one from scratch.
- Re-admission: set a real exit condition. A quarantined test earns its way back into the blocking suite only after a set number of consecutive clean runs, not the moment someone believes they fixed it.

Frequently Asked Questions
How many failures should it take before a test gets flagged flaky?
There is no universal number, but most teams settle on a version of this: if a test produces two different results, pass and fail, across the same commit within a rolling window of 10 to 20 runs, flag it. Fewer runs than that and you are reacting to noise. More, and a genuinely flaky test sits unflagged for too long.
Does quarantining flaky tests just hide real bugs?
It can, if quarantine has no exit condition and no owner. A quarantined test that never gets root-caused is a bug that got a permanent pass, not a temporary one. The practice only works when every quarantined test has a ticket, a tag, and a deadline attached to it.
What is a healthy flaky test rate to aim for?
Zero is not realistic at any meaningful scale, timing and network variance exist in every real system. Treat a rising trend as the actual signal to watch, not an absolute number: a suite holding steady at 1 to 2 percent flaky runs is very different from one climbing toward the double-digit rates the Bitrise data above shows becoming common.
Bottom line
None of the five causes above are exotic. Race conditions, network timing, state bleed, hardcoded sleeps, and unstable selectors are the same five things breaking test suites at companies of every size, which is exactly why a structured playbook works better than fixing flaky tests one panicked ticket at a time. Quarantine buys the time to fix them properly instead of re-running until green. And the reason any of this is worth the effort is not abstract: a flaky test waved through as probably fine, just re-run it, is exactly the kind of gap that lets a real defect slip past a merge, which our breakdown of the cost of defects in software testing lays out in real numbers.