
Flaky tests erode confidence in CI pipelines, causing developers to repeatedly re‑run builds and miss real failures. This blog walks through six concrete actions you can take to detect, isolate, and eliminate flakiness, from measuring instability to quarantining problematic tests.
Measure Before You Fix
Flaky tests are tests whose outcomes change between runs without any code modification. Because they mask genuine regressions, the first step in remediation is to make the flakiness observable and quantifiable.
Conceptual workflow
- Execute the full test suite repeatedly on the same commit.
- Persist each run’s JUnit XML report.
- Parse the reports to identify tests that have both passed and failed across the runs.
- Calculate a failure‑rate for each flaky test (failed runs ÷ total runs).
- Rank the tests by descending failure‑rate to surface the worst offenders.
Practical implementation
Below is a minimal Bash loop that runs pytest ten times and stores the JUnit output in a reports/ directory. The || true ensures the loop continues even when a run exits with a non‑zero status.
# Run the suite 10 times and keep each JUnit report
for i in $(seq 1 10); do
pytest tests/ --junitxml=reports/run-${i}.xml || true
done
After collection, a short Python script can aggregate the results:
import xml.etree.ElementTree as ET
from collections import defaultdict
failure_counts = defaultdict(int)
total_counts = defaultdict(int)
for file in glob.glob('reports/run-*.xml'):
tree = ET.parse(file)
for case in tree.findall('.//testcase'):
name = f"{case.get('classname')}.{case.get('name')}"
total_counts[name] += 1
if case.find('failure') is not None:
failure_counts[name] += 1
flaky = {name: failure_counts[name] / total_counts[name]
for name in total_counts
if 0 < failure_counts[name] < total_counts[name]}
# Rank by failure rate
for name, rate in sorted(flaky.items(), key=lambda x: x[1], reverse=True):
print(f"{name}: {rate:.2%}")
The output lists each flaky test with its failure percentage, allowing engineers to prioritize the highest‑rate items first.
Key considerations when measuring flakiness
- Number of repetitions: Ten runs provide a balance between detection confidence and resource consumption; increase the count for large test suites or when failure rates are low.
- Isolation of environment: Ensure the same commit is built on identical infrastructure (same Docker image, same environment variables) to avoid conflating environmental variance with test instability.
- Automation support: Many CI platforms (GitHub Actions, GitLab CI, Azure Pipelines) can archive JUnit artifacts automatically, simplifying the collection step.
- Continuous tracking: Store historical failure‑rate data to detect trends; a test that gradually becomes more flaky may indicate a degrading dependency.
By systematically measuring, ranking, and addressing flaky tests, teams can restore confidence in their CI pipeline and reduce the “re‑run and hope” behavior that erodes software quality.
Replace Fixed Sleeps with Explicit Waits
Hard‑coded sleep() calls introduce a deterministic pause that assumes a specific system state will be reached within the given interval. In a continuous‑integration (CI) environment, resource contention, variable network latency, and container start‑up times often exceed that assumption, causing the test to fail intermittently. This nondeterministic failure pattern is a classic source of flaky tests, as described in community reports on flaky‑test mitigation.
Instead of waiting a fixed duration, tests should wait for an explicit condition that signals readiness. An explicit wait monitors the target state and proceeds as soon as the condition is satisfied, or fails after a configurable timeout. This approach reduces idle time on fast runs while providing resilience on slower CI runners.
- Condition‑driven waiting – poll for a UI element’s visibility, clickability, or an API health endpoint response.
- Timeout control – define a maximum wait (e.g., 10 seconds) that bounds test duration.
- Immediate continuation – the test proceeds as soon as the condition is met, improving overall pipeline throughput.
For Selenium‑based UI tests, replace a static pause with WebDriverWait and an appropriate expected condition:
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
# Flaky: fixed 2‑second sleep
# time.sleep(2)
# button = driver.find_element(By.ID, "submit")
# Stable: explicit wait up to 10 seconds
button = WebDriverWait(driver, 10).until(
EC.element_to_be_clickable((By.ID, "submit"))
)
For service‑level integration tests, poll a health endpoint until it returns a successful status code:
import time
import requests
def wait_for_service(url, timeout=15, interval=1):
end_time = time.time() + timeout
while time.time() < end_time:
try:
if requests.get(url).status_code == 200:
return True
except requests.RequestException:
pass
time.sleep(interval)
raise TimeoutError(f"Service at {url} not healthy after {timeout}s")
Adopting explicit waits eliminates the arbitrary timing assumptions that cause flakiness on busy CI runners, leading to more deterministic test outcomes and faster feedback cycles.
Make Tests Independent of Each Other
When a test passes only after another test has executed, the suite is relying on hidden shared state. This state can be introduced unintentionally through resources that persist beyond the lifetime of a single test case.
- Database rows or user accounts that are created by one test and read or modified by another.
- Global variables, singletons, or module‑level caches that retain values between runs.
- Files in a common temporary directory that are not removed after the test finishes.
- External services or network sockets that remain bound after a test completes.
These dependencies become especially problematic when tests are executed in parallel, shuffled, or retried in a continuous‑integration (CI) pipeline. A test that assumes a particular row exists may fail if the suite is reordered, and a global flag left true by a previous test can cause spurious failures in unrelated cases.
Isolating state with fixtures
In pytest, fixtures provide a deterministic way to set up and tear down resources for each test. By yielding a resource and performing cleanup after the yield, the fixture guarantees that no residue leaks into subsequent tests.
import uuid
import pytest
@pytest.fixture
def isolated_customer(api_client):
# Create a unique customer for the test
customer = api_client.create_customer(name=f"test-{uuid.uuid4()}")
yield customer
# Ensure the customer is removed regardless of test outcome
api_client.delete_customer(customer.id)
Tests that need a customer simply request the fixture:
def test_customer_can_place_order(isolated_customer, api_client):
order = api_client.create_order(customer_id=isolated_customer.id)
assert order.status == "pending"
Because each invocation receives a fresh record, the test is independent of any other test that also uses isolated_customer. The same pattern applies to temporary files, in‑memory caches, or mock servers.
Detecting hidden dependencies
- Run the suite with the
pytest‑randomlyplugin to shuffle test order; failures indicate order‑sensitive state. - Execute tests in parallel (e.g.,
pytest -n auto) and observe any nondeterministic behavior. - Inspect CI logs for intermittent failures that correlate with resource contention.
By consistently using fixtures for data creation and cleanup, and by regularly exercising the suite in random or parallel modes, engineers can eliminate hidden shared state, reduce flakiness, and maintain reliable, maintainable test pipelines.
Control Time, Randomness, and External Services
Automated tests that rely on the system clock, random number generators, or third‑party APIs are a common source of nondeterminism. When a test asserts “created today” it can fail at midnight, a seeded Random that is not reproducible can produce different outcomes on each run, and a live external service may be slow or unavailable, causing intermittent failures. Stabilizing these dependencies requires isolating the test from the real source of variability while preserving a minimal set of true integration checks.
Freezing or injecting time
- Replace direct calls to
System.currentTimeMillis(),Instant.now(), or language‑specific time APIs with an abstractClockinterface. - In unit tests, provide a
FixedClockthat returns a constant timestamp, e.g.:Clock fixed = Clock.fixed(Instant.parse("2023-01-01T00:00:00Z"), ZoneOffset.UTC); service.setClock(fixed); - Reserve a separate integration stage that runs a small suite against the real clock to verify time‑sensitive logic such as expiration handling.
Seeding randomness
- Expose the random generator through a constructor or setter so tests can inject a deterministic instance.
- Use a fixed seed:
Random rng = new Random(12345L); service.setRandom(rng); - Document the seed value in the test file to make the intent explicit and avoid accidental changes.
Mocking external services
- Wrap API calls in a client interface and provide a mock implementation that returns canned responses.
- Frameworks such as
WireMockorMSWcan simulate HTTP endpoints with configurable latency and error codes, allowing tests to verify retry and fallback logic without contacting the real provider. - Maintain a limited “real‑service” test set that runs in a dedicated pipeline stage, exercising the actual third‑party sandbox under controlled conditions.
By applying these three controls—clock injection, deterministic randomness, and service mocking—engineers can eliminate the majority of flaky behavior while still satisfying compliance requirements (e.g., NIST or ISO 27001) that demand verified integration with external systems. The result is a faster, more reliable CI pipeline where failures are indicative of genuine defects rather than environmental noise.
Quarantine, Don’t Ignore Flaky Tests
Flaky tests are tests whose outcomes change without any code modification, eroding confidence in continuous‑integration (CI) pipelines. Before applying any mitigation, teams must first identify which tests are flaky and quantify their instability. A common technique is to execute the full test suite multiple times on the same commit and record the results; any test that both passes and fails across runs is flagged as flaky. CI platforms that collect JUnit XML reports can automate this measurement and rank flaky tests by failure rate.
Once a flaky test is identified, the recommended containment strategy is a quarantine workflow. In pytest, a test is marked for quarantine with a custom marker:
@pytest.mark.quarantine
def test_payment_callback_timeout():
# test implementation
...
The marker enables two distinct CI stages:
- Blocking stage: runs all tests except those marked
quarantine(pytest -m "not quarantine"). Failures here prevent the merge. - Non‑blocking stage: runs only quarantined tests (
pytest -m quarantine || true). Results are reported but do not block the pipeline.
To keep quarantine from becoming a permanent repository of broken tests, the process must include ownership, tracking, and regular review:
- Owner assignment: each quarantined test is linked to a responsible engineer who is accountable for fixing it.
- Ticket creation: a work item (e.g., a Jira ticket) is opened at the moment a test is quarantined, containing the test name, failure pattern, and owner.
- Review cadence: the quarantine list is examined in a recurring meeting (e.g., bi‑weekly). Tests that remain flaky beyond an agreed threshold are escalated for deeper investigation.
Practical implementation steps for a new project might look like this:
- Instrument the CI pipeline to collect flaky‑test metrics for each build.
- Introduce the
@pytest.mark.quarantinemarker in the test suite. - Configure two CI jobs: one with
-m "not quarantine"(required) and one with-m quarantine(optional). - Automate ticket creation using a script that parses the quarantine job’s JUnit report.
- Document the quarantine policy in the team’s engineering handbook and enforce the review schedule.
By isolating flaky tests in a non‑blocking stage while maintaining visibility through tickets and ownership, teams preserve pipeline reliability, avoid “re‑run and hope” cycles, and ensure that flaky tests are systematically resolved rather than silently ignored.
Treat Automatic Retries as a Last Resort
Flaky tests are those that pass and fail intermittently without code changes, often because of timing, shared state, or external dependencies. When a test fails sporadically, teams instinctively add a retry mechanism to keep the CI pipeline green. While retries can reduce immediate noise, they also mask the underlying instability, allowing genuine defects to slip through the review process.
Retry plugins work by re‑executing a failing test a configurable number of times. The evidence shows that this practice should be treated as a last resort. Unlimited retries create a false sense of reliability, encourage developers to click “re‑run” reflexively, and erode trust in the pipeline. Instead, the pipeline should surface flaky behavior so it can be addressed directly.
Recommended retry policy
- Limit retries to one. A single retry provides a safety net for transient infrastructure hiccups without concealing systematic problems.
- Log every attempt. Include the original failure message, the retry count, and timing information in the CI logs. Example:
TEST FAILURE: test_payment_callback_timeout
Attempt 1: Timeout after 30s
Attempt 2 (retry): Passed
- Mark retried tests as flaky. Store the test identifier in a “flaky registry” (e.g., a database or a dedicated file) and surface it on dashboards so owners can prioritize fixes.
- Enforce ownership. Each flaky entry must have an assigned engineer and an open ticket that tracks remediation progress.
Implementing the policy can be done with a minimal wrapper around the test runner. In pytest, a custom plugin could enforce the limit and emit structured JSON logs:
@pytest.hookimpl(tryfirst=True)
def pytest_runtest_makereport(item, call):
if call.when == "call" and call.excinfo is not None:
if not getattr(item, "has_retried", False):
item.has_retried = True
# trigger a single retry
item.rerun()
else:
# record as flaky
log_flaky(item.nodeid, call.excinfo.value)
By treating retries as an exception rather than the norm, teams keep the focus on eliminating root causes—such as fixed sleeps, hidden shared state, or uncontrolled external services—while preserving fast, trustworthy feedback loops.
Looking for Custom Software or AI Solutions?
Appworks Technologies designs, builds, and scales production enterprise platforms, microservices, and AI agent workflows tailored to your business goals.
Editorial Policy & Research Methodology
Our findings are based on rigorous internal research, verified industry benchmarks, and direct technical implementation experience from our enterprise client projects. All statistics and technical claims are reviewed by senior engineers before publication to ensure accuracy, transparency, and helpfulness for our readers.
