Free CI tools

Flaky tests: suite and retry probability calculator

If 100 independent tests each have a 1% chance of flaking, the chance of at least one flake in a suite is about 63.4%. This free calculator helps teams explore that compounding effect and a bounded whole-suite retry scenario. The inputs are hypothetical: equal per-test probabilities, independent tests, and independent attempts. The output cannot diagnose a failure or establish your actual CI reliability.

Explore flaky-test probability

Model independent tests and whole-suite attempts with hypothetical inputs. Calculations stay in your browser; no signup or upload is needed.

Enter 0–1; 0.01 means 1%. Assumes every test has the same probability of flaking, independently of the others.

Enter a whole number from 1 to 1,000,000. Shared services, clocks, or state can invalidate the independence assumption.

Observed or planned first attempts in your chosen period, excluding retries. Enter a whole number from 0 to 1,000,000.

Enter a whole number from 0 to 10. One retry allows two attempts total. Assumes each full-suite attempt is independent.

At least one flake in a suite
63.4%
Expected initially flaked executions
63.4
Every allowed attempt flakes
40.19%

Illustrative model: equal per-test probabilities, independent tests, and independent attempts. Expected counts are averages, not guaranteed outcomes. Real bugs and correlated infrastructure failures are outside this model. Whole-suite retries differ from retrying only a failed test.

Shared links contain your entered numbers. Check them before sharing.

Start with a precisely defined event

This model asks whether at least one test flakes during a suite execution. It does not count failed assertions or calculate the probability that a code change contains a bug. Each test is assigned the same hypothetical flake probability, p, for each execution. The suite contains N tests. Independence means that learning one test’s outcome does not change the model’s probability for another test. Those are modelling choices to make the arithmetic understandable, not observations about your repository.

A test avoids a flake with probability one minus p. Under independence, all N tests avoid a flake with probability (one minus p) raised to N. Subtract that from one to obtain q, the probability that at least one test flakes. The formulas below describe our illustrative model. Martin Fowler’s discussion of nondeterministic tests provides background on why inconsistent results deserve investigation; it does not establish these assumptions or validate the numbers you enter.

The independent-test and independent-attempt model
q = 1 − (1 − p)^N
Expected initially flaked executions = R × q
Probability every allowed whole-suite attempt flakes = q^(K + 1)

p = per-test flake probability
N = independent tests per suite
R = initial suite executions, excluding retries
K = extra whole-suite retries allowed

Sources: Martin Fowler: Eradicating Non-Determinism in Tests

Why 1% per test becomes 63.4% per suite

Use the default scenario of 100 tests, each with a 0.01 probability of flaking. The probability of all tests avoiding a flake is approximately 0.3660, leaving approximately 0.6340 for at least one flake. That is a 63.4% suite probability, not a 1% suite probability. It is also not 100%: adding one hundred individual 1% probabilities double-counts outcomes in which more than one test flakes. The complement calculation counts each affected suite execution once.

For 100 initial suite executions, multiply q by 100 to get an expected count of about 63.4 executions with at least one flake. An expected count is a model average across repeated comparable experiments. An actual set of 100 executions has a whole-number result and need not contain exactly 63 or 64 affected executions. Entering zero initial executions makes the expected count zero while leaving the per-suite probability unchanged. The workload size and the probability of an individual event are separate inputs to the interpretation.

Try one test with a probability of 0.5 as a simple check: its suite probability is 50%. With two independent tests at that probability, the suite probability becomes 75%. These small examples expose the calculation without implying anything about a real test population. Larger values explore sensitivity to assumptions; they do not replace evidence about whether the tests share state or whether the entered probability is representative.

What the retry result means

One extra retry allows two full-suite attempts in this model. The probability that both contain at least one flake is q multiplied by q. With the default inputs, that is approximately 40.19%. With no extra retries, the result equals the initial suite probability. The calculator allows at most ten extra retries and evaluates only that finite attempt budget. This output is an unconditional probability before the initial attempt, not the probability of a second failure after you have already observed a first one.

Whole-suite repetition is not the same as a framework retrying only a failed test. Playwright’s documentation describes retries of failing tests and distinguishes first-attempt passes, passes after a retry, and failures after the retry budget. Its worker lifecycle and serial-group behavior also affect what executes again. Check your runner’s actual retry scope before comparing its reports with this tool. The documentation explains framework behavior; this calculator does not reproduce that runner’s scheduling or compute its exact eventual suite verdict.

A passing retry cannot tell this model whether the first result came from a timing problem, a real product defect, or a transient infrastructure issue. Likewise, several failed attempts do not prove a deterministic bug. Do not convert the retry result into permission to ignore an initial failure. It illustrates the consequences of assumed independent whole-suite attempts, which is a narrower question than whether a change is safe to release.

Sources: Playwright: Retries

Check where independence breaks

Shared clocks, services, and state can connect outcomes that the formula treats as separate. Consider a hypothetical suite whose tests all depend on one service: if that service becomes unavailable, many tests may fail together. Repeating the suite immediately may encounter the same outage. A common clock boundary can similarly affect several time-sensitive checks at once. In both examples, the situation influencing one outcome also informs the others, so multiplying independent probabilities no longer describes the experiment.

Equal probabilities are another restriction. A suite with one frequently unstable test and many stable tests is different from a suite in which every test has the same small chance of flaking. Entering an average does not generally preserve the probability of at least one flake. The direction and size of error depend on the actual probabilities and their dependence. This tool intentionally does not infer either from a test count; treat the result as a scenario until you have checked whether those assumptions are useful.

Fowler identifies isolation, asynchronous behavior, remote services, time, and resource leaks as sources of nondeterminism. Use those categories to organize an investigation, rather than assuming every inconsistent result is an independent random event. The calculator is most useful for explaining why many small assumed risks can accumulate. It is less useful as a forecast when a shared dependency or a lasting environmental condition dominates the failures.

Sources: Martin Fowler: Eradicating Non-Determinism in Tests

Record evidence before estimating a flake rate

For an investigation, collect attempts at the same revision with comparable environment, dependencies, test selection, and execution settings. Record the original failure as well as the retry. Keep the observed period and the number of eligible executions explicit. A count assembled from changing code, mixed runner configurations, or only the failures people chose to rerun can answer a different question from the one the calculator models. An unexplained difference should remain unexplained in the record, rather than being automatically classified as a flake.

Attach a failure signature and a minimal reproducer to any proposed flake classification. Note which condition changes the outcome and which alternatives remain plausible. If you cannot reproduce the behavior, preserve the available logs and uncertainty. A test runner’s retry label can locate an attempt worth examining, but the label does not establish its cause. Separate a measured observation from the hypothetical per-test probability entered here, and retain the underlying counts so another person or agent can assess the interpretation.

  • Identify the revision, environment, test selection, and observation period.
  • Keep first-attempt results and retry results separately, including failed retries.
  • Record shared dependencies and any condition that changed between attempts.
  • Preserve a reproducer or mark the cause as unresolved.
  • State whether the calculator input is a hypothetical scenario or an estimate from defined observations.

Sources: Playwright: Retries

Choose a repair task, then check its effect

Use the scenario to explain a problem, then choose a specific investigation based on the evidence. For example, isolate the input that changes a reproduced outcome and verify the proposed repair against that reproducer. Compare equivalent attempts before and after the change, retaining the original failure record. A lower displayed number after editing a probability field proves only that the hypothetical scenario changed. It does not prove a test became reliable or that the delivery process became safer.

If repeated execution also creates a cost question, use the separate CI cost calculator with observed job executions and your own rate. Keep that accounting separate from this probability model. For teams where coding agents investigate delivery failures, Agentic Pipeline’s public feature guide describes agent-readable logs and run controls. Those capabilities can be evaluated as part of the evidence workflow; they are not an automatic classifier for flaky tests. Access remains closed beta, and this guide makes no promised reduction in failures, retries, or execution cost.

Sources: Agentic Pipeline: CI cost and rerun waste calculator · Agentic Pipeline: Public feature guide

Sources and verification

  • Martin Fowler: Eradicating Non-Determinism in Tests

    Fowler explains inconsistent tests and discusses isolation, asynchronous behavior, remote services, time, and resource leaks as causes to investigate. The probability formulas are this page’s model. Verified .

  • Playwright: Retries

    Playwright documents failing-test retries, outcome categories, worker restarts, and serial-group retries. These semantics differ from independently repeating an entire suite. Verified .

  • Agentic Pipeline: CI cost and rerun waste calculator

    The separate execution-cost tool uses job counts, rounded job duration, a reader-entered rate, and an observed extra execution ratio. Verified .

  • Agentic Pipeline: Public feature guide

    The public guide describes agent-readable CI/CD evidence and run controls; access is closed beta. Verified .

Browse all resources