Free CI tools

Test parallelization break-even calculator

Test parallelization helps this model when the time saved by dividing parallel work exceeds the added overhead. Enter a single-worker baseline, a parallelizable fraction, a worker count, and extra overhead to explore that tradeoff. The default hypothetical scenario changes 600 seconds to 225 seconds with four workers. That is a modeled 2.67× speedup, not a measured result or a guarantee for your test suite.

Explore parallelization break-even

Compare an idealized parallel run with a single-worker baseline. Inputs are scenarios, not measured speedup claims; calculations stay in your browser.

Enter 0–86,400 seconds for the same workload on one worker. Zero is valid, but speedup is then not applicable.

Enter 0–1; 0.9 means 90% of baseline time can be divided evenly among workers. This is an assumption to check.

Enter a whole number from 1 to 256. One worker preserves the baseline and ignores additional parallel overhead.

Enter 0–86,400 seconds added once to elapsed time when using more than one worker. Exclude setup already counted in the baseline.

Modeled elapsed seconds
225
Modeled speedup (×)
2.67
Reserved worker-seconds
900

Amdahl-style scenario with equal work distribution and fixed added overhead. Reserved worker-seconds assume every worker is reserved for the whole modeled run; this is not actual CPU use or billing. Zero-baseline speedup is Not applicable. Validate against observed scheduling, imbalance, and the critical path.

Shared links contain your entered numbers. Check them before sharing.

Separate serial time, parallel work, and overhead

Start with the elapsed time for one worker to complete a defined workload. The parallelizable fraction is the share of that baseline time you assume can be divided evenly between workers. The remaining share stays serial. Additional overhead represents elapsed time introduced by the parallel arrangement, beyond work already in the baseline. The tool adds that overhead once when more than one worker is selected. At one worker it returns the original baseline, which keeps the comparison anchored to the starting workload.

This is an Amdahl-style model: increasing parallel capacity does not remove the part that must remain serial. Amdahl’s original paper discusses sequential overhead and the irregularities that complicate parallel execution. Here we use a deliberately small model to make a proposed tradeoff explicit. Its parallel fraction and overhead are inputs supplied by you. The paper is background for the limitation, not evidence that a particular test suite has the default fraction or will achieve the displayed speedup.

Elapsed-time and reserved-capacity model
For one worker: modeled elapsed time = baseline
For more than one worker:
  elapsed = baseline × (1 − parallel fraction)
          + baseline × parallel fraction ÷ workers
          + additional overhead

Speedup = baseline ÷ modeled elapsed time
Reserved worker-seconds = workers × modeled elapsed time
A zero baseline makes speedup Not applicable.

Sources: Gene M. Amdahl: 1967 paper (Carnegie Mellon copy)

Work through the 600-second example

The defaults assume a 600-second single-worker run, a parallel fraction of 0.9, four workers, and 30 seconds of extra overhead. Ten percent of the baseline remains serial, contributing 60 seconds. The parallel portion contributes 540 divided by four, or 135 seconds. Adding 30 seconds gives a modeled elapsed time of 225 seconds. Dividing the 600-second baseline by 225 gives approximately 2.67. All these values come from the hypothetical inputs; none is an observed benchmark.

Four reserved workers over 225 seconds represent 900 reserved worker-seconds. This is larger than the single-worker scenario’s 600 reserved worker-seconds even though its elapsed time is shorter. The measure assumes every worker remains allocated throughout the modeled run, including its serial work and overhead. It is not actual CPU consumption, active execution time on each worker, or billable usage. A real scheduler that allocates or releases workers at different times will have a different capacity profile.

Try the endpoints to check your interpretation. With no parallelizable work and no extra overhead, adding workers leaves elapsed time unchanged. With fully parallel work and zero overhead, the modeled speedup equals the worker count. A zero baseline produces an undefined speedup comparison, so the tool reports Not applicable rather than zero or infinity. Extra overhead can still produce a positive modeled duration in that zero-baseline scenario when multiple workers are selected, but it cannot make the speedup ratio meaningful.

Find the overhead that cancels the benefit

For more than one worker, break-even occurs when extra overhead equals the baseline multiplied by the parallel fraction and by one minus the reciprocal of the worker count. This follows directly by setting the modeled elapsed time equal to the baseline and rearranging the formula. In the default scenario, that threshold is 600 × 0.9 × 0.75, or 405 seconds. At 405 seconds of overhead, the modeled run takes 600 seconds and the speedup is one. Less overhead produces a shorter modeled duration; more produces a longer one.

A speedup below one is useful information, not an invalid calculation. It means the proposed inputs describe a slower arrangement than the baseline. For example, leave all 600 seconds serial and add 600 seconds of overhead: the result is 1,200 seconds, or a 0.5× speedup. Do not adjust the fraction until the output looks attractive. Instead, use the scenario to identify which assumption requires evidence before you change the runner configuration or allocate additional capacity.

Map the model to your runner’s actual work units

A worker is a unit in this model, not automatically a machine, a CPU core, or a CI job. Playwright documents worker processes, configurable worker limits, and the distinction between running files in parallel and running tests within a file in parallel. Its sharding guide describes distributing work across separate shards. Identify which level you are changing before choosing a worker count. Multiplying two configuration values without checking the resulting execution structure can obscure the capacity being requested.

The parallel fraction is a fraction of time, not a fraction of test names. Ninety short tests and ten long tests do not imply that 90% of execution can be distributed evenly. A long indivisible file, a required serial sequence, or an uneven shard can leave other workers waiting. Playwright’s sharding documentation explains how file-level and test-level distribution affect balancing. Inspect the work units your actual configuration distributes and measure their durations before assuming that adding a worker divides the parallel portion perfectly.

Sources: Playwright: Parallelism · Playwright: Sharding

Compare the prediction with the critical path

Define the measurement boundary before comparing runs. Decide whether your baseline includes dependency setup, test execution, artifact collection, and report merging. Then use the same boundary for the parallel run. Record the revision, test selection, environment, worker configuration, cache condition, and concurrent workload. A comparison that changes both the test set and the worker count cannot isolate the effect of parallelization. Keep the evidence that explains any difference rather than reducing every result to a single speedup number.

Look at the critical path: the sequence of dependent work that determines when the run can finish. A fast group of tests may complete early while the final slow job still holds the result open. Queue delays can add elapsed time before execution starts, and shared services can make job durations grow when more workers compete. The calculator does not schedule tasks, model those dependencies, or predict contention. Its fixed overhead field is a scenario input, not a simulation of how overhead changes at every worker count.

Use a controlled comparison to test whether the model is useful. Keep the required checks intact, try a clearly identified configuration, and inspect both the completed work and its timing. If more workers create failures or resource pressure, diagnose that behavior before treating a quick successful rerun as proof. A smaller measured duration matters only when the run still performs the intended verification. The useful next action may be improving an unbalanced work unit or removing duplicated setup rather than choosing a larger worker count.

Sources: Gene M. Amdahl: 1967 paper (Carnegie Mellon copy) · Playwright: Sharding

Use the result as a planning comparison

Save the assumptions beside the result when discussing a proposed change. The copied link contains the numeric inputs, and the JSON export preserves unrounded values and the model’s limitations. Displayed results are rounded for readability; a zero-baseline speedup remains null in JSON and Not applicable in readable output. Compare several plausible assumptions if the parallel fraction or overhead is uncertain. A single precise-looking output should not hide uncertainty in the inputs that produced it.

For a question about how many jobs a matrix creates, use the separate matrix budget calculator. For execution costs, use observed workload and your applicable rate in the CI cost calculator instead of treating reserved worker-seconds as an invoice. Agentic Pipeline’s public feature guide describes run controls and delivery evidence that coding agents can inspect. Evaluate those capabilities against a real delivery task; access is closed beta. This tool does not promise a performance gain from any product or recommend a worker count without workload evidence.

Sources: Agentic Pipeline: Matrix budget calculator · Agentic Pipeline: CI cost and rerun waste calculator · Agentic Pipeline: Public feature guide

Sources and verification

Browse all resources