CI/CD operations
GitHub Actions timeout: bound retries and cleanup
Diagnose where a run waits before raising a GitHub Actions timeout. A job timeout bounds execution on its runner; a step timeout bounds one step, while queueing, dependencies, approvals, and external deployment confirmation are separate clocks. Retry only a known transient operation, record its first failure, and reserve time for cleanup and target verification. Cancellation may interrupt cleanup, so an uncertain deployment needs a follow-up check against the destination.
Identify which clock is running
A workflow that takes an hour is not necessarily a job that ran for an hour. It may wait for an event, a runner, an upstream job, a concurrency group, or an environment approval before the affected job executes. GitHub records workflow, job, and step states separately; inspect those records before changing timeout-minutes. If a job never reached a runner, extending its execution allowance does not cure the wait. If a prerequisite job is still running, tune that dependency or its fan-out instead of the downstream job. If a deployment provider accepted a request but has not confirmed a live revision, the uncertainty is outside the build step's clock.
Name each deadline in the incident note: event-to-run delay, queue delay, dependency wait, job execution, individual step duration, and delivery verification. The workflow run itself has a separate overall limit that includes waiting and approvals. A self-hosted job can also sit in a queue until its queue limit, which is not the same as its execution cap. A timeout setting is useful only when it bounds the operation you intend to bound. Use the run graph and step logs to find where elapsed time accumulated, then choose the narrowest control that exposes a meaningful failure.
Keep the revision and run identity with every timestamp. Two runs for the same pull request can be pending or running at once, and a cancelled old run must not be mistaken for the current candidate. Record when the job started, when the last useful log appeared, which step was active, and whether a delivery operation had already begun. This evidence tells you whether the action is to add runner capacity, fix a stuck dependency, bound a command, or inspect an external target. Raising every timeout to the maximum only delays that decision.
Sources: Actions limits — GitHub Docs · Using workflow run logs — GitHub Docs · Workflow syntax for GitHub Actions — GitHub Docs
Know the current GitHub limits without flattening them
GitHub's current workflow syntax gives jobs a default timeout of 360 minutes, but a larger configured value cannot exceed the runner's execution limit in practice. The current Actions limits list a six-hour execution cap for GitHub-hosted jobs and a five-day cap for self-hosted jobs. Those are different limits, not a universal 360-minute ceiling. Step timeout-minutes is a separate setting: GitHub currently caps an individual step at 360 minutes for both runner types and requires a positive whole number. Use the current documentation and the actual runner class when reviewing a long-running job.
A workflow run has its own overall limit that includes waiting and approval time, and a self-hosted job has a queue limit before automatic cancellation. Neither is extended by giving a single job a longer execution allowance. GITHUB_TOKEN is another clock: GitHub documents that it expires at job end or after its effective maximum lifetime, with a shorter practical span on hosted jobs and a 24-hour refresh boundary on self-hosted jobs. A self-hosted job can therefore remain alive after its GitHub token can no longer authenticate. If a long job needs GitHub API access late in its life, design that authentication explicitly rather than assuming the job timeout governs token validity.
Do not choose a timeout by asking only how long GitHub permits. Set it from an expected operation and a failure budget. A test that normally takes four minutes but once hangs on network setup should fail in a bounded, explainable period. A release that calls an external provider needs its own provider request deadline and a way to query the result. If the provider is still acting after the CI job times out, the job's terminal state does not prove that nothing shipped. Treat the external operation as unresolved until its recorded revision and target state are reconciled.
Sources: Workflow syntax for GitHub Actions — GitHub Docs · Actions limits — GitHub Docs · GITHUB_TOKEN — GitHub Docs
Calculate a worst-case retry budget before enabling retries
Here is a synthetic budget for a flaky external dependency probe, not a benchmark or a recommended default. The job needs up to three minutes of setup. It may make an initial attempt and at most two retries, each capped at eight minutes. It waits two minutes after the first failure and four after the second, then reserves five minutes for local cleanup and seven for scheduling or teardown margin. The arithmetic is three plus twenty-four plus six plus five plus seven, or forty-five minutes. A job timeout of forty-five minutes matches this worst-case plan; a step timeout of eight minutes can bound each attempt only if the attempt is a distinct step or the invoked command enforces its own deadline.
| Component | Minutes | Evidence to preserve |
|---|---|---|
| Setup | 3 | Runner and dependency initialization result |
| Initial attempt plus two retries | 8 + 8 + 8 = 24 | Each attempt's result, especially the first failure |
| Backoff between attempts | 2 + 4 = 6 | Reason each retry was permitted |
| Local cleanup | 5 | Whether resources were actually released |
| Margin | 7 | Remaining time before the job cap |
| Total | 45 | Terminal result for the same revision and run |
The table is a ceiling for one selected transient probe, not permission to repeat every failing test. Retry a failure only when its signature is known to be transient and the operation is safe to repeat. An assertion failure, changed dependency, permission denial, or wrong revision is usually evidence to fix, not noise to bury. Save the first causal log before the retry changes the environment or replaces a clearer error with a timeout. Record attempt number, start time, failure class, and backoff; the final green attempt should not erase the fact that two failures occurred.
For a deploy operation, idempotence matters more than a retry count. If the first request may have reached the provider, query the operation identifier or target revision before sending another request. A second request can duplicate work or roll the target again, even if the first CI step reported a network timeout. Reserve cleanup time for local resources, but use target-side evidence to close the delivery. If the provider cannot tell you whether the first operation applied, stop automatic retries and surface an unresolved delivery state.
Sources: Workflow syntax for GitHub Actions — GitHub Docs · Using workflow run logs — GitHub Docs · Workflow cancellation reference — GitHub Docs
Apply a job cap and narrower step cap
The partial YAML below illustrates only the GitHub timeout keys. It is not a complete workflow: it omits the trigger, checkout, permissions, retry implementation, and the command that probes a dependency. When the reviewed step is supplied, the eight-minute step cap will apply to that step; it will not automatically wrap a loop inside its command into eight-minute attempts. If retries happen in one script, the script must enforce its own per-attempt deadline and backoff limit. The forty-five-minute job cap still covers all executed steps together, including setup and cleanup.
# Partial job only; supply the actual trigger, runner, and reviewed step.
jobs:
probe:
timeout-minutes: 45
steps:
- name: Bounded dependency check
timeout-minutes: 8
# Add the repository's reviewed run or uses step here.A step timeout kills the process when it reaches the limit; it is not a request to finish gracefully before the deadline. The job timeout can cancel the whole job earlier if setup and previous attempts have already consumed its budget. Set an application-level deadline shorter than the outer step cap when the command can close connections, write a useful error, and free its own resources. For an external deployment, add a target query that can run after an uncertain response; do not rely on a shell exit code to say what the provider did.
Sources: Workflow syntax for GitHub Actions — GitHub Docs · Workflow cancellation reference — GitHub Docs
Treat cancellation and cleanup as separate outcomes
GitHub's cancellation process reevaluates job and step conditions, sends messages to runners, and then signals processes that need to stop. It can forcibly terminate work that does not exit within the cancellation period. A cleanup step conditioned to run after failure may still be skipped or interrupted when the job is cancelled, and a process may receive a signal before it writes its final log. Do not equate the presence of a cleanup step with proof that cleanup completed. Inspect the step's result and, for external resources, query the resource itself.
There are two useful terminal questions. First, is the runner job finished, failed, timed out, or cancelled for this revision? Second, did the external thing it touched finish, roll back, remain active, or become unknown? A timed-out test process may leave a local service that the runner teardown handles. A timed-out deploy request may leave a production operation still progressing. The latter requires a bounded follow-up from the delivery controller or operator, with the target revision read back. A green later run cannot retroactively prove what happened to the earlier external operation.
When a timeout repeatedly fires, classify the elapsed phase and the first causal error before increasing the number. If the job waits for capacity, repair routing or capacity. If a dependency stalls, bound and instrument that call. If retries consume the budget, narrow the retryable errors or reduce attempts. If cleanup itself times out, move recovery to a mechanism that can act after the runner is gone. The outcome to optimize is a fast, correct diagnosis and a verified shipment, not a longer period before an ambiguous red badge.
Sources: Workflow cancellation reference — GitHub Docs · Using workflow run logs — GitHub Docs · Actions limits — GitHub Docs
Sources and verification
- Workflow syntax for GitHub Actions — GitHub Docs
GitHub documents job timeout default and runner-cap interaction, step timeout maximum and integer rule, and timeout-minutes syntax. Verified .
- Actions limits — GitHub Docs
GitHub lists current hosted and self-hosted job execution caps, workflow run and approval limits, and self-hosted queue limit. Verified .
- GITHUB_TOKEN — GitHub Docs
GitHub documents token expiry at job end and effective maximum lifetimes for hosted and self-hosted jobs. Verified .
- Workflow cancellation reference — GitHub Docs
GitHub explains conditional reevaluation, runner cancellation messages, process signals, and eventual forced termination. Verified .
- Using workflow run logs — GitHub Docs
GitHub documents job and step log inspection and step execution timing for diagnosing a workflow run. Verified .