GitHub Actions

GitHub Actions logs: find the first causal failure

Start a GitHub Actions log investigation with the exact workflow run, attempt, job, step, and source SHA. Read the ordered steps and find the earliest failure that explains later symptoms, rather than the last red line in the archive. Keep warnings and cleanup failures in context. A rerun tests the original event's SHA and ref again; a source fix needs a new commit and new run. Preserve a short, redacted handoff with the causal line, evidence for the hypothesis, and the next check.

Fix the run and attempt before reading a log line

A GitHub Actions failure has an identity more precise than a workflow name. Record repository, event, run ID, run attempt, job ID, step name or number, source SHA, and the result of that attempt. The same workflow can execute for a pull request, push, or manual event; two runs can share a display title while testing different commits. A partially rerun workflow can also combine jobs whose logs belong to different attempts. If the agent omits the attempt, it may diagnose a fresh green step beside an older failing step and describe a sequence that never happened.

GitHub's workflow log documentation supports viewing and downloading job logs and linking a specific line. It also notes that a log archive from a partial rerun includes only jobs rerun in that attempt; previous attempts hold the other jobs. Begin with the run summary and job graph to identify which job failed, then open its failed step and neighboring earlier steps. Search is a navigation aid, not a substitute for order: an error word in post-job cleanup can appear after the original failure. Keep the nearby timestamps and step boundaries so the explanation can be checked by another agent.

For API-driven retrieval, GitHub's workflow-jobs endpoints expose job and step logs through short-lived redirect URLs. The documented download link expires after one minute, so store the stable run, job, and step identifiers in the handoff rather than a temporary redirect. Private repository reads require appropriate access, and fine-grained tokens need Actions read permission for the job-log endpoint. A missing or expired download link is a retrieval problem, not evidence that the workflow produced no logs. Respect the repository's configured retention period when deciding whether evidence may still be available.

Sources: Using workflow run logs — GitHub Docs · REST API endpoints for workflow jobs — GitHub Docs · Removing workflow artifacts — GitHub Docs

An ordered failure that ends with misleading noise

Synthetic run 8421, attempt 1, source SHA 1111111111111111111111111111111111111111, job 73; no real workflow or customer data.
Time and stepIllustrative log signalDiagnostic reading
09:02 — Restore cacheCache key miss; step succeededExpected performance variation, not a failed gate
09:03 — Build packageType error: property `currency` missing on `PriceInput`; process exits nonzeroFirst causal failure: source/build contract at this SHA
09:04 — Run testsSkipped because the build job path failedNo test result for this attempt; do not mark tests passing
09:05 — Always-run package uploadPackage directory not foundDownstream artifact symptom of failed build
09:06 — Complete jobPost-job cache save skipped and temporary directory cleanup warnedLater cleanup noise to record, not the root cause of the failed build

The first line containing a warning is the cache miss, but the first causal failure is the build step's nonzero type error. The absent package follows from the failed build. The skipped tests provide no evidence that behavior was correct or incorrect. Cleanup messages may deserve their own follow-up if they reveal a separate problem, yet they do not explain why this build stopped. The agent should open the source location for `PriceInput` at the recorded SHA, compare the changed code with the type contract, and reproduce the compiler diagnostic on that revision before proposing a fix.

Suppose a patch adds the missing property and creates a new commit represented here by SHA 2222222222222222222222222222222222222222. The relevant proof is a new run for that new revision with the build and tests selected. Rerunning run 8421 would test the original source again, even if a branch now points elsewhere. Conversely, if the first causal line had been a clearly transient registry transport error, a bounded rerun of the original revision could help distinguish infrastructure noise from a source defect. Neither choice means retry until green; the diagnostic hypothesis determines the next experiment.

Sources: Using workflow run logs — GitHub Docs · Re-running workflows and jobs — GitHub Docs · REST API endpoints for workflow jobs — GitHub Docs

Trace cause through step order and exit status

Read from the earliest failed dependency on the job's path. Identify the command or action step that returned a failure, the first diagnostic explaining it, and which later steps were conditional, skipped, or allowed to run for cleanup. An application stack trace can be useful, but the final stack frame is not always the cause. A package-not-found message can be a consequence of an earlier build failure; a test-report upload can fail because no tests ran. Verify this relationship from the run's step ordering and recorded outputs, not from the emotional weight of the last red message.

If multiple jobs fail independently, keep separate hypotheses. A matrix may have one environment with a compiler failure and another with a runner provisioning error. Do not flatten them into a single generic failed CI summary. Record which jobs share the same source SHA and which attempt each belongs to, then decide whether one change can plausibly explain both. A broad search for the word error across a combined archive may mingle jobs and time zones; use the job graph and step identities to rebuild the causal order before assigning a root cause.

Separate observation from inference in the handoff. Observation: build step exited nonzero with a type diagnostic at this line on this SHA. Inference: the changed input type likely omitted a field required by the package. Next proof: compile the candidate fix and see the relevant tests run on the new SHA. That structure gives the next agent a falsifiable path. It also avoids claiming that every later warning is harmless: if cleanup itself failed independently, note it as a second finding after the original build cause has been established.

Sources: Using workflow run logs — GitHub Docs · REST API endpoints for workflow jobs — GitHub Docs

Choose a rerun only when it tests the hypothesis

GitHub documents rerunning a whole workflow, failed jobs, or a specific job. A rerun uses the original event's `GITHUB_SHA` and `GITHUB_REF`, and the privileges of the original triggering actor. It creates another attempt of the same run, not a new run for a newly pushed fix. This is useful when the hypothesis concerns a transient dependency, flaky test, or runner fault on unchanged source. It is not a way to validate a code correction. If the source changes, review the new commit and follow its own run and checks.

Preserve attempt one before starting attempt two. Compare the same job and step across attempts, including dependency versions or external service conditions if those were material. If the rerun passes, the first failure still happened; record what changed between attempts and whether the underlying reliability issue is understood. A single pass does not prove a test is stable. If the rerun fails in the same place, stop spending attempts and investigate the deterministic path. GitHub imposes rerun windows and limits, but operational policy should be narrower than exhausting an available retry quota.

When only failed jobs are rerun, the newest archive will not contain every job from the original attempt. GitHub explicitly directs readers to previous attempts for logs of jobs that were not rerun. A green rerun result therefore needs context: which jobs actually executed again, and which successful jobs were inherited from the earlier attempt? For a required check, inspect its evaluated SHA and current conclusion rather than using a screenshot of one rerun step. The purpose of a rerun is to gain evidence about a defined failure mode, not to erase a red history entry.

Sources: Re-running workflows and jobs — GitHub Docs · Using workflow run logs — GitHub Docs

Use extra logging briefly and keep the handoff safe

When ordinary logs do not distinguish competing causes, GitHub offers runner diagnostic and step debug logging. Runner diagnostics add process and worker logs to the archive; step debugging increases event detail. Enable the smallest useful scope for a controlled rerun, then return to normal verbosity after the investigation. Extra output can make a short causal chain harder to read and can increase the amount of sensitive context present in a stored log. It should answer a specific question, such as whether the runner failed before the build step started.

Do not print tokens or secret values to make a failure easier to explain. GitHub masks many secrets, but its own documentation says redaction is not guaranteed for transformed values and is limited to secrets used within the current job. Review a log excerpt before copying it into a ticket or agent message; replace credentials and personal data with neutral placeholders while retaining the error structure and line location. A temporary download URL for an API log is also a poor long-term handoff link because it expires quickly. Use stable run and job references subject to repository access.

The final handoff can fit in a few lines: run and attempt, event and source SHA, failed job and step, first causal diagnostic, later consequences, confidence level, and proposed next proof. For an agent using Agentic Pipeline, the same discipline applies to its own run and delivery evidence: identify the exact failing stage and revision before taking another action. This page describes GitHub Actions' log surfaces specifically; it does not imply that another platform has identical API endpoints or archive behavior. Aggregate CI/CD monitoring answers where failures cluster over time, while this workflow explains one failure well enough to fix or rerun it safely.

Sources: Enabling debug logging — GitHub Docs · Secrets — GitHub Docs · REST API endpoints for workflow jobs — GitHub Docs

Sources and verification

Browse all resources