Deployment operations

CI/CD monitoring: trace runs to verified deployments

CI/CD monitoring should connect a triggering revision, run, check result, build artifact, deployment operation, and observed target revision. Watch each transition separately: queued, building, testing, delivering, and verified live. A green check or a merged commit does not by itself count as a deployment. Treat a missing operation or target readback as missing evidence, not a measured zero or a success. Alert when a transition fails, remains stale under a defined policy, or completes without proof that the intended revision serves the target.

Keep one evidence chain across the pipeline

A useful CI/CD monitoring record starts with identifiers rather than a wall of charts. Capture repository, triggering event, source SHA, run ID and attempt, check name and reporting app, check conclusion and evaluated SHA, build digest, merge revision, destination environment, provider operation ID, and observed target revision. These fields answer different questions. The check says whether a specific evaluated revision passed a gate; the artifact says what was built; the operation says what the provider accepted; the target marker says what is serving. Correlating them avoids interpreting a green check on one revision as proof that another revision reached production.

GitHub's Checks API associates check runs with commits and includes the reporting app, status, and conclusion. Its Deployments API records a deployment's SHA and environment, while deployment statuses can represent progress and outcomes. Those are useful source records for repositories that use them, but the provider and application still need to supply their own delivery and live-version evidence. The API objects do not guarantee that traffic is reaching the intended binary. Keep the join keys explicit even when events come from separate systems; a display label or branch name can be reused and is a weak substitute for exact identities.

OpenTelemetry publishes CI/CD semantic conventions for spans, metrics, and logs. Their current CI/CD status is Release Candidate, so record the convention version if adopting its names and plan for possible evolution. The run-resource guidance also warns that run-level attributes on metrics can create high cardinality and cost; detailed run IDs belong in trace or log records and a drill-down ledger, while aggregate metrics use bounded dimensions. The guide's ledger is a design worksheet, not a claim that any product dashboard already implements all these fields.

Sources: REST API endpoints for check runs — GitHub Docs · REST API endpoints for deployments — GitHub Docs · REST API endpoints for deployment statuses — GitHub Docs · Semantic conventions for CI/CD — OpenTelemetry · CI/CD resource conventions — OpenTelemetry

Follow a partial delivery without inventing a success

Illustrative correlation ledger. R1, M2, P3, H2, and D2 are labels, not real IDs or measurements.
Observed recordKnown conclusionMissing next proof
R1 attempt 1: source H2; check from expected app green on H2; build D2H2 passed the recorded check and produced D2Whether merged revision M2 contains that candidate and triggered delivery
Merge M2; production delivery P3 submitted with D2A delivery was requested for productionP3 terminal outcome and target revision
P3 response unknown; one target reports H2/D2, another has no version markerAt least one observed target has H2/D2; fleet state is incompleteOperation readback, remaining target identity, and critical journey
Critical journey not yet runLive behavior is unverifiedSafe probe result tied to the observed revision

The synthetic sequence is intentionally untidy. R1's green check applies to H2, while M2 is the later merge revision. The monitoring record should show how H2 and D2 relate to M2 instead of assuming the revisions are byte-identical. P3 may have changed one target even though its response was lost. Counting this row as a failed deployment, a completed deployment, or zero deployments would each overstate what is known. The honest state is partial or unknown delivery until P3 and the remaining target can be reconciled.

Once the provider reports a terminal result, read each relevant target's version marker and exercise the changed user path. If one target still serves an earlier revision, record mixed state and the traffic scope rather than collapsing the fleet to a single green value. If the marker cannot be read, that is an instrumentation gap; it is not proof of no delivery. Carry the gap into the alert and investigation queue with an owner. A release checklist handles the initial candidate and a rollback guide handles recovery choices; ongoing monitoring keeps the chain visible across every run and environment.

Sources: REST API endpoints for check runs — GitHub Docs · REST API endpoints for deployments — GitHub Docs · REST API endpoints for deployment statuses — GitHub Docs

Measure phases with named populations

Separate queue, build, test, delivery, and live verification. Queue delay begins when a run is eligible to execute and ends when work actually starts. Build duration covers the build task, not waiting for a runner. Test duration covers the selected test tasks, with canceled and skipped work labeled rather than silently counted as passing. Delivery duration needs a provider start and terminal result. Time to verified live ends only when the target revision and agreed journey are observed. Missing timestamps make a duration unknown; they do not make it zero.

Example metric definitions; choose time windows and exclusions for your own service before comparing results.
MetricDenominator or populationInterpretation guard
Queue delay distributionRuns admitted to the selected queue in the windowShow sample count and canceled-before-start runs separately
Build/test failure shareAttempts that reached the named build or test stageKeep infrastructure error, cancellation, and assertion failure distinct
Delivery terminal outcome shareProvider operations submitted for the named environmentUnknown or still-running operations remain their own states
Verified-live coverageDelivery attempts expected to produce a target readbackReport missing readbacks as missing, not verified or zero-time

A synthetic weekly report might say four of five eligible production delivery attempts have target readback, while one remains unresolved. That is a sample count in an invented example, not a product benchmark or a recommended target. If the team changes which events are eligible, the numerator and denominator must change together. Compare like environments and like trigger classes. For example, a pull-request check and a production deploy have different meanings, even when the same repository emits both. Use breakdowns to locate a bottleneck, then open the underlying run and operation records before proposing a fix.

Sources: Semantic conventions for CI/CD — OpenTelemetry · CI/CD resource conventions — OpenTelemetry · REST API endpoints for deployments — GitHub Docs

Use delivery metrics only after defining a deployment

DORA currently describes five software delivery metrics: change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. They are higher-level outcome measures than a per-stage queue chart. Define the deployment event from observed delivery to the intended production target, including how mixed or unresolved operations are classified. A merge count is not a deployment count, and a green build is not a successful change in production. The deployment-frequency calculator can explore a hypothetical rate; actual monitoring needs observed event data for the same service and time window.

For change lead time, keep the commit-to-production event pairing and expose the sample count. For change fail rate, define what counts as a production impairment and which deployment introduced it. For failed deployment recovery time, timestamp the impairment and the verified restoration, rather than using the moment someone clicks rollback. For deployment rework rate, identify unplanned corrective deployments by a documented classification rule. These choices require judgment; a dashboard cannot infer every causal relationship from exit codes. Publish the definitions beside the trend so teams know when a policy change makes two periods incomparable.

Do not convert an unobserved deployment into a zero-duration success to keep a graph complete. Show coverage: number of eligible events, number with complete source-to-target linkage, and number unresolved. Trend a metric only when its measurement quality is visible. A sudden improvement in lead time caused by missing target timestamps is a telemetry incident, not a performance win. DORA's metrics are useful for learning where delivery is constrained, but one number cannot substitute for the run-level evidence required to investigate a failure today.

Sources: A history of DORA's software delivery metrics — DORA · REST API endpoints for deployments — GitHub Docs · REST API endpoints for deployment statuses — GitHub Docs

Alert on an actionable transition or stale proof

An alert should identify a condition, an owner, and the next evidence to inspect. Examples include a required check that fails on the release candidate, a build that cannot produce the expected artifact, a provider operation that reaches a terminal failure, a target that reports the wrong revision after provider success, or a missing live readback after a defined waiting period. The waiting period is a service policy based on its observed rollout shape, not a universal number. A queue that is long but behaving as expected might warrant a capacity review; a release operation stuck beyond its policy needs an immediate investigation path.

For each alert, include repository, environment, run and attempt, source SHA, artifact digest, provider operation ID when known, last observed state, and the precise missing transition. Link to the relevant record rather than copying full logs into an alert. Keep log retention bounded and redact tokens, secret values, personal data, and provider credentials before storage or forwarding. Raw CI output can contain sensitive values even when a platform has masking rules. The purpose of the alert is to guide a safe read of evidence and a bounded next action, not to dump the entire environment into a notification.

Review alerts for two failure modes: noise from repeated symptoms of one operation, and silence when the instrumentation itself fails. Correlate updates for P3 into one incident rather than paging separately for every poll. Also watch a proof-coverage signal so missing deployment or target telemetry cannot make the system appear healthy. Close the alert only after the condition resolves or is explicitly handed to a recovery owner. The last line of the record should say what revision is live and whether the affected journey works; a green CI tile alone does not answer either question.

Sources: Semantic conventions for CI/CD — OpenTelemetry · CI/CD resource conventions — OpenTelemetry · REST API endpoints for deployment statuses — GitHub Docs

Sources and verification

Browse all resources