Deployment operations

Software deployment checklist: candidate to verified live

A deployment checklist should carry one candidate from exact source revision through build identity, required checks, destination approval, provider operation, and observed live behavior. Stop when any link is missing. A green CI result does not prove that a provider accepted the release, and a completed provider operation does not prove the intended revision serves a critical user journey. Record an uncertain provider response as unresolved, inspect the operation and target, then decide whether another action is needed.

Make the worksheet about one candidate

A useful software deployment checklist is a record of one candidate, one destination, and the evidence that connects them. Give the candidate an exact source SHA and a build artifact digest, then name the run and job that produced it. Record the required check name, expected source, evaluated revision, and result. The candidate is eligible to proceed only when those fields agree. A green icon on an older commit or a similarly named job is not evidence for the candidate being released. This worksheet is meant to be filled from actual run and provider records, not completed by assumption.

Next define the destination and the kind of proof it can return. The release operation may report accepted, progressing, succeeded, failed, or an unknown response after a network interruption. The target may report a deployed revision, image digest, release ID, or another durable version marker. Choose a marker that can be compared to the candidate. Also choose a critical user journey that the release is supposed to preserve: for example, a signed-in read or a payment-free task that exercises the changed path. A generic health endpoint can say the process is alive while the feature users need is broken.

The worksheet is not a guarantee of zero risk. It makes unresolved questions visible before they become a false success claim. Record who or what is responsible for each stop condition and how that actor will recover. For a small team, the owner may be the same agent or on-call operator across stages, but the next operation should still be explicit. A checklist without a stop state invites a rushed release when evidence is missing; a checklist with only preflight checks cannot prove what happened after delivery.

Sources: Deployment environments — GitHub Docs · Viewing deployment history — GitHub Docs · Deployments — Kubernetes Documentation

Before delivery: bind code, artifact, and policy

Start with source identity. If an agent changed the repository, select the reviewed merge revision rather than a nearby branch tip whose meaning can move. The build artifact should have a recorded digest and origin from that revision. An artifact name alone is weak identity, especially across attempts or matrix jobs. Check that the required verification result applies to the evaluated revision and expected reporter. If the repository uses merge queue, the integration candidate may be a queue revision rather than the pull request head. Stop if the build source, required check, and intended release candidate do not align.

Review the destination policy. GitHub environments can restrict which branches deploy, require review, and hold environment secrets until the configured protection rules pass. Those settings apply when the job actually references the environment. Provider-side permissions are a separate boundary: the deployment identity should be able to reach only the intended target. Document the environment name and provider account or cluster, without copying a secret value into the worksheet. A selected environment in a form is not itself proof that the provider authorized the candidate.

Check compatibility with data and the previous version before sending a release. Does the candidate change a schema, stored record, event contract, or external API in a way that the old and new versions cannot both tolerate during transition? Which migration is already applied, and which one must precede or follow code rollout? Identify a known previous candidate, but do not label it rollback-ready until its code and data assumptions have been reviewed. A prior image existing in storage is not proof that it can safely run against today's data.

Choose concrete prechecks rather than a long universal list. In the synthetic example below, the release cannot start without source H2, digest D2, required check verify from the expected app, production policy approval, a compatible data plan, and a target readback method. H2 and D2 are labels, not real cryptographic values. If any field is unknown, record the missing evidence and stop before requesting the provider operation. This is the cheapest moment to resolve ambiguity.

Sources: Deployment environments — GitHub Docs · Viewing deployment history — GitHub Docs · Deployments — Kubernetes Documentation

A three-stage release worksheet with stop conditions

Illustrative release worksheet; replace labels and owners with actual run and provider evidence.
Stage and evidenceStop conditionOwner and recovery step
Before: H2 source, D2 digest, required check source/result, production policy, data compatibility, prior candidate H1Any identity mismatch, missing check, unreviewed data change, or unavailable previous candidateRelease agent: resolve evidence or change plan before submission
During: provider operation P7, target production, submitted D2, progress stateProvider rejects, stalls, or returns an unknown responseDelivery operator: query P7 and target before another write
After: terminal P7, live H2/D2 marker, critical journey resultTarget marker differs, journey fails, or proof cannot be readOn-call owner: contain impact and choose a recovery action from observed state

The table separates three verdicts. Before says the candidate is eligible to attempt delivery. During says what the provider is doing with that request. After says what the target serves and whether the critical behavior works. Do not compress these into one green or red flag. A passed build can coexist with a rejected provider request. A provider success can coexist with a target that still serves the old revision, routes only some traffic to the new revision, or fails the changed user path. Each stage needs its own source of truth.

Write the recovery step beside the stop condition while the system is healthy. The exact mechanism depends on the provider and application. Kubernetes Deployments report rollout progress and completion through Deployment state, and their rollout history concerns the Pod template. AWS CodeDeploy can represent rollback as a new deployment with a new operation ID. Neither behavior is a universal undo switch for database changes or external side effects. The worksheet asks who will decide and what evidence they will inspect; the next guide can handle the more detailed rollback decision.

Sources: Deployments — Kubernetes Documentation · Redeploy and roll back a deployment with CodeDeploy — AWS · Stop a deployment with CodeDeploy — AWS

During delivery: reconcile an uncertain operation

Submit the candidate once and record the provider operation identifier, destination, artifact digest, and time. Follow that specific operation to a terminal result where the provider exposes one. If the network call times out after submission, the absence of a response is not proof that the provider rejected the request. A blind retry could start a second operation against a target already changing. Query the operation by its identifier and read the target state before deciding whether a second write is warranted. If the provider gives no usable identifier, record the uncertainty and inspect its deployment history and target rather than manufacturing success from a CI exit code.

A provider may report progress without declaring the rollout complete. Kubernetes documents that a Deployment can stall and surface a progress-deadline condition; checking rollout status is distinct from issuing the update. CodeDeploy documents that stopping a deployment can leave instances in an indeterminate state in some cases. These are provider-specific examples of why a sent command, a successful API response, and terminal delivery evidence should be kept separate. Use the status semantics of the actual platform; do not import a Kubernetes success condition into an unrelated service.

If the release is paused or stopped, preserve the candidate and operation evidence. Do not assume the previous version is still fully live. Some targets can be mixed during rollout, and a stopped operation can leave partial effects. The next safe operation may be to hold traffic, finish a repair, redeploy a known candidate, or roll back according to the provider and data state. The immediate checklist task is narrower: identify what is running now and who has authority to act next.

Sources: Deployments — Kubernetes Documentation · Stop a deployment with CodeDeploy — AWS · Viewing deployment history — GitHub Docs

After delivery: verify the live revision and user path

When the provider reports a terminal success, read back the target's version marker and compare it with H2 and D2. GitHub's deployment history can show associated commits, workflow logs, and deployment statuses, which are useful links in the record. That history is not a substitute for observing the destination itself. A service may expose a release label or image digest; a web app may need a version endpoint or another approved deployment record. If the target cannot report a version, say that the live-revision proof is missing and make it a design gap to close.

Then exercise the critical journey that the change could break. A health probe is useful for liveness and readiness but may not traverse authentication, a write path, or the changed integration. Select a safe test with a clear expected result and an environment-appropriate account or fixture. Record which revision served the request if the system can expose it, especially while traffic shifts. This is a bounded post-deploy verification step, not a claim that every future user action has been tested. Continued monitoring belongs to the operating plan after the initial release verdict.

Close the worksheet with one of three honest states: verified live, failed with a known target state, or unresolved. Verified live needs the intended revision and critical journey result. Known failure needs the observed target state and a recovery owner. Unresolved means a required readback or provider result is missing; it calls for investigation, not an immediate repeat of the deploy. Keep the run, check, artifact, provider, and target references together so the next agent can resume without guessing which candidate reached the environment.

Sources: Viewing deployment history — GitHub Docs · Deployments — Kubernetes Documentation · Stop a deployment with CodeDeploy — AWS

Sources and verification

Browse all resources