Deployment operations

Deployment rollback: recover code without guessing about data

A deployment rollback is a new recovery decision, not a reversal of time. First identify the release operation, actual target revision, traffic state, and any data or external effects. Contain impact while that state is uncertain. Restore earlier code only if it can safely read and write the data that exists now; otherwise choose a bounded forward repair or another provider-specific recovery path. After the action, verify the target revision and the affected user journey. A provider success alone does not close the incident.

Decide from the target state, not the deploy button

The first question during a failed release is what is serving traffic now. A deploy command may have been accepted, timed out after acceptance, failed before changing any target, or changed only part of a target. Preserve the source revision, artifact digest, operation ID, destination, and time before another write. Read the provider's operation status and the destination's version marker. If the answer remains unknown, label it unknown. Repeating the same deploy or invoking rollback without knowing what changed can create a second operation on top of the first.

Contain impact according to what the system actually supports. A team may pause a rollout, stop new traffic to a bad slice, disable a feature behind an existing control, or hold further deployments while it investigates. These are examples of possible controls, not universal commands. Check whether the control itself has side effects and whether the provider has already moved traffic. A failed probe on one new instance and a failed critical journey for all users imply different urgency and may call for different containment. Record the scope and the person or agent authorized to choose the next write.

Choose among rollback, forward repair, or continued containment only after three checks: the current target state is known, the candidate action is compatible with current data and external contracts, and there is a way to prove the result. The last previously successful binary is not automatically a known-good recovery candidate. It might depend on a database field already removed or an external service contract that changed. Conversely, a small correction to the current version might be safer than restoring old code when the data transition has passed its reversible phase. The decision should name the evidence, not just the preferred button.

Sources: Deployments — Kubernetes Documentation · Stop a deployment with CodeDeploy — AWS · Expand-and-contract migrations — Prisma Documentation

Work through a data-compatible rollback window

Consider an illustrative account record with an old field `display_name` and a new pair, `given_name` and `family_name`. Version V1 reads `display_name`. An expansion adds the two new fields without removing the old one. Version V2 initially writes both representations and reads the new fields with a reviewed fallback. A backfill reconciles older rows, and checks confirm that both V1 and V2 can handle records written during the transition. If V2 then fails because of an unrelated rendering bug, returning to V1 can be a viable code recovery action while the old field is still maintained. The team still needs to verify the data assumptions and the live target.

Now change the scenario: the contract step has removed `display_name`, or V2 has stopped maintaining it while new records are created. V1 can no longer be assumed safe, even if its image digest is available. Restoring V1 may turn a visible rendering bug into failed reads or writes. The safer action may be to contain the affected path and release a tested V2 repair, or restore compatibility before moving code. This is why the rollback window belongs in the migration plan before the release. Prisma's expand-and-contract guide describes adding and backfilling the new representation before removing the old one; the example here is our own scenario, not a Prisma command sequence.

Code rollback also does not undo data already committed by the new version. A corrected application cannot infer which records to erase merely because an older artifact returned. Nor can it automatically cancel an email, refund a payment, or reverse a message delivered to another system. Those effects need their own reconciliation or compensation rules, grounded in provider records and business policy. During recovery, avoid a broad data reversal based only on deploy time: users may have made legitimate changes after that point. Identify the precise affected operations, preserve their IDs, and separately decide what data action is justified.

Sources: Expand-and-contract migrations — Prisma Documentation · Redeploy and roll back a deployment with CodeDeploy — AWS

An illustrative recovery decision record

Synthetic example; V1/V2, D1/D2, and P7/P8 are labels, not real revisions or provider commands.
ObservationDecision testNext action and proof
P7 submitted V2/D2; response lost; target revision unknownHas P7 completed or changed any target?Hold a second write; query P7 and read target/traffic state
Target partly V2/D2; critical journey fails; old field still maintainedCan V1/D1 read and write current records?Contain bad slice; if compatibility is proven, submit provider-specific recovery as P8
Target partly V2/D2; old field removed or no longer maintainedWould V1/D1 break current records?Do not restore V1; contain and test a compatible forward repair
P8 reports success or repair reports successDoes the target serve the intended digest and pass the affected journey?Read target marker and exercise the journey before closing recovery

The first row is deliberately unresolved. A lost response from P7 is not evidence that V2 never reached production. Querying operation and target state may reveal no change, a completed release, or a mixed rollout. The next row permits code rollback only while V1 remains compatible with data V2 produced. The third blocks that same action after a destructive contract. The last row separates the provider's terminal status from actual recovery. Use a fresh operation ID for the recovery attempt where the provider offers one, and keep P7 and P8 in the same incident record.

The table is a decision aid, not an automated runbook. Its rows do not claim that every deployment platform exposes traffic slices or a pause control. Replace the synthetic markers with the real provider's operation identifiers, source SHA, artifact digest, environment, and target readback. If the system cannot reveal which revision is serving, recovery proof has a known gap; saying so is safer than declaring success from a completed API request. The release checklist covers candidate eligibility before delivery; this record begins only when the live state or intended behavior is in doubt.

Sources: Deployments — Kubernetes Documentation · Stop a deployment with CodeDeploy — AWS · Expand-and-contract migrations — Prisma Documentation

Know what the provider's rollback actually changes

For a Kubernetes Deployment, rollout history tracks revisions created by changes to the Pod template. Scaling alone does not create such a revision. Returning to a previous Deployment revision changes the workload template and starts a rollout; it does not restore a database schema, reverse a queue message, or prove that the application journey works. Inspect the selected template and image, the resulting workload status, and the actual service path. A stalled Deployment can report a progress-deadline condition while the controller continues trying; treat that signal as a reason to investigate, not a complete account of user impact.

For AWS CodeDeploy, rollback redeploys an earlier application revision as a new deployment with a new deployment ID. AWS also documents limits to cleanup: its EC2/on-premises agent does not reconcile arbitrary actions taken by prior lifecycle scripts. Stopping an EC2/on-premises deployment can leave instances in an indeterminate state. Therefore the recovery operation itself needs tracking, and files, processes, and data outside the revision may need separate inspection. These details describe CodeDeploy, not every provider. A platform's rollback label is shorthand for its own documented operation, never proof that all effects were undone.

Do not paste a generic rollback command into a production procedure without the provider, target, revision, and authorization context. A command that chooses the prior workload template on one platform might create a fresh deployment on another. Instead record the proposed action as an intent: restore compatible candidate V1/D1 to environment production, or deploy compatible repair V3/D3, after observing P7. Attach the actual reviewed provider operation when it is chosen. This preserves the distinction between advice in a guide and an executed production change.

Sources: Deployments — Kubernetes Documentation · Redeploy and roll back a deployment with CodeDeploy — AWS · Stop a deployment with CodeDeploy — AWS

Close recovery with revision and behavior evidence

A recovery action has three useful verdicts: accepted, terminal at the provider, and verified at the target. Record the new operation ID and follow it to a terminal result when the provider supports that. Read back the deployed artifact digest or another approved version marker from the destination, including each active slice if traffic is mixed. Then run a safe check of the critical journey that failed, plus a data compatibility check relevant to the migration. A generic process health response may miss a broken read, write, or external integration. Stop with an unresolved verdict if the revision or journey cannot be observed.

If recovery chose V1, confirm it remains able to read new records and that its writes preserve fields needed by a later V2 or V3 rollout. If recovery chose V3, confirm the repaired path and the unchanged paths most exposed by the fix. Record any external side effects still awaiting reconciliation separately from the code verdict. A successful code release does not close a payment, notification, or data correction issue by implication. Name the next owner and the evidence required to close each remaining item.

Finally, preserve the sequence: initial source and digest, P7 outcome, observed target state, compatibility decision, chosen recovery candidate, P8 outcome, target readback, and critical journey result. That sequence lets the next agent explain why rollback was safe or why a forward repair was necessary. It also exposes a missing proof link for future design work without claiming a one-click recovery guarantee. Once the immediate user path is stable, ongoing monitoring can look for delayed effects; it should not replace the bounded verification that closes this recovery attempt.

Sources: Deployments — Kubernetes Documentation · Redeploy and roll back a deployment with CodeDeploy — AWS · Expand-and-contract migrations — Prisma Documentation

Sources and verification

Browse all resources