CI/CD practice

CI/CD best practices for small teams shipping with coding agents

The most useful CI/CD practice for a small team using coding agents is to keep one change's evidence connected from requirement to live behavior. Define the user rule, make a small change, add the smallest durable proof that could catch a real mistake, and run a fast, trusted build. Carry the exact source and artifact identities through policy checks and a bounded provider operation. Read back the live revision and exercise the changed path, then leave a recovery owner and clear unresolved state. Measure feedback and delivery costs from real events; do not trade away the proof needed to ship safely.

Start with a release journey, not a list of tools

A small team can assemble many CI/CD tools and still lose the central question: did the intended change reach users safely? Draw the path for one ordinary change. Someone or an agent states the behavior, edits a small part of the codebase, runs focused proof, submits it for review, integrates it, builds a candidate, passes release policy, asks a provider to apply it, and reads the live target. Each arrow has an input, an output, and an owner. If an arrow is invisible, a green tile farther downstream may hide a missing release or a different revision. The best practices in this guide improve those arrows rather than prescribe a universal vendor stack.

Imagine an illustrative workspace product adding a pause control for digest emails. The requirement is narrow: while a member's digest setting is paused, the scheduled digest should not be sent; account security messages still may be sent. A coding agent can implement the setting, but the team needs evidence for the boundary between digest and security messages. The change might pass a unit test yet fail after a queue worker uses a different setting read path. It might deploy successfully yet leave one worker on an older revision. Those possibilities define the release journey's proof points. No customer, benchmark, production run, or feature of Agentic Pipeline is being represented by this example.

DORA describes continuous delivery as keeping software in a deployable state so it can be released on demand. DORA's continuous integration guidance calls for frequent integration, quick automated feedback, and authoritative packages from the CI build. These are useful goals, but the team's implementation should be explicit about what a pass proves. A merged change is not automatically a deployed one. A deployable candidate has not necessarily been applied. A provider's successful operation does not by itself show that the live path behaves as intended. Write those distinctions into the agent handoff so responsibility does not disappear at a job boundary.

Sources: Continuous integration capability — DORA · Continuous delivery capability — DORA · Deployment automation capability — DORA

Make the change small enough to review and repair

Give the agent a behavioral boundary before it edits. In the digest example, the desired rule identifies who can pause, which scheduled message is affected, and which messages are outside scope. Ask for a concrete before-and-after example: a paused member's digest is suppressed, while the same member's security notice is unaffected. That example is an oracle for review and testing. It also limits implementation sprawl. A broad instruction such as improve notification settings invites an agent to change storage, templates, scheduling, and preferences at once, making both cause and rollback harder to reason about.

Integrate the smallest complete slice that preserves existing behavior. If a new setting needs a stored field, decide how old workers and new workers coexist during rollout. If there is no safe compatibility path, name that as a release dependency before merging. A small diff does not guarantee safety, but it makes the changed contract easier to inspect and the failing revision easier to isolate. DORA connects small, frequent integration with faster feedback. For an agent, that means the next branch should not quietly accumulate a second feature while the first one waits for a test result. Keep source identity and review scope aligned.

A focused review checks both code and the claim the agent will later make. Did the implementation actually branch on the digest preference at the point where the worker sends? Does a different route send the same digest? Could the new field default incorrectly for existing members? These are review questions grounded in the requirement, not a ritual count of files changed. The agent should carry unresolved assumptions into the release record rather than turning them into implied approval. A durable decision note can be brief: the changed rule, the selected proof, the data compatibility assumption, and the exact revision that reviewers accepted.

Sources: Continuous integration capability — DORA · Test automation capability — DORA

Add the smallest durable proof that tests the rule

Choose tests by the failure they could catch. A pure function that decides whether a digest is eligible can be tested with paused and active settings, but that alone might miss a worker that never calls it. A narrow integration test at the scheduled-send boundary can prove the queue path respects the setting without launching the entire application in a browser. A post-deploy probe can confirm the live worker revision and a safe representative journey. The exact mix depends on the architecture; a fixed percentage of unit versus end-to-end tests is not a substitute for identifying the highest-risk seam. The separate guide on testing AI-generated code handles oracles and counterexamples in more depth.

A meaningful test should fail for a plausible wrong implementation. In the example, a test that merely sees a preference field in a response would pass even if the worker ignores it. A better counterexample sets pause to true, invokes the scheduling boundary with an eligible digest, and observes that no digest send is requested; a companion case confirms active digest behavior and a separate case ensures security messages are not suppressed. These cases are illustrative, not executable code or a claim about a particular mail provider. The agent can then run the narrow test against a deliberately wrong branch or inspect its assertion to make sure it would catch the omission.

Keep tests durable and proportionate. If the rule will be easy to break during the next refactor, automate it in the repository's normal suite. If a one-time visual detail can be checked quickly and has no durable contract, a new brittle browser test may add more maintenance than proof. DORA's test-automation guidance emphasizes fast, reliable suites and feedback throughout delivery. Do not turn that into an arbitrary universal deadline for every service; measure your own feedback path and find the stage that delays useful results. An agent handoff should name the exact tests that ran, tests that did not run, and the requirement each test covers.

Sources: Test automation capability — DORA · Continuous integration capability — DORA

Keep feedback fast without hiding its cost

Break the pipeline into queue, setup, build, focused tests, broader tests, and release checks. When a run feels slow, measure which stage consumes time and which failures each stage finds. A long queue is a capacity or scheduling problem, not evidence that the tests themselves are slow. A repeated dependency download may be a caching opportunity, but a cache hit is not proof that the lockfile or build output is correct. A large matrix can shorten wall time while increasing total compute and log volume. The useful optimization is one that improves time to a trustworthy answer at an acceptable operating cost.

For a small team, start with the few gates that can change a release decision. Put fast, deterministic checks early so a source error stops wasted downstream work. Move slow environment-specific proof to the stage where its result is meaningful, but do not remove it merely to make the dashboard green. If a test is flaky, record its failure rate and causal patterns before deciding on a retry policy. A rerun can distinguish a transient infrastructure fault from a deterministic source failure; repeated red-until-green attempts do not repair either. The CI cost, matrix budget, flaky probability, and parallelization tools linked below are for exploring hypothetical tradeoffs, not substituting for measured run data.

Measure a baseline from real runs with sample counts: median or distribution of time from trigger to first useful failure, time to all required checks, total work consumed, and the fraction of attempts that must repeat. Separate canceled attempts and missing timestamps from completed ones. If an optimization speeds up the typical run but delays the rare failure that blocks production, the team may have improved a chart while worsening the decision loop. DORA's delivery research ties fast feedback to keeping software deployable, but there is no credible single timeout or spend target for every codebase. Set a service-specific threshold from observed needs and revisit it when the workload changes.

Sources: Continuous integration capability — DORA · Test automation capability — DORA · Continuous delivery capability — DORA

Carry a bounded, trusted candidate through handoffs

The release candidate should be an identifiable object, not a phrase such as latest green. Record the source SHA, run and attempt, required check name and expected reporter, build digest, target environment, and release policy result. If the change was tested at a pull-request head but merged into a different revision, show the relationship and verify the revision that the release rule requires. A passing check on one commit cannot be silently transferred to another. The artifact digest binds the selected bytes; a friendly artifact name or a cache key alone does not. Keep these identifiers together so a second agent can continue without guessing which output belongs to which run.

Treat untrusted inputs as untrusted through the release boundary. A pull request can alter source, test fixtures, and artifact contents. The privileged delivery step should not execute arbitrary code or consume an artifact from a less-trusted run merely because it was uploaded under the expected name. Rebuild or verify the candidate under the trusted policy that governs production, and compare the source and build identity with the expected repository and workflow. SLSA's artifact-verification guidance focuses on checking provenance against expectations; provenance can support an origin claim but does not prove that the application is correct or safe for users.

Bound what the handoff authorizes. The build stage may need read access to source and a write to a test result. The release stage may need a credential for one destination. In GitHub Actions, an environment can hold a job behind protection rules and restrict access to its secrets when the job references that environment. An OIDC token request needs `id-token: write`, but that permission alone does not grant provider write access; the cloud role's trust conditions and permissions determine what it can do. Other CI/CD systems have different mechanisms. The principle is to make the release identity and destination policy explicit, then verify that the actual operation used them.

Sources: Verifying artifacts — SLSA specification · Deployment environments — GitHub Docs · OpenID Connect reference — GitHub Docs

Observe delivery and the live path separately

Submitting a deploy request is a state transition, not the end of the journey. Record the provider operation ID, submitted digest, environment, and current result. If the request times out after submission, the operation may still be running; query the provider and target before sending a duplicate. If a provider reports success, read the target's revision marker and compare it with the candidate. A rolling service may briefly serve multiple revisions, so a single host's marker may not describe the whole target. The release verdict should say what was observed and where, including unresolved slices. A generic healthy-process response cannot stand in for the changed user rule.

The digest-email example needs a safe live check. After the target reports the intended revision, exercise an approved fixture or non-user-impacting path that reaches the scheduled-send decision and confirms a paused digest is suppressed while an allowed message path still works. If the platform cannot safely probe that behavior in production, define a narrower live signal and state its limit. The team may combine a target revision readback, worker execution record, and a monitored canary account. These are design options, not promises that a particular product offers a universal probe. The outcome should be tied to the specific candidate and target, not just a health dashboard's current color.

DORA's deployment-automation guidance includes a deployment test as part of the delivery process. That is compatible with a broader operating check after the provider finishes. The guide on software deployment checklists walks through candidate, operation, and target evidence in detail, while the monitoring guide covers ongoing trends. Here the practical rule is simple: do not collapse CI, provider, and live proof into one boolean. An agent should be able to state integrated and tested, submitted, applied, or verified live with the record supporting each claim.

Sources: Deployment automation capability — DORA · Deployment environments — GitHub Docs · Continuous delivery capability — DORA

Name recovery ownership before something breaks

A safe release path includes a known stop state and a recovery owner. Before delivery, decide who can pause further writes, what target readback is available, and which previous candidate may still be compatible with today's data. A previous binary is not automatically a safe rollback target after a schema contraction, a message-format change, or an external side effect. If a release fails midway, preserve the first operation and observed target state before choosing a new action. The rollback guide treats that decision in depth; this pillar's point is that the owner and evidence path must exist before the team is under time pressure.

Recovery is not always an undo button. The safest next step could be to hold traffic, repair the current candidate, or redeploy an earlier compatible one, depending on what is live. A code rollback does not erase data written by the new code or retract a notification already sent. If the digest change accidentally sends a message, restoring the previous worker cannot unsend it. The team needs a separate user-impact and reconciliation decision. A provider operation may also return an unknown status; uncertainty calls for investigation, not an immediate repeat request. Give an agent authority to gather evidence and a clear boundary for production writes.

After recovery, verify the new target revision and the critical path that failed. Keep the original and recovery operation IDs, artifact digests, and test results together so the next incident review can distinguish a bad build, a policy gap, a provider failure, and a missing live signal. Ongoing monitoring should watch delayed effects, but it cannot replace a bounded post-recovery proof. The value of this record is not a promise that every failure will be quick. It is that the team can make the next decision from observed state rather than from a green check that belongs to an earlier stage.

Sources: Deployment automation capability — DORA · Continuous delivery capability — DORA

Improve one bottleneck at a time from measured evidence

Once the path works, review it over a real sample of changes. Count eligible changes, how many reached each stage, how long each transition took, and how many ended without a target readback. Missing telemetry is unknown, not zero duration or zero failures. DORA's current delivery measures include change lead time, deployment frequency, failed deployment recovery time, change fail rate, and deployment rework rate. Those are outcome measures, not a substitute for run-level diagnosis. Define what counts as a production deployment and failure for this service before comparing periods or using a calculator's hypothetical output as observed performance.

Pick the next improvement from the largest evidenced gap. If many agent changes fail at the same requirement boundary, improve the oracle and review prompt. If most delay sits in queue, revisit capacity or scheduling. If production operations succeed but live revision proof is absent, add a target marker before adding more release automation. If credentials are broad, narrow the release identity. If recovery repeatedly waits on one expert, codify the read path and handoff so another authorized agent can diagnose safely. A long generic checklist will not identify these priorities; a complete evidence chain will.

Agentic Pipeline's public promise is a build-and-ship loop: Auto Profile is the default when a repository has no recipe, testing pushes and shipping eligible merges, while an existing repository-owned recipe remains authoritative. Treat that as the intended path and inspect the actual run and delivery evidence for each change. Do not infer a universal one-click rollback, dashboard field, or production outcome from a product description. For a small team, the practice is to keep the agent's claim bounded to what the current run proves, then make the next missing proof visible. The related testing, MCP, deployment, and monitoring resources below go deeper on those handoffs.

Sources: A history of DORA's software delivery metrics — DORA · Continuous integration capability — DORA · Continuous delivery capability — DORA · Agentic Pipeline: Public feature guide

Sources and verification

Browse all resources