Agent delivery workflows

MCP for CI/CD: a safe build-to-deploy loop

Use MCP for CI/CD as an evidence loop: identify the exact revision and run, read the failure before changing code, make one bounded repair, rerun the checks, follow the delivery, and verify the revision that actually reached production. A tool response is a source of facts or an operation result; it does not replace repository permissions, a release policy, or proof that the deployed code is live.

Begin with the run, not a proposed fix

A coding agent needs to answer four questions before touching a failed change: which revision was tested, which attempt failed, where it ran, and what evidence supports the failure. The same repository can have several pushes and several retries in flight. A green result from an older commit cannot clear a newer red one. A queue entry is not a failed test. Read the platform's run status and the relevant log or diagnosis, and keep the returned identifiers together in a small evidence record. If the result lacks a revision, resolve that ambiguity before acting.

MCP gives a client a standard way to discover a server's tools and call them by name with schema-defined arguments. It does not prescribe a universal CI/CD tool catalog. One server may expose a status read and a log read; another may add a rerun or a long-running follow operation. Discover the connected server's actual tools and inspect the output fields before assuming a workflow exists. Treat a tool description as a description of capability, not evidence that a particular run succeeded.

For Agentic Pipeline, the listed names include agentic_pipeline_status, agentic_pipeline_log_get, agentic_pipeline_diagnose, agentic_pipeline_rerun, and agentic_pipeline_follow. The practical order is status for the exact pushed revision, diagnosis of the identified failure, a narrow log read if more context is needed, and a rerun or new push only after a repair. The platform also exposes agentic_pipeline_platform_health so an agent can distinguish a platform incident from a repository defect. The actual response, rather than a remembered tool name, decides the next step.

Sources: Tools — Model Context Protocol specification

A worked failure ledger that prevents stale-green decisions

Imagine a branch push called revision R1. These labels are illustrative placeholders, not real commit hashes, run IDs, or MCP call arguments. The platform reports that run A, attempt one, on a Linux build environment failed in a pricing test. The log excerpt says the discount cap assertion expected 9,500 cents payable but received 9,000. An agent should attach the excerpt to run A and R1, then compare it with the product requirement. It should not assume that an unrelated prior green run proves R1 safe.

Illustrative evidence ledger; replace every label with identifiers returned by the connected CI/CD system.
EvidenceNext operationProof needed
R1 · run A · attempt 1 · Linux test · cap assertion failedInspect requirement and repair the cap logicDiff changes only the bounded pricing behavior
R2 · run B · attempt 1 · Linux test · checks passedFollow R2 through merge and deliveryRequired checks belong to R2, not the older R1
Merge revision M · deployment reported completeRead the release result and live revisionVerified live revision equals the intended merge revision

After fixing the cap, the agent creates a new revision R2 and observes run B. It must not overwrite run A in the ledger: the failed attempt explains why the change happened, while run B proves the replacement revision. If run B goes green, delivery still has a separate question. A deployment can fail after tests pass, can be skipped under a declared rule, or can finish without enough evidence to establish the live revision. The final row therefore records the merge revision and the system's actual release evidence rather than merely copying the CI verdict.

This ledger is intentionally small. Keep the repository, revision, run or task identifier, attempt, environment, terminal state, and one relevant evidence link or excerpt. Add a new row when the revision or attempt changes. The discipline makes an agent's next operation auditable: a reader can see why it edited code, why it reran, and why it called delivery complete. If a tool result cannot supply one of those facts, state the missing fact and obtain it from the appropriate surface before making a stronger claim.

Make one bounded repair and prove the resulting revision

A useful repair begins with the failing requirement, not with a command to make the pipeline green. Confirm that the failed assertion represents intended behavior; then change the smallest relevant code path and retain the assertion. If the failure is an infrastructure outage or a temporarily unavailable dependency, changing application code may add risk without addressing the cause. If the run is queued, it has not judged the change at all. Read the platform state and the specific failure before deciding whether to edit, wait, or rerun.

Prove the repair against the exact new revision. Local tests can shorten the loop, but the pushed run is the evidence for the revision that teammates will review. Record any retry as a separate attempt rather than treating it as a continuation of the original failure. A rerun of unchanged code can be appropriate for a transient problem; it cannot prove that a code edit was tested unless the run points to the edited revision. Where a repository has a required check, verify which revision that check actually covers before allowing the merge.

Agentic Pipeline defaults to Auto Profile when a repository has no recipe file, while a repository-owned recipe at the triggering revision takes precedence. The agent's work is still to supply meaningful assertions for the application and to repair its own change. The pipeline can run supported checks and report their results; it cannot know an unstated product rule. Keep the repair scoped to the failed behavior and let the next run show whether the change resolved it without breaking the rest of the repository.

Sources: About protected branches — GitHub Docs

Follow delivery until the outcome is terminal

Long-running work needs a stable handle and a terminal result. The current MCP Tasks extension, identified as io.modelcontextprotocol/tasks, lets a capable client receive a task handle and retrieve later status when both sides support it. The extension defines working, input-required, completed, failed, and cancelled states. This is an optional negotiated capability, not a promise that every MCP server or client implements tasks. Other servers may return a regular tool result or provide their own documented follow mechanism.

For Agentic Pipeline, agentic_pipeline_follow tracks the pushed revision through its delivery journey. A compatible client may receive an MCP task; a client without that extension has the documented agentic_pipeline_task_get path to resume the follow operation. Retain the returned handle, read the terminal verdict, and distinguish a running journey from a completed one. A momentary lack of new log lines is not a terminal outcome. When an operation asks for input or reports a failure, handle that state explicitly rather than summarizing it as success.

After merge, a green check answers whether the required checks passed for a revision. It does not answer whether a deployment occurred or what code is live. Read the delivery outcome and its stated basis. A declared shipping-disabled policy can explain why no deployment was requested; report that policy and stop short of a live claim. A missing or unresolved delivery target leaves the delivery evidence unresolved, even when CI is green. If a release completed without proof of the live revision, say that verification remains open. The words in the final report should be no stronger than the evidence returned.

Sources: Tasks — MCP Tasks extension · About protected branches — GitHub Docs

Keep tool access and release authority separate

A model may be able to discover and propose a tool call, but the connected service and its authorization policy decide what that call may do. MCP's authorization specification requires protected servers to validate access tokens and their intended audience. A tool's annotation, including a read-only hint, is metadata for client behavior; the tools specification tells clients not to trust annotations from untrusted servers. A label cannot grant repository permission, approve an irreversible action, or prove that a returned result belongs to the intended tenant.

Give the agent the narrow access needed for the task: read the run, inspect the failure, and request only the authorized repair or rerun. Check the repository and revision in each response, especially when an agent operates across multiple projects. If a write requires approval under your team's release policy, obtain it before the write. The MCP transport does not supply that policy by itself. Keep credentials and private log contents inside the approved service boundary, and avoid copying a full log into a public report when a short, relevant excerpt will explain the decision.

The useful outcome is a traceable claim: this revision failed for this reason; this bounded change produced a new revision; its checks passed; the delivery reached a terminal state; and the live revision was verified, or verification remains open. That sequence is valuable even when the connected server uses different tool names or lacks optional task support. The test is whether the agent can show the evidence for each transition, state uncertainty honestly, and stop where its authority ends.

Sources: Authorization — Model Context Protocol specification · Tools — Model Context Protocol specification · Tasks — MCP Tasks extension

Sources and verification

Browse all resources