Testing agent changes
Testing code written by AI agents
Test code written by an AI agent against a requirement that exists independently of the implementation. Run the repository's normal checks, then add a small set of examples at the requirement's boundaries and at least one counterexample the agent did not select. Review the changed behavior and its failure paths before treating a green run as evidence to ship.
Start with a requirement the code cannot redefine
AI code testing begins before the test runner starts. Write the expected behavior in terms of an input, an observable result, and a constraint that matters to the user or system. A coding agent can produce a plausible function and an equally plausible test for that function. If both were inferred from the same mistaken assumption, passing tests only show agreement between two artifacts. The original request, a reviewed policy, or an existing contract must decide the expected result. Record that source where the next reviewer can see it.
Separate semantic requirements from implementation details. A requirement such as ‘a discount never exceeds five dollars’ is about the result; the shape of a conditional branch is only one way to achieve it. A test that asserts a helper's internal call sequence can pass while the final invoice is wrong. Conversely, a refactor may preserve the invoice while breaking a test that was coupled to private structure. Choose assertions on externally meaningful values at the narrowest stable boundary that can express them.
For each changed behavior, ask who supplied the oracle. An agent-generated test can still be useful, particularly for routine edge cases, but have someone or something independent validate the key expectation. That could be a product rule quoted from a reviewed specification, a known fixture from an older implementation, or a hand-calculated example. Independence matters more than the identity of the test author. Do not promote a test just because it looks comprehensive or because its prose sounds confident.
Worked example: a capped discount that passes the happy path
Consider an illustrative order function that takes a nonnegative subtotal in cents and applies a ten percent promotional discount, capped at 500 cents. The requirement is the oracle: payable cents equal the subtotal minus the smaller of ten percent of the subtotal and 500 cents. Suppose an agent implements the ten percent calculation but omits the cap. Its first test uses a 1,000-cent subtotal, expects 900 cents payable, and passes. That test is correct yet cannot distinguish the intended rule from the defective implementation.
| Subtotal | Required payable | Uncapped implementation | What it reveals |
|---|---|---|---|
| 1,000 cents | 900 cents | 900 cents | Both implementations agree below the cap |
| 5,000 cents | 4,500 cents | 4,500 cents | The discount reaches the cap |
| 10,000 cents | 9,500 cents | 9,000 cents | The uncapped version is 500 cents too low |
The 10,000-cent case is a counterexample to the agent's implementation. Its expected 9,500-cent result comes from the cap, not from copying the function's output. Add it as a regression test with a name that states the policy. The exact cap boundary also deserves a check: at a 5,000-cent subtotal, ten percent equals 500 cents. Inputs just above it catch a branch that applies the cap one step late. The point is not to collect many inputs; it is to choose inputs that separate plausible interpretations of the rule.
The illustrated subtotals make ten percent a whole number of cents, so their arithmetic is unambiguous. A real pricing system has more decisions: rounding, taxes, stacked promotions, refunds, invalid subtotals, and currency rules. Those require their own approved requirements before a test can assert an answer. Do not infer a tax order or rounding policy from this miniature example. When the specification is silent, mark the gap and obtain a decision from the owner of that behavior rather than letting generated code silently establish policy.
Choose counterexamples that challenge the change
A useful counterexample targets a way the change could be wrong while still passing its first demonstration. For a bounded value, test below, at, and above the boundary. For a permission check, vary the actor and the resource owner independently. For a parser, try malformed input that resembles a valid case. For a state transition, retry an operation or change the order of events. These are prompts for reasoning, not a fixed checklist: select the cases that follow from the actual contract and the changed code path.
Authorization deserves special care because the visible success case may say little about denial. If a generated endpoint allows a project member to read a document, also attempt to read a document owned by another project with the same role. Then try a different role against the original document. OWASP's authorization testing guidance explicitly distinguishes cross-user and cross-role access. An allowed case and a denied case should be asserted separately; an HTTP response that merely has the right shape is not proof that data stayed within its boundary.
Property-based testing can widen the search once the property itself is trustworthy. For the discount, one property is that the discount lies between zero and 500 cents for all valid subtotals. Another is that payable never exceeds subtotal. Hypothesis describes generating inputs within a declared range and finding cases that violate a stated property. Yet a property that omits the cap would preserve the original mistake. Review the invariant first, constrain generated data to the real domain, and retain any discovered counterexample as an explicit regression case.
A second technique is to introduce a deliberate fault and see whether the tests notice. Mutation testing tools such as PIT change program behavior, run tests against the changed version, and report whether a test detects the change. A surviving boundary mutation suggests that the suite may not distinguish the policy's critical branch. It does not automatically prove a production defect: some mutants are equivalent or outside the feature's meaningful domain. Use a survivor to ask a precise question about your assertions, not as a single score that decides release safety.
Sources: Bypassing Authorization Schema — OWASP WSTG · Introduction to Hypothesis — Hypothesis documentation · Basic Concepts — PIT Mutation Testing
Turn a passing run into a reviewable decision
Run the repository's existing checks on the exact revision under review. Compilation, type checks, tests, static analysis, and the relevant application flow each answer a different question. GitHub's guidance for reviewing AI-generated code starts with automated tests and static analysis, then asks reviewers to verify the change against project intent. Preserve the failing output when a check fails; ask the agent to repair the failure without deleting a meaningful assertion or narrowing the test until it becomes trivially green.
Review the diff as a behavioral change. Identify the input surface, altered outputs, persistent side effects, error handling, and permissions. Compare those with the independent requirement and the counterexamples. Then check that the tests would have failed against the plausible wrong version. In the discount example, the 1,000-cent assertion would pass on both versions; the 10,000-cent assertion separates them. A reviewer can understand that evidence without treating a coverage percentage as a proxy for correctness.
Keep the feedback loop small enough that failures remain attributable. If an agent changes pricing, authorization, and deployment configuration in one batch, a single green result obscures which requirement each test protects. Ask for a bounded change, an explanation of each new assertion's source, and the command and revision that produced the result. For a user-visible flow, exercise that flow in a representative environment as well. A unit test can establish a calculation while a browser or API path reveals a missing wiring step.
A green check is a necessary delivery signal only for the behavior it actually examines. A required status check can prevent a protected-branch merge until its configured checks pass, as GitHub documents, but it cannot invent the absent cap test. Agentic Pipeline's role in this loop is to run repository checks and expose their results to the coding agent; the team still owns the product rule and its assertions. When a counterexample fails, repair the implementation or clarify the requirement, rerun on the resulting revision, and record the evidence before shipping.
Sources: Review AI-generated code — GitHub Docs · About protected branches — GitHub Docs
Sources and verification
- Review AI-generated code — GitHub Docs
GitHub recommends automated tests and static analysis, followed by review of generated code against project context and intent. Verified .
- Bypassing Authorization Schema — OWASP WSTG
OWASP identifies horizontal and vertical privilege escalation and describes tests across resource and role boundaries. Verified .
- Introduction to Hypothesis — Hypothesis documentation
Hypothesis generates inputs from a defined strategy and uses stated properties to find counterexamples. Verified .
- Basic Concepts — PIT Mutation Testing
PIT runs tests against mutated code and reports whether each mutation was caught, survived, or was not covered. Verified .
- About protected branches — GitHub Docs
Configured required status checks must pass before a protected branch can accept a merge. Verified .