Testing practice
A practical test pyramid for code written by AI agents
Use the test pyramid as a placement decision, not a required ratio. For each plausible failure in an agent-written change, find the smallest stable boundary that can observe it: a domain rule, a data or authorization integration, a service contract, or one critical browser journey. Keep a higher-level test only when it proves a connection lower-level tests cannot. Reject tests that merely repeat the implementation or duplicate the same assertion at every layer. A durable suite should tell the next agent what broke, not only that a broad flow turned red.
Treat the pyramid as a map of proof boundaries
The test pyramid is a picture of different test scopes, not a universal procurement list for tools. Ham Vocke's Practical Test Pyramid argues for many focused tests and fewer broad tests because broad flows are slower to diagnose and more expensive to maintain. Google's Testing Overview adds a useful distinction: test size describes the resources needed to run it, while scope describes how much behavior it validates. A narrow UI component test might need a browser but still verify a small rule. A service endpoint test might touch many lines of code yet run entirely on one machine. Choose by the failure the test must expose and the stability of its boundary, then name the layer for communication.
AI agents make this decision more important because they can generate tests quickly. More generated tests can create a false sense of rigor when they all assert the same helper's return value or copy the implementation's branch conditions. A useful test has a reason to exist: it would fail for a plausible wrong change and its failure would point toward a manageable area. The test's oracle should come from the requirement or an independent contract, not from the exact code the agent just wrote. The companion guide on testing AI-generated code develops those independent counterexamples; this guide asks where each counterexample belongs in a maintainable suite.
Avoid declaring that a team needs a fixed percentage of unit, integration, and browser tests. Even Google's published percentages are presented as its own rough mix, and architecture changes which seams matter. A tiny static site and a multi-tenant service need different integration proof. The durable decision is whether a lower layer catches a failure quickly without mocking away the very boundary at risk. Keep one broader check when it adds evidence about wiring, user navigation, or deployed configuration. Remove or narrow it when it merely repeats cases already proved below.
Sources: The Practical Test Pyramid — Martin Fowler site · Testing Overview — Software Engineering at Google · Best Practices — Playwright
Map one agent-written change to four distinct failures
Consider an illustrative project-notes app. A coding agent adds an Include archived control to search. The requirement is that a member sees only notes from the current workspace; archived notes appear only when the control is enabled; a member without access never sees another workspace's note. The agent changes a search predicate, a database query, an API response, and the page control. A single end-to-end test could click the control and find one note, but it would be hard to tell whether a later failure came from the predicate, tenant filter, response shape, or UI wiring. The map below gives each risk a more direct proof.
| Plausible wrong change | Smallest meaningful proof | Why the next layer still matters |
|---|---|---|
| Archived predicate inverted | Domain test: active and archived records with Include archived on/off | It cannot prove the data query calls the predicate |
| Workspace filter omitted in query | Data/authorization integration: two workspaces, one requester, real query boundary | It cannot prove the browser sends the intended workspace context |
| API omits `archived` flag or cursor | Consumer contract test on the response fields the page uses | It cannot prove the page renders or navigates correctly |
| Toggle is hidden, mislabeled, or disconnected | One critical browser journey using visible controls and isolated accounts | It proves the user path, not every domain edge case |
The first test belongs near the domain rule because the input and expected output can be stated without a server. The second needs the real query boundary because an in-memory fake that ignores workspace scoping would conceal the fault. The third checks a producer-consumer agreement, such as fields and cursor semantics; it should not rerun every search permutation. The fourth asks whether a user can operate the control and see a permitted result. Each test proves something the preceding test cannot. Together they provide a causal trail when an agent repairs the feature later.
Use two or three concrete records in the example: an active note and an archived note in workspace A, plus a distinctive archived note in workspace B. When Include archived is off, the A member sees only A's active note. When it is on, the member sees A's active and archived notes, never B's note. The B note is a counterexample at the authorization boundary, not merely another search fixture. These labels are illustrative, not a customer data set or an executable test. The implementation should create isolated records and clean them up through its normal test harness.
Sources: The Practical Test Pyramid — Martin Fowler site · Testing Overview — Software Engineering at Google · Testing Strategies in a Microservice Architecture — Martin Fowler site
Put stable domain rules close to their inputs
At the lowest useful boundary, a test states a rule and its counterexample. For the archived filter, define the expected set of notes for Include archived false and true using a tiny fixed collection. The test should be independent of whether the implementation uses a loop, a query builder, or a helper class. If the agent copies the predicate's code into the test and compares two identical expressions, both can share the same mistake. A better expected result is a small literal set derived from the requirement. The test should fail if archived items leak into the default view or disappear when explicitly requested.
Keep domain tests fast and deterministic by controlling inputs rather than relying on current time, shared accounts, or a remote service. If a rule depends on time, pass a clock or fixed instant at the boundary the application already owns. If it depends on feature policy, use an explicit policy value. This makes a failure reproducible for the next agent. Google's Testing Overview treats size and scope separately and emphasizes small tests for speed and determinism; the lesson here is not that all domain logic must be extracted into artificial helpers. The smallest stable entry point might be an existing service method that keeps the rule coherent.
Do not create a separate unit test for every line in the implementation. A test that asserts the name of a private method or the order in which two mocks were called can break during a harmless refactor while missing a wrong result. Ask whether the observation matters to a caller. If the search predicate changes but the returned note set remains correct, a behavior test should stay green. If a persistence call's ordering is itself a safety contract, then it deserves explicit proof at the boundary where ordering matters. The distinction is between protecting behavior and freezing incidental structure.
Sources: Testing Overview — Software Engineering at Google · The Practical Test Pyramid — Martin Fowler site
Use integration tests where data and authority meet
The multi-tenant query is not fully proved by a pure predicate. The risk is that a repository method omits a workspace condition, joins the wrong table, or applies the archived filter after an unsafe broad read. Exercise that repository or service boundary with two workspaces and an authorized requester. Assert on returned identities, including absence of the distinctive B note. Use the application's normal test database or a sufficiently faithful boundary so query semantics are real. A mock that returns only safe A records regardless of the query would make the test pass while the production query remains unsafe.
Authorization deserves a negative case, not just a happy-path result count. A count of two could be satisfied by the wrong two records. Verify exact membership and the access decision for a requester from the other workspace. If the API accepts a workspace ID, test that a requester cannot choose B's ID to override authorization. The test can stay narrow: one request boundary and one data store, with email, search ranking, and browser layout excluded. It is more diagnostic than a long browser journey that happens to notice the same leak only after several unrelated steps succeed.
Integration tests should also cover serialization where the data crosses a process boundary. A database column renamed in code may still deserialize old rows. A response field may be omitted by a serializer even when the domain result is correct. Ham Vocke's Practical Test Pyramid calls out databases, files, queues, and external APIs as important integration points. Test the relevant one in this feature, not all integrations in the system on every change. If a dependency cannot be run locally, choose a controlled test double for the consumer side and a separate contract or provider check that proves the double's assumptions; do not call a production service from every automated test.
Sources: The Practical Test Pyramid — Martin Fowler site · Testing Strategies in a Microservice Architecture — Martin Fowler site · Testing Overview — Software Engineering at Google
Protect the contract the consumer actually needs
A service contract test is useful when two components can change independently. In the notes app, the page may require each result's ID, title, archived state, and a cursor for the next page. A narrow consumer contract test should assert those fields and the relevant semantics, such as whether an archived result carries the flag the page uses. It need not snapshot every JSON property or render a browser. Fowler's microservice testing discussion frames contract tests around what the consumer depends on, which keeps them smaller than a full system test while catching incompatible interface changes.
Do not assume a contract test proves authorization. If the provider returns the right shape but includes B's note, the schema can be perfectly valid while the result is a security failure. Conversely, a data integration test may prove scoping while missing a serializer that drops the cursor. Keep those proofs separate when the risks are separate. The agent should name the producer and consumer versions or revisions under test and explain whether the provider verified the same contract that the consumer expects. A fabricated mock response with no provider verification is an assumption, not end-to-end compatibility evidence.
When an API field changes, choose a compatibility test at the boundary where old and new clients overlap. A new page could read `archived` while an old page ignores it, which may be safe. Removing a field that the old page still reads is different. A contract test can encode the required overlap without asserting the provider's private data structures. This makes agent-written changes easier to review: the question becomes whether the externally visible agreement still holds. It also gives the next agent a targeted failure when a provider update breaks the page, rather than a generic browser timeout several steps later.
Sources: Testing Strategies in a Microservice Architecture — Martin Fowler site · The Practical Test Pyramid — Martin Fowler site
Keep one browser journey for the user-facing connection
The browser test should prove what the lower layers cannot: a member can find the Include archived control, operate it, and see an allowed archived note without seeing another workspace's note. Start from an isolated signed-in test account or approved fixture, navigate as a user would, and assert visible behavior. Playwright's guidance recommends user-visible locators and isolated tests because implementation details and shared state make UI checks brittle. The test need not repeat every domain combination. One critical journey confirms that the API, page wiring, and rendered state meet at a usable path.
A weak browser test might click a CSS class named `.archive-toggle` and assert that a hidden internal state variable changed. That can pass while the visible result remains wrong, or fail after a harmless redesign. A stronger check finds the control by its visible label or role, activates it, and observes the expected note title in the results. If the page updates asynchronously, use the framework's web-first assertion that waits for the observed state instead of a fixed sleep. Playwright documents retrying assertions for dynamic pages; this does not make an eventually passing wrong result acceptable, so the assertion still needs a clear expected outcome and bounded timeout.
Browser tests have a real cost: fixtures, account state, navigation, rendering, and timing all become possible failure sources. Keep the journey long enough to test the user connection and short enough to diagnose. If the only purpose of five browser cases is to enumerate archived true/false combinations already covered at the domain boundary, remove those duplicates. If one browser case catches a missing label or a disconnected control that no lower test sees, keep it. The relevant measure is added confidence per maintenance cost, not an arbitrary count of browser tests. When it fails, collect the trace or step evidence needed to distinguish app behavior from environment setup.
Sources: Best Practices — Playwright · Assertions — Playwright · The Practical Test Pyramid — Martin Fowler site
Reject mirror tests and separate environment failures
Generated tests often mirror implementation details because those details are easiest for an agent to read. Suppose the code has `if (includeArchived) return all; return active`. A test that repeats that exact branch with the same input collection may pass even if the original requirement was to include only archived notes the member may access. Require the test author to write the expected record identities independently, including the B note that must stay absent. Reviewers can then compare the oracle to the requirement rather than compare two pieces of matching code. The original testing-AI-code guide focuses on constructing that independent oracle; this guide places it at the correct boundary.
Do not elevate every failure to a browser test. If the query leaks B's record and a narrow data test catches it, adding a second browser test for every workspace combination may consume time without finding a new class of error. Keep a browser authorization check only if it proves a distinct user-facing path or a context-propagation seam that the data test cannot see. Ham Vocke's practical pyramid explicitly warns against repeating the same assertions at higher layers. A higher-level test earns its place by testing a connection, configuration, or user outcome that a lower one cannot observe.
Some failures belong to the environment rather than the code layer. A test database may fail to start, an identity fixture may expire, or a browser worker may lose its session. Mark that as setup or infrastructure evidence rather than calling the search rule wrong. Rerun only when the cause is plausibly transient and the rerun will test the same revision under a known condition. A consistently failing database setup deserves repair, not a retry quota. This classification preserves trust: the team can tell whether a red result implicates domain behavior, integration, contract, browser wiring, or the test environment itself.
Sources: The Practical Test Pyramid — Martin Fowler site · Testing Overview — Software Engineering at Google · Best Practices — Playwright
Let the suite evolve when it finds a real gap
A pyramid is not finished when the initial change lands. When a broad test catches a bug that lower layers missed, ask whether the failure can be represented at a narrower stable boundary. If the browser finds that the API drops `archived`, a contract test may capture the regression more directly while one browser journey continues to check wiring. If a production incident reveals that the test database used a different query mode, repair that fidelity gap. Do not simply add another broad test around the incident and leave the original blind spot intact. Each regression should improve the shortest useful feedback loop.
Review tests that fail during unrelated refactors. Some are signaling real contract changes; others freeze internal method names, CSS selectors, or incidental call order. Keep the former and rewrite or remove the latter. Track repeated failures by cause and the time needed to diagnose them, not only by how often a suite is green. A test that is frequently red for environment setup can drown out a rare real authorization failure. Google's testing guidance emphasizes reacting quickly to broken tests to preserve confidence. For a small team with agents, a clear owner for repairing test infrastructure is as important as the agent that writes new tests.
Test selection can follow changed boundaries without pretending that a static file list proves safety. A change to the archived predicate should run domain tests; a repository query change should include data and authorization integration; an API serializer change should include the consumer contract; a UI control change should include the browser journey. Shared libraries and migrations may broaden the set. Record which tests actually executed for the exact revision and why skipped tests were out of scope. If the selection mechanism is uncertain, run the broader set rather than assert that unrun checks passed. The related flaky-test and parallelization resources help explore feedback tradeoffs after the suite's proof value is understood.
Sources: Testing Overview — Software Engineering at Google · The Practical Test Pyramid — Martin Fowler site · Best Practices — Playwright
Give the next agent a test map it can trust
An agent handoff should fit the failure map on one page: changed requirement, plausible wrong outcomes, selected boundary for each, exact test command or suite identifier, evaluated revision, and any missing proof. For the notes example, the domain test protects inclusion rules, the integration test protects workspace isolation, the contract test protects fields consumed by the page, and the browser journey protects navigation and visibility. The next agent can start from the failing boundary rather than reread an entire generated suite. A passing result is limited to the behavior and revision it actually checked.
This testing map is narrower than the full CI/CD release journey. The best-practices guide connects source, build, policy, provider, and live verification. A test pyramid helps choose durable proof within that journey; it does not show by itself that an artifact was deployed or that a live target serves the tested revision. For Agentic Pipeline, Auto Profile is the default when a repository has no recipe and an existing repository-owned recipe takes precedence. The platform can execute supported checks and ship eligible merges, but the repository still needs tests that encode its own product rules. The public feature guide describes the build-and-ship path; inspect the actual run and delivery evidence for a specific change.
A useful review question closes the loop: if the agent made one plausible mistake at each boundary, which test would fail first, and would its failure explain what to fix? If a failure has no test, add the smallest durable proof that can see it. If five tests all fail for the same copied predicate, remove duplication after the distinct connection is covered. If a browser journey is the only way to see a mislabeled control, keep it and make its fixture stable. The objective is not a triangular chart with attractive numbers. It is a suite that gives fast, credible evidence about behavior while leaving clear limits for deployment and live verification.
Sources: The Practical Test Pyramid — Martin Fowler site · Testing Overview — Software Engineering at Google · Best Practices — Playwright · Agentic Pipeline: Public feature guide
Sources and verification
- The Practical Test Pyramid — Martin Fowler site
Ham Vocke's practical pyramid distinguishes granular tests, integration boundaries, limited end-to-end journeys, and avoiding duplicated assertions across layers. Verified .
- Testing Overview — Software Engineering at Google
Google distinguishes test size from scope and discusses determinism, smaller tests, and the cost of large-test failures. Verified .
- Testing Strategies in a Microservice Architecture — Martin Fowler site
Consumer contract tests verify the provider behaviors and fields a consuming component actually depends on. Verified .
- Best Practices — Playwright
Playwright recommends tests of user-visible behavior, resilient locators, and independent test isolation. Verified .
- Assertions — Playwright
Playwright's web-first assertions retry for asynchronous visible state instead of relying on fixed waits. Verified .
- Agentic Pipeline: Public feature guide
The public feature guide describes Auto Profile default, existing-recipe precedence, and the build-and-ship journey; access remains closed beta. Verified .