What I Need Before I Approve an Agent’s Work
$ grep -n "^##" 2026-09-before-approving-agent-work.md
A useful review bundle connects each requirement to evidence for the exact work being approved.
“Implemented, tested, ready to ship” leaves most of the review work untouched. Which requirement? Tested how? Against the candidate in this pull request, or the version already deployed?
I would send a bundle back if those questions required reconstructing the agent’s conversation. Its summary should point to the evidence needed for a decision.
A small worked example
Consider a review bundle for a blog’s publication rules. The example below shows what the bundle should contain; it is a template, not a completed test report.
The requirement has several parts: drafts stay out of public listings and feeds; scheduled posts stay out until their publication time; the eligibility check changes when that time is reached; metadata and preview images must not reveal a draft.
Keep the candidate, checked properties and remaining gaps separate:
| Item | Evidence to include |
|---|---|
| Candidate | Exact commit and test-suite revision |
| Environment | Local checkout, recorded Bun version, env-file loading disabled |
| Scheduling assertions | Past, future, exact-boundary and malformed directives; drafts remain excluded |
| Publication assertions | Draft metadata and both OpenGraph paths stay hidden; full text export includes public articles once |
| Result | Complete command output, including pass and failure counts |
| Outside these checks | Browser rendering, production build, deployed revision or cache behavior in production |
For a Bun project with scheduling and publication suites, the targeted command could be:
bun --no-env-file test tests/lib/blog-scheduling.test.ts tests/lib/blog-publishing.test.ts
In an actual review, the bundle should link the reviewer to the test source and complete log for that candidate. The pass count is useful only alongside the assertions.
One distinction matters immediately: an exact-time test can exercise a function supplied with a fixed clock. It does not show that every cached page on the live site refreshes at that instant. If the release requirement includes visible behavior through the browser, that proof is still needed.
Keep “not run” separate from “failed”
A targeted suite can pass while leaving browser rendering and deployment unverified. Record those checks as not run; do not count them as failures or silently treat them as passes.
Keep separate checks for a branch’s production build and the deployed site. Either can pass while saying nothing about the other. A healthy production page may be serving the previous revision. A passing local build cannot show what production serves.
This is the practical extension of the independent done-check: name both the property checked and the object checked. “Browser tests passed” needs a target. “Approved” needs a candidate.
GitHub provides ruleset options for dismissing stale approvals when the diff changes and selecting the expected app that supplies a required status check. Those controls can help bind evidence to a review. They cannot tell you whether the check covers the requirement, and their availability does not mean they are enabled in every repository.
Ask for a decision the evidence supports
For a proposed release, I want the bundle to say what changed, which requirements each check covers, which failures remain, and which checks were skipped or never run. Then it should name the decision requested.
A passing run of the local checks above would support a narrow conclusion: those publication assertions passed against that candidate. I would not use it alone to approve a claim that the deployed scheduling behavior had been verified.
A bundle is an index into evidence. The reviewer still has to inspect the relevant code, judge the coverage and decide which missing proof matters. Put that remaining decision where they can see it, before the reassuring sentence about everything being done.
$ subscribe --newsletter
Practical AI engineering, in your inbox
Field notes for technical leaders building agents, evaluation systems, governance, and production infrastructure.
Related
The 5-Step Loop: Why Your Agent Fails at Step 4
ReAct gave us a three-step loop. Production hardened it into five. The two new steps — Plan and Verify — are where everything that goes wrong, goes wrong. And the field has now named the worst offender.
When Should an Agent Ask Permission?
A useful agent acts within delegated authority, prepares consequential actions for review, and asks again when the action changes.
Two Papers That Puncture the Hype
One paper shows frontier models degrade as context grows — even on trivial tasks. The other shows reasoning models hit a wall and think less as problems get harder. Read carefully, both point at the same engineering response.