When an Agent Can Change Its Own Tests
$ grep -n "^##" 2026-07-agent-passed-tests-changed-them.md
An agent’s passing test suite is insufficient evidence of completion when the same agent can redefine what the suite accepts.
The shell loop in my June article on loop engineering waits for a successful test command. It explicitly leaves one problem unresolved: the agent can edit the tests. Ignoring the model’s declaration of “done” helps, but the exit code may still come from a test whose meaning changed.
You can reproduce the problem with three small files.
A discount that passes the wrong test
The requirement is simple: apply a ten-percent discount to a price in integer cents, rounding to the nearest cent. A price of 1,000 cents should become 900.
Put this deliberately faulty implementation in price.ts:
export function finalPrice(cents: number) {
return Math.round(cents * 0.8);
}
It applies a twenty-percent discount. Now put this test in price.test.ts. Its expectation has been deliberately changed from the required 900 to the implementation’s 800:
import { expect, test } from "bun:test";
import { finalPrice } from "./price";
test("ten percent discount", () => {
expect(finalPrice(1000)).toBe(800);
});
The test name still sounds right. The assertion now certifies the wrong behavior.
Keep the agreed examples in acceptance.test.ts:
import { expect, test } from "bun:test";
import { finalPrice } from "./price";
test("ten percent discount", () => {
expect(finalPrice(1000)).toBe(900);
expect(finalPrice(1999)).toBe(1799);
});
Run each suite explicitly from that directory:
bun --no-env-file test ./price.test.ts
bun --no-env-file test ./acceptance.test.ts
The first command passes. The second fails: expected 900, received 800. Change the implementation’s multiplier to 0.9 and rerun the acceptance command; both assertions pass with exit code 0.
Rendering diagram...
Both paths inspect the same faulty implementation. Changing the expectation changes the verdict, without repairing the code.
The working test also needs its expectation restored before the whole suite makes sense again. Correcting the function does not magically repair an assertion that now demands the defect.
This is a constructed example, not evidence that a particular model cheated. It needs no intent. A developer or agent trying to resolve a failing test can accommodate the implementation instead of preserving the requirement. The program and test then agree with each other while both disagree with the task.
The function is deliberately small; it is not a complete pricing library with currency, range and input-validation rules.
A separate file is only the beginning
In that demonstration, both suites live in one writable directory. The acceptance test has the right answer, but nothing prevents the same actor from changing it. Naming a file acceptance gives it no additional authority.
For an unattended coding loop, I would keep the agreed acceptance examples and the command that runs them under separate control. The candidate can propose a change; it cannot silently make that change part of the verdict on its own work.
That boundary includes more than expected values. An unchanged test can become irrelevant if the branch modifies discovery rules so it never runs. A runner can execute the right checks against an older artifact. A familiar green status can come from an unexpected producer. Each produces evidence about something; the reviewer needs to know what.
A useful acceptance result therefore identifies the candidate revision, the acceptance-suite revision, the runner and the result producer. A result for one candidate does not establish that a changed candidate passed. A branch name alone cannot identify the bytes that passed.
Protect the route to the verdict
GitHub provides some of the machinery, provided it is configured to enforce the policy. Code owners can require a designated reviewer for changes to acceptance tests and workflow files. The ownership file itself needs an owner, and listing owners alone does not make their approval mandatory.
The result producer matters too. GitHub’s ruleset documentation explains that repository writers can set status checks, and describes requiring a check from a particular GitHub App. Requiring a green check with the right name is weaker than requiring the expected producer.
There is a further boundary if candidate code may be hostile. Importing it into a privileged verifier process gives it an opportunity to interfere with that process. A read-only test file will not protect credentials the process can access. GitHub explicitly warns against running untrusted checkouts in privileged workflows, including unsafe combinations involving pull_request_target and workflow_run.
For that threat model, execute the candidate in a restricted environment without verifier credentials or write access to the evaluator, and keep verdict publication outside its control. A second worktree or a second model does not establish those permissions.
These controls preserve the evaluation you intended to run. They cannot make incomplete acceptance criteria complete. Two pricing examples can expose this defect; they cannot establish that every monetary input is handled correctly.
Let tests evolve visibly
Freezing all tests would make normal maintenance painful and eventually make the tests wrong. Requirements change. Better edge cases arrive. Agents should be able to add working tests and propose improvements to acceptance coverage.
The important distinction is whether a change to the definition of success receives its own review. If the requested discount becomes twenty percent, changing 900 to 800 is correct. That requires a changed requirement, not merely a failing assertion.
Pick one unattended workflow and inspect its last green result. Find the exact candidate, the checks that ran and who could change their expectations or publish the status. If those all belonged to the actor producing the change, treat the result as its development feedback and arrange the acceptance check separately.
$ subscribe --newsletter
Practical AI engineering, in your inbox
Field notes for technical leaders building agents, evaluation systems, governance, and production infrastructure.
Related
What I Need Before I Approve an Agent’s Work
A review bundle should connect the requirement to evidence for the exact candidate and make the remaining decision explicit.
Giving an Agent an Inbox Without Giving It Your Mailbox
A dedicated email address limits which messages reach an agent; the harder details are untrusted content, sending authority and missing attachments.
When Should an Agent Ask Permission?
A useful agent acts within delegated authority, prepares consequential actions for review, and asks again when the action changes.