When Should an Agent Ask Permission?
$ grep -n "^##" 2026-07-when-should-agent-ask-permission.md
An agent should ask when the next action exceeds its authority, not merely because it has reached the next step.
Suppose you ask an agent to fix a bug and prepare a pull request. It reads the relevant code, makes a change and runs the tests. Asking permission at each of those steps would make you operate the workflow by remote control. Then it decides to email the affected customer. That is a different decision, even if the email is helpful.
The awkward part is deciding where the original instruction stops. “Be careful” does not answer that. Neither does “ask before anything risky”: the agent still has to decide what risky means, and you still have to interpret a vague approval request.
Authority comes before reversibility
I would evaluate the next action in this order: is it within the existing delegation; what can it affect; how can the effects be reversed? The first question determines whether the agent may proceed. The others determine whether its authority is specific enough for the consequences.
Take the same hypothetical maintenance task:
| Proposed action | Decision |
|---|---|
| Read the relevant source and tests | Proceed within the assigned task and available access. |
| Edit the fix on the agreed branch; run local checks | Proceed if those operations are covered by the task. |
| Send a customer an explanation | Prepare the message; establish authority to send it to that recipient. |
| Deploy the fix to production | Check whether the delegation covers that environment and candidate; otherwise request the specific release decision. |
These are decisions under the stated assignment, not universal tool classifications. Reading can expose sensitive information. Running a test can contact a real service. A branch edit is normally recoverable, but pushing a secret to a public repository has an effect that resetting the branch cannot undo.
Reversibility helps size the decision. It cannot supply missing permission. You can delete an email from your sent folder; that does nothing to the recipient’s copy.
My earlier lifecycle principles put human gates at intake, irreversible actions and merge. Those are useful locations. They still need an action-level definition: which resource, which effect, under whose authority?
Prepare the decision before asking
When a new decision is needed, the agent should first finish the work it already has authority to do.
For the customer email, that means drafting the message, identifying the recipient and showing any attachments. “May I contact the customer?” asks the human to approve an idea. “Send this message to this address, with this attachment?” presents something they can inspect.
Here is an illustrative approval description for that email:
| Field | Approved action |
|---|---|
| Operation | Send one email |
| Recipient | customer@example.com |
| Content | The displayed draft, revision 3 |
| Attachments | None |
| Validity | One send; renewed approval if recipient or content changes |
The interface should show the actual content. A revision identifier or digest can bind the machine’s execution to it, but a hash is not something a human can meaningfully review.
If the agent then adds an attachment, approval no longer covers the proposed send. If it changes the recipient, same result. The implementation needs to compare the operation it is about to execute with the approved operation, not accept a remembered “yes” from somewhere in the conversation.
There is a familiar version of this in code review. GitHub offers a ruleset option to dismiss stale approvals when the diff changes. That option is not a universal default. It illustrates the right relationship: approval attaches to the work that was inspected.
Give the tool a smaller set of powers
A conversational instruction cannot be the only thing preventing an out-of-scope action.
OWASP’s Excessive Agency guidance recommends downstream authorization, limited permissions and human approval for high-impact actions. Its email example is particularly relevant: an agent intended to summarize messages can become a route for leaking information if it unnecessarily has permission to send them.
For an email integration, I would use separate credentials for the inbound relay and the agent’s tools. The relay credential should permit submission of incoming mail without conferring sending authority. That separation limits capabilities; it does not establish a runtime approval check for every outgoing message. An instruction to require an independent user request remains agent policy unless the server verifies approval for the specific action.
Those are different controls, and the distinction matters. A narrow credential restricts available operations. An approval policy decides when a permitted operation should be used. A runtime check can enforce that policy for the specific action.
For OAuth integrations, RFC 9700 recommends restricting token privileges to the minimum needed for the use case. Apply that principle before trying to make the prompt more emphatic.
Repeated work needs standing authority
The opposite failure is an agent that receives a clear delegation and keeps returning it to the user for renewal.
If the assignment permits a branch, a bounded change and local validation, a failed test is normally a reason to investigate within that scope. It is not automatically a reason to ask permission to continue. If the investigation reveals that the fix requires deleting customer data, the proposed effect has changed; that is a new decision.
For a recurring workflow, write down the permitted action class, destinations, limits and escalation conditions. An agent that prepares a weekly internal report might be allowed to read specified sources and update a draft, while distribution remains a separate decision. Another workflow might explicitly authorize sending that report to a fixed internal list. Both can be reasonable. The authority needs to say which one you chose.
Keep missing information separate from missing permission, too. An agent may already be authorized to send a message but need the correct address. Guessing the address would not complete the delegation.
Before an agent acts, it should be able to specify this operation, on this target, with these effects, under this authority. When the action exceeds that authority, use those details to make the approval request concrete.
$ subscribe --newsletter
Practical AI engineering, in your inbox
Field notes for technical leaders building agents, evaluation systems, governance, and production infrastructure.
Related
The 30 Principles for Agentic Engineering — Part 2: The Lifecycle
Principles 6–14. How work moves through an agentic engineering team: the ticket as contract, AI distillation with human curation, three gates, verification before done, characterisation tests, the 1.2× capacity rule, the J-curve, and telemetry.
The 5-Step Loop: Why Your Agent Fails at Step 4
ReAct gave us a three-step loop. Production hardened it into five. The two new steps — Plan and Verify — are where everything that goes wrong, goes wrong. And the field has now named the worst offender.
What I Need Before I Approve an Agent’s Work
A review bundle should connect the requirement to evidence for the exact candidate and make the remaining decision explicit.