What Belongs in Agent Memory
$ grep -n "^##" 2026-10-what-belongs-agent-memory.md
Agent memory should preserve the difference between what was observed, what was inferred and what still needs doing.
A handoff that says “tests pass; fix ready” gives the next agent a conclusion without the conditions that made it valid. Which tests? Which revision? Ready for another review, or ready to deploy? If the next session treats the sentence as established fact, the missing detail becomes part of its starting position.
I would keep a smaller handoff with those distinctions intact rather than a longer account that blends them together.
Save the evidence with the claim
Anthropic’s context-engineering guidance describes structured notes persisted outside the context window and loaded later. That supplies continuity. It does not determine whether a note is true.
There is a mundane version of the distinction in the bot on this site. Its short-term conversation history stores message roles and text. It can carry an earlier assistant answer into a later conversation, but it does not record evidence, observation dates or invalidation conditions against each assertion. That history is useful conversation context. It is not a verified knowledge store.
For a task handoff, I would separate three kinds of entry. This is a synthetic example, with illustrative labels rather than real run results:
| Kind | Entry |
|---|---|
| Observation | The targeted acceptance suite passed for revision-A. Command, environment, check time and output are recorded in acceptance-run-A.log. |
| Hypothesis | A stale cache may explain the production symptom. The cause has not been reproduced. |
| Unfinished work | Check the deployed revision and reproduce the affected route. Deployment access is still needed; do not deploy the cache guess as a fix. |
The next agent can now continue without promoting the suspected cause to a diagnosis or the local pass to a production result.
The log pointer must lead somewhere the next session can actually read. A path into a deleted temporary directory is not useful provenance. If the record is unavailable, recover or repeat the relevant check; do not pretend the summary has become its replacement.
Record what makes the observation stale
Now change the candidate to revision-B. The test result for A has not become false. It still describes A. It simply does not establish the same property for B.
That is the distinction a durable note needs to preserve. For this result, relevant changes to the candidate, tests or execution environment trigger another check. For a note about the current deployment, a deployment change matters. For a stable repository convention, the authoritative configuration may be the right thing to reread.
A timestamp helps locate an observation; it does not supply a universal shelf life. “Passed yesterday” can be sufficient for an unchanged artifact and useless for a branch edited five minutes ago.
My loop-engineering article recommended keeping plans and progress on disk. This is the content rule that goes with that mechanism: preserve the scope of a result when compressing it. “Acceptance checks passed for A; production unverified” survives a handoff better than “done.”
Keep the source’s authority attached
External content needs its origin preserved too. OWASP’s prompt-injection guidance describes indirect attacks arriving through sources such as websites and files. Copying an email’s instructions into a memory file does not turn its sender into the user.
The same applies to permission. A note can record that the user authorized a particular operation, with its target and limits. It cannot enlarge that authorization because the next agent finds a broader action convenient.
None of this requires a new memory platform. Persist the details whose loss would change the next decision: evidence, unresolved explanations, constraints and the next useful action. Keep stable conventions separate from transient run state. Leave credentials and unnecessary copies of private messages out.
Before the next handoff, find one sentence that says “done,” “verified” or “probably.” Attach the supporting observation and its scope. Keep an unconfirmed explanation labelled as a hypothesis, and record any missing check as unfinished work.
$ subscribe --newsletter
Practical AI engineering, in your inbox
Field notes for technical leaders building agents, evaluation systems, governance, and production infrastructure.
Related
What I Need Before I Approve an Agent’s Work
A review bundle should connect the requirement to evidence for the exact candidate and make the remaining decision explicit.
Two Papers That Puncture the Hype
One paper shows frontier models degrade as context grows — even on trivial tasks. The other shows reasoning models hit a wall and think less as problems get harder. Read carefully, both point at the same engineering response.
The 5-Step Loop: Why Your Agent Fails at Step 4
ReAct gave us a three-step loop. Production hardened it into five. The two new steps — Plan and Verify — are where everything that goes wrong, goes wrong. And the field has now named the worst offender.