Loop Engineering: The Loop Was Never the Hard Part
$ grep -n "^##" 2026-06-loop-engineering-claude-code.md
I didn't write the first draft of this post. A loop did.
I typed one command, handed it a folder of research, and walked away. While I made dinner, a fleet of agents argued about it — researchers pulled sources, a strategist picked the angle, a writer drafted, a fact-checker tore through the numbers, an editor scored the prose against a rubric. State got written to disk between every stage so the next agent could pick up where the last one died. I came back to a draft, not a blank page.
When half my feed started telling me this month to "stop prompting and start looping," my reaction was: yes — and you're selling the wrong half. I've been shipping loops for a year. Early versions of Gluon, my own agent orchestrator, ran without a circuit breaker; a single run once blew through $500 in tokens before I noticed. Another lost two hours in retry hell, Claude convinced the error was its fault when the system just needed a restart.
What loop engineering is — and who said what
The term has a clean provenance, and most of the videos get it wrong. Boris Cherny, who heads Claude Code at Anthropic, is the catalyst. On stage at the Acquired Unplugged event hosted by WorkOS on 2 June 2026 he said: "I don't prompt Claude anymore. I have loops that are running. They're the ones that are prompting Claude and figuring out what to do. My job is to write loops." Five days later Peter Steinberger — the creator of OpenClaw, now at OpenAI — posted the line that went viral: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." And the same day, Addy Osmani gave the idea its name: "Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead." His sharpest line is worth tattooing on the wall: the agent forgets, the repo doesn't.
It's a real progression, not an obituary for what came before:
Prompt engineering -> Context engineering -> Harness engineering -> Loop engineering
(steer one output) (feed the right info) (stable exec env) (the system prompts
the agent for you)
Prompt engineering didn't die; it moved up a level — to the seed, the spec, the test cases, the definition of done. A precise instruction matters more in a loop, because it now compounds across hundreds of unwatched steps instead of one.
Strip away the marketing and a loop is a plain object: a cron job with a decision-maker in the body. Something wakes the agent, it reads state, acts, observes, checks the result against an objective "done," and either stops or goes around again — persisting what it learned to disk. The lineage runs straight back to the ReAct paper (Yao et al., 2022): reason, act, observe, repeat. The 2026 version adds three things the early autonomous-agent experiments lacked: context resets, a separate verifier, and a hard stop.
The loop is the easy 20%
Geoffrey Huntley's Ralph loop fits on one line:
while :; do cat PROMPT.md | claude-code ; done
An infinite loop piping a prompt file into a coding agent, fresh context every iteration. Huntley's July 2025 post describes using it to build a programming language.
But a while loop with no exit condition is a very confident token furnace. The hard 80% is the three things the one-liner doesn't have:
- A done-check the agent can't fake. Not "the model said DONE." Something external and objective.
- A verifier that never grades its own homework. The thing that produced the work cannot be the thing that signs off on it.
- A hard ceiling so a runaway can't eat your weekend.
A loop can increase production without improving correctness. The bottleneck moves from writing prompts to verifying output before it ships. When generation outpaces review, the loop points a firehose at that backlog.
The Claude Code primitives
The headless CLI starts with claude -p "<prompt>" --max-turns 10 --max-budget-usd 5.00. Those flags bound one invocation; an outer loop starts a new allowance each time. Set tool permissions deliberately and track aggregate spend separately.
Hooks can enforce checks at lifecycle events; a Stop hook can block completion. Subagents provide separate conversations and tool permissions. Git worktrees separate working directories, though they are not a security boundary. The Claude Agent SDK exposes the agent loop through query() for a programmatic runner.
/loop schedules recurring prompts within a session. /goal, released on 11 May, keeps Claude working toward a stated condition. Cloud routines run without an open laptop; --bg starts a local background agent. These are different kinds of persistence. None establishes that the output meets your acceptance criteria.
I got the existence of /goal and routines wrong in an earlier draft because I trusted explainer videos over the docs. The dated releases settle that question; they do not certify my verifier.
Worked example 1: agent-until-green, where the shell is the judge
Let the shell check the test exit status. This illustrative loop stops on an agent error or a passing suite; it does not protect the tests from edits or bound the duration of an individual agent call:
#!/usr/bin/env bash
# loop.sh — illustrate an exit-status gate, not a production runner
set -euo pipefail
MAX="${1:-50}"
ITERATION=0
while [[ $ITERATION -lt $MAX ]]; do
ITERATION=$((ITERATION + 1))
echo "=== ITERATION $ITERATION ==="
# Fresh context window every iteration (Ralph pattern).
# PROMPT.md tells the agent: read IMPLEMENTATION_PLAN.md, do the next
# task, run the tests, log to PROGRESS.md, exit.
if ! claude -p < PROMPT.md; then
echo "Agent command failed; stopping for review." >&2
exit 1
fi
# Check the suite after a successful agent call, ignoring any DONE claim.
if bun test > /tmp/verify.txt 2>&1; then
echo "Local suite passed after $ITERATION iteration(s); review the diff."
exit 0
fi
echo "Still red. Continuing."
tail -20 /tmp/verify.txt
done
echo "Hit the iteration ceiling without going green." >&2
exit 1
Each iteration gets fresh context. The prompt directs the agent to re-read IMPLEMENTATION_PLAN.md and PROGRESS.md from disk. That keeps the plan available without carrying the entire previous conversation forward.
The agent's claim of "DONE" is ignored. Success here means the local suite passed. If the agent can edit that suite, it can also change the evidence. Keep acceptance checks and their runner outside its write permissions, inspect changes to tests, and add per-run time and spend limits before using this unattended. A green result only covers the behavior those checks exercise.
Worked example 2: Maker/Checker — what a second call buys you
When tests leave questions about the design, a second review can help. Anthropic's evaluator-optimizer pattern uses one call to generate a response and another to evaluate it, with clear criteria and measurable benefit from iteration.
In this sketch, the evaluator receives the task and proposed solution without the generator's conversation history. That is a design choice, not an Anthropic requirement that system prompts must differ. Separate calls do not guarantee unbiased judgment; the evaluator needs evidence and can still miss the same defect.
# Two calls sharing the task and solution, but no conversation history.
gen = client.messages.create(
model="claude-opus-4-5",
max_tokens=1024,
messages=[{"role": "user", "content": task}],
)
solution = gen.content[0].text
evaluation = client.messages.create(
model="claude-opus-4-5",
max_tokens=1024,
system="Evaluate the proposed solution against the task requirements. "
"Reply PASS if it fully satisfies the task, or FAIL with specific gaps.",
messages=[{"role": "user",
"content": f"Task: {task}\n\nSolution:\n{solution}\n\nEvaluate."}],
)
A different model is an option to test, not a guarantee of independent errors. Compare reviewers against known defects before choosing a cheaper one. Waiting a few hours does not, by itself, give an automated checker new evidence.
I run four or five Claude Code agents across my projects, and one reviews the others' diffs. It catches real things; I'd be slower without it. But I would not use that review alone as approval for a sensitive change. A loop that grades its own output is still AI reviewing AI. Cross-model verification can screen for defects; it does not establish independent accountability.
I learned the done-check the expensive way, on Gluon. My first instinct was to ask Claude to output "DONE" when finished. It doesn't work — Claude gets creative, "done" because one section completed, or "I've finished the implementation, no tests yet" when the requirement was implementation plus tests. So I built a multi-signal completion detector: four independent signals, a 60% confidence threshold, exiting only when enough align. The whole architecture exists because the agent will happily tell you it's done in the middle of not being done.
Cost and failure rate
First, correct the number you've seen, because it anchors every breathless thread and it's misattributed. The viral $1.3 million-per-month token bill — a CodexBar screenshot showing $1,305,088.81 and about 603 billion tokens over 30 days — is Peter Steinberger's, tied to his OpenClaw experiment, a project OpenAI already sponsors. Not Pieter Levels, as the threads keep claiming. The screenshot is a usage display, not an audited invoice, and it says nothing about anyone else's budget. It's not the aspiration the videos imply — it's the cautionary exhibit, what a token furnace looks like when someone else is footing the bill.
For the rest of us, measure the actual workload: cost per accepted result, failed attempts, verifier calls and concurrent agents. A cheap generation that takes three expensive repairs may lose to a dearer first pass. Count the review time too.
Microsoft's AI Red Team updated its taxonomy of agentic failure modes on 4 June 2026, adding goal hijacking alongside existing concerns such as memory poisoning and human-in-the-loop bypass. A completion check needs to consider what the agent was supposed to do, not merely whether it kept running.
Anthropic's December 2025 internal study surveyed 132 of its engineers and researchers. Respondents reported using Claude in roughly 60% of their work, while most said they could fully delegate only 0–20%. Those are self-reports from one company, not a measure of all developers or autonomous-agent success rates. The famous "90 to 95% of code is AI-generated" stat is mangled in transit the same way: Garry Tan's actual claim was that 25% of one YC cohort, Winter 2025, had 95% of their lines LLM-generated. Twenty-five percent of one batch is not "all top startups." Keep the qualifier.
Loops shine on binary, objective tasks; on nuanced, creative work they remain slop machines. My own productivity number sits at something close to 1.3×, not 10×, and it took a few sprints of watching the bug queue catch up with me before I trusted even that. Loops didn't change the multiple. They changed where the time goes — out of typing prompts, and into reading output.
When to use a loop
Keep state on disk so the verifier has a durable record to inspect. Set hard stops because the loop can be confidently wrong. Engineer the seed: a vague spec poisons an objective check.
Rendering diagram...
The easy 20% is the blue boxes. The hard 80% is the two yellow ones.
Before using a loop, check four conditions. Is the task recurring? Can "done" be objectively verified? Can you afford the wasted runs? Does the agent have the tools to do and verify the work? Any "no" and the task isn't a loop candidate.
What my own pipeline still can't do
In the pipeline that drafted this article, a fact-checker and an editor assess the writer's output against a rubric, with their own independent context. State lives in a body-of-work.md on disk. There are hard stops. That separation produces something worth editing, and I trust it enough to walk away while it drafts.
The pipeline checks citations, compares claims with research and flags prose patterns. Those checks can miss errors; a passing report is not proof that the article is sound. It also cannot tell whether the post landed: whether the argument was right, whether it changed one CTO's mind. I could build a delayed scraper that feeds published engagement back into body-of-work.md, the same trick that closes the fuzzy LinkedIn loops. But even wired up, "was the argument correct?" remains a judgment call. There is no objective done-check for true.
That's the same thing the auditors and the regulators keep asking for: a named human who can say "yes, this is right, ship it." The loop can't be that person. It was never going to be.
So the loop didn't remove me from the work. It moved me to the only seat that was ever load-bearing — the one where someone decides the output is true before it goes out. Stop prompting and start looping, by all means. Just don't mistake the loop for the job. I came back from dinner to a draft, and then I did the only part that was ever hard.
$ subscribe --newsletter
Practical AI engineering, in your inbox
Field notes for technical leaders building agents, evaluation systems, governance, and production infrastructure.
Related
Ralph Loop: Teaching AI Agents to Work Autonomously (Without Burning Your Budget)
How Gluon's Ralph Loop enables autonomous Claude execution with built-in safety rails — circuit breakers, multi-signal completion detection, and cost controls that scale from simple tasks to complex workflows.
When a Bank Open-Sources Its AI, It Ships the Control Layer — Not the Model
Santander open-sourced its AI control layer in June 2026 — the harness, the guardrails and the governance — but the early public activity leaves adoption unproven.
What I Need Before I Approve an Agent’s Work
A review bundle should connect the requirement to evidence for the exact candidate and make the remaining decision explicit.