Skip to content

Production Agents Are Outgrowing the while(true) Loop

#ai-agents #agent-architecture #event-sourcing #llm-agents #agentic-systems #ci-cd

Open any agent codebase. Claude Code, Codex, most of the open-source ones. You'll find the same skeleton:

javascript
let state = {};
while (true) {
  const plan = await llm.plan(state);
  const results = await runTools(plan);
  state = updateState(state, results);
  if (isDone(state)) break;
}

It's the natural first design. The model is the brain, the loop is the heart, and state is whatever bag of objects you've accumulated. It works great for demos.

Then you live with it.

I spent a weekend with Qwen 3.8 27B on a pair of 3090 Tis, running it through Cline and ZooCode with a few MCP servers. The model burned through tokens, then either finished the task wrong or looped forever trying to fix something that wasn't broken. I found the same failure mode in every local-model thread: tons of tokens, incorrect finishes, and loops that never terminate.

Blame the loop, not the model. The while(true) shape is fine for a demo, but it has four structural cracks that show up the moment you run an agent for real.

The four cracks in the loop ​

Interruption is a hack. Kill the process mid-turn, or let a tool hang, and you've got a half-finished iteration sitting in state. You either throw it away or patch it in place. Either way, state is lying to you about what actually happened.

Retries are special cases. A tool fails? Write a try/catch, maybe retry, maybe pass the error to the model, remember to add it to state. After a few months you've got a dozen ad-hoc branches for "what if this turn didn't finish cleanly."

Parallel tool calls are awkward. The model wants to call read and grep at the same time. Now the loop has to sequence them or spawn promises and reassemble the result before the next llm.plan(). More state to manage.

You can't rewind. state is a pile of mutable objects. You can't branch from an earlier point in the conversation without reconstructing it by hand. You can't replay what the model actually saw. Good luck debugging a long session.

These aren't implementation details. They're the direct result of the shape: one mutable state object trying to stand in for the entire history of the conversation.

Dimensionwhile(true) loopEvent-sourced agent
InterruptionHalf-finished turn corrupts stateInterruption is just another event
RetriesAd-hoc try/catch branchesReplay the failed event
Parallel toolsSequence or manually reassembleEach tool call is an independent event
Rewind / branchReconstruct state by handFork a new thread_id from any sequence
DebuggingReplay what the model sawRe-apply the event log
Multi-agent handoffShare a mutable objectSend a tell event between logs
Operational costOne in-memory objectA real database in the hot path

The log is the state ​

A few months ago I started building a personal agent called Pizza with a different premise: the log is the source of truth, and the state is just a projection of that log.

Every message, tool call, result, and file edit becomes one row in an EventStore:

sql
CREATE TABLE events (
  sequence INTEGER PRIMARY KEY,
  event_id TEXT,
  type TEXT,
  payload_json TEXT,
  caused_by TEXT,
  thread_id TEXT
);

The runtime doesn't hold state in memory. It reads the tail of the log, dispatches the next event to a handler, and appends the result back. A turn is a state transition, not another loop iteration.

The UI, the LLM context, and the session tree are all queries over the same log. To see what happened, you read the events. To branch the conversation, you start a new thread_id from an earlier sequence. To replay, you re-apply the events.

Once you commit to the log, a bunch of otherwise hard features stop being special cases. There's no "new chat" button, because there's no hard boundary between tasks. Instead of a JSON menu of read_file, write_file, grep, and git tools, the model gets one cli tool. Built-in commands like read and write use structured internal handlers; everything else, grep, sed, git, npm, python, goes straight to the user's shell. The model has to learn shell, but it gets to compose real pipelines. And because every invocation is one event, the log stays consistent.

Quick Take: The while(true) loop is a fine starting point, but production agents need a durable record of what happened, not a bag of mutable state.

The same shape enables multi-agent collaboration. One agent can send a tell event to another agent in a different workspace. The receiving agent handles the task in its own event log and writes a result back. Project files from workspace B never leak into A's log, only the request and response.

Event sourcing isn't free. You now have a real database in the hot path. Replay cost grows with log size, and you'll eventually want periodic snapshots. Anything outside the log is an invisible side effect and a bug waiting to happen. But the hard things, forks, replays, multi-agent collaboration, debugging, become ordinary database operations.

The comments on the original post mostly agreed: event logs are the part most agent loops rediscover too late, and replay without idempotency keys is a second execution bug waiting to happen.

Safety gates: the part nobody demos ​

The event log solves auditability. It doesn't solve the problem of an agent doing something destructive. That's where the second pattern in this wave of agent tooling comes in: read-only by default, with explicit gates for writes.

The claude-ads project is a good example. It's a Claude-first tool for paid-media operations: audits, campaign plans, creative briefs, monitoring, reports. Every command defaults to drafting, not doing. /ads launch --draft produces a campaign mutation plan without touching the account. All adapters are read-only by default. Applying a change requires a tested and enabled capability for the exact operation, explicit account and object IDs, a human-readable before and after diff with blast radius, owner approval within account-defined ceilings, an idempotency key, an audit destination, rollback, and a verification window. Missing ceilings mean no write. Permanent deletion isn't supported at all.

The project's control-plane directory records product boundaries, dated sources, claims, capabilities, and safety rules. No source means no platform claim. No approval and rollback means no account mutation. That's the part that never makes it into the demo video. It's also the part that determines whether an agent survives contact with a production account.

PostHog's self-driving mode runs on the same principle. Signals in product data, errors, rage clicks, failed queries, become researched reports and pull requests that a human reviews and merges. The agent proposes. The human disposes.

Claude Tag: agentic CI/CD in production ​

Anthropic's CI team runs an on-call agent called Claude Tag. It lives in Slack, holds memory across the on-call channel, and acts in real time to events. When an alert escalates to an incident, an orchestration agent spins up executor subagents to investigate each dependency: Grafana, the log store, PagerDuty, GitHub, Kubernetes, all wired via MCP connectors. Executors report back, and the orchestrator synthesizes a SITREP.

The numbers tell the story.

Key Numbers

  • 14 minutes: median time from incident open to first evidence-grounded analysis
  • 4 minutes: fastest root-cause identification in a first report
  • 15 minutes: typical time to first SITREP
  • 8x: code shipped per engineer per quarter, 2021 to 2025
  • 617 lines: the investigation skill for shadow divergence bugs
json
{
  "type": "bar",
  "title": "Claude Tag incident response timing",
  "x_label": "Metric",
  "y_label": "Minutes",
  "caption": "Claude Tag publishes its first evidence-grounded analysis a median of 14 minutes after an incident opens, names root cause in as little as 4 minutes, and typically posts a SITREP within 15 minutes.",
  "data": {
    "labels": ["Median first analysis", "Fastest root cause", "First SITREP"],
    "series": [
      {
        "name": "Minutes",
        "values": [14, 4, 15]
      }
    ]
  }
}

14 minutes is fast enough that the on-call engineer can stay in bed and review a grounded hypothesis in the morning. 4 minutes is faster than most humans can even open the right dashboards. The 8x shipping figure is the context: when engineers ship 8x more code, the CI failure rate scales with it, and humans can't keep up. Agentic CI is the only way to match agentic coding.

Two details make this work. First, a lessons.md file that the agent appends to after every incident. Every new investigation starts by reading it, so the first hypothesis is grounded in what happened recently. When a pattern shows up enough times, it gets promoted into the investigation skill itself. Second, the investigation skills are markdown files committed to a GitHub repo, versioned like code. The engineer built a 617-line skill for shadow divergence bugs by troubleshooting turn-by-turn with Claude during a real incident, then having it write the file from that experience. 617 lines encodes every step of a typical investigation: what to check, in what order, and what each result means.

That's the abstraction-discovery pattern in production, which I'll get to in a minute.

The research frontier: agents that build their own tools ​

There's a paper out of arXiv that formalizes what lessons.md is doing. Procedural Content Metageneration via Program Search and Continual Abstraction Discovery evolves complete Python generators through LLM mutation and crossover, searching over procedural content generators for Sokoban, Zelda, Dangerous Dave, and Lode Runner. The twist is Continual Abstraction Discovery (CAD): extracting reusable primitives from high-fitness programs into a run-specific helper module.

The experiment is a clean 2x2: CAD on or off, crossed with access to a fixed hand-written domain API. The completed dataset has 160 complete runs, with at least ten 50-generation runs in every cell. CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs, and the agent repeatedly rediscovers validation, reachability, and structural utilities.

The practical read: agents that accumulate their own abstractions beat agents that start from scratch every run. The helper module is the research version of lessons.md, of the investigation skill, of the plugin skill directory. The plugin marketplace Anthropic ships for Claude Code has the same shape: skills, agents, and commands bundled as versioned, installable units that any agent can load. The agent isn't just executing a loop, it's building a library of what it learned. That's the direction the whole field is moving.

Common Pitfalls ​

Treating the loop as the architecture. The while(true) shape is a scheduling mechanism, not a data model. If your agent's entire history lives in one mutable object, you've built a system that can't be replayed, branched, or audited. The event log is the part most agent loops rediscover too late.

Replaying without idempotency keys. This is the one that bites in production. If you replay an event log after a crash and a tool call was already executed on the remote side, you get a double execution. My team ran into exactly this. We now attach a deterministic intent fingerprint to every mutation event before dispatch. If a recovered worker sees an unconfirmed external action, it doesn't re-invoke the tool, it enters a reconciliation phase and checks whether the fingerprint already exists on the remote system.

Letting the agent write by default. Read-only is a product decision, not a limitation. The claude-ads pattern of requiring approval, idempotency, verification, and rollback before any write is the difference between an agent you demo and an agent you trust with a production account.

Running a small local model with high reasoning effort. Qwen 3.8 27B is a capable model, but agentic loops amplify its weaknesses. On 2x3090 Tis with a 50k context window, it runs tons of tokens, then either finishes wrong or loops trying to fix something that isn't broken. Set the reasoning effort to medium, keep the context under control, and match the model to the task. A 27B model is not going to match Claude Code for complex multi-file edits, and that's okay.

Skipping the snapshot layer. Event logs grow. After a few thousand events, rebuilding context from the log on every fork gets slow. You need periodic materialized snapshots, and you need to design for them from the start, not bolt them on when sessions start timing out.

One Thing to Remember ​

Every production agent covered here shares one design bet: the agent's memory is external, versioned, and inspectable. Whether it's an event log, a lessons.md file, or a run-specific helper module, the agent that remembers what it did and why beats the agent that starts fresh every turn. Build the memory first, and the loop becomes an implementation detail.

What This Means for Production Agents ​

If you're building an agent that runs longer than a single session, adopt an event-sourced core. The log-as-state pattern turns forks, replays, and multi-agent handoffs into database operations, and it's the only pattern here that survives contact with a crash.

If you're constrained by a local model, skip the complex agentic workflows and match the model to the task. A 27B model with medium reasoning effort handles bounded jobs like config editing, but it will loop and burn tokens on open-ended multi-file work. Use the small model for the narrow slice, and route the hard stuff to a frontier model.

One thing to watch: the abstraction-discovery direction is moving fast. Claude Tag's lessons.md, the plugin skill directories, and the CAD paper are all converging on the same idea: agents that build and reuse their own tools. Expect skill libraries that agents write for themselves to become a standard part of agent runtimes within the next year.