Skip to content

The Agent Loop Is the New Model: Repo0, Cursor /goal, and the State of Agentic Coding

#agentic-coding #code-generation #ai-assistants #cursor #local-llm #repository-generation

A year ago, "AI coding assistant" meant autocomplete with a chat panel. This week's releases make that look like a different era. Repo0 generates entire software projects from natural-language requirements while keeping the architecture modular. Cursor shipped /goal, which turns "fix all flaky tests and make CI green" into a standing order the agent works until it's done. DeepSeek Harness v0.1.1 gave its agents eyes. Stampli cut launch hours by 68% with Codex. Replit made its agent free.

The thread connecting all of it: the model is no longer the bottleneck. The agent loop is. How you structure goals, context, vision, and architectural state determines output quality more than which checkpoint you load. Each signal matters on its own. Together they point in one direction: agentic coding is moving from "write this code" to "own this outcome."

Here are the numbers that matter this week:

68%: Stampli's cut in launch hours using Codex and ChatGPT Work 35%: Cursor's merged PRs created by cloud agents +29.74 points: Repo0's Pass Rate gain over RPG on RepoCraft 90k vs 67k: context kept uncompressed by PI Agent vs OpenCode on a 100k window

Repo0: architecture as a first-class state ​

Most code-generation agents assume the repository structure already exists. That works when the task is "add a feature to this codebase." It falls apart for zero-to-all generation, where the agent starts from an empty directory and a paragraph of requirements, and has to grow an entire project without turning it into a pile of tangled modules.

Repo0, from a new arXiv paper, treats architecture as an explicit, evolving state. It maintains a Dual-DAG: a requirement-level DAG, a component-level DAG, and an alignment relation between them. Starting from natural-language requirements, the agent iteratively evolves component boundaries through structural actions guided by modularity metrics, until the architecture converges. Only then does test-driven development generate the code.

The results are hard to wave away. On six real-world repositories from RepoCraft, using GPT-5 mini and DeepSeek V3.2, Repo0 beats RPG, the strongest repository-planning baseline, by up to 20.08 percentage points on Functionality Coverage and 29.74 points on Pass Rate. A 29.74-point pass-rate jump is the difference between generating plausible code and generating code that actually runs.

The ablations back up the design. Drop the Dual-DAG state, or skip the modularity-guided evolution, and performance falls off. Architecture is load-bearing in zero-to-all generation. It keeps a growing codebase from collapsing under its own weight.

Same model, two harnesses, wildly different results ​

While the paper side pushes architecture, the local LLM community is doing the empirical work on harnesses. A Reddit comparison caught my attention because it matches what I found in my own testing. I ran the same Qwen 3.8 27B checkpoint through two agent harnesses, PI Agent and OpenCode. The model fits on a single RTX 3090, no cloud GPU needed. Same model. Same hardware. Different loops.

PI Agent used fewer tokens, had no hard 32k output cap, didn't freeze, and compressed context far less aggressively. With a 100k context window and a 32k output setting, OpenCode started compressing at 67k of context. PI Agent held out until 90k. That 23k of extra uncompressed context is the difference between the agent remembering the function signature it wrote three steps ago and silently re-deriving it wrong.

The 32k output cap matters more than people think. Long refactors get truncated mid-edit. The agent starts a change, hits the cap, and the tail of the edit never happens. You get a half-migrated file and no error message.

The takeaway: benchmark harnesses, not just models. A checkpoint that looks weak in one agent loop can look strong in another.

Quick take: the model is no longer the differentiator in agentic coding. The loop around it is.

Vision is becoming a standard agent sense ​

DeepSeek Harness v0.1.1 added DeepSeek-V4-Flash-Vision-Exp, a multimodal visual understanding model. The /goal and /plan commands now accept image input, MCP and ACP support persistent image attachments, and PTC Mode forwards nested images. The agent can see screenshots, not just read text.

The Reddit thread makes the same point from the local side. The user's rule: always use a vision module, because the model uses vision to assess its own output quality. They offload the vision model to RAM, since 3 seconds per screenshot is fine when you only capture a few during a codegen or debugging loop. GPU vision at 0.3s doesn't matter when you're not doing real-time rendering.

This feels like a small change until you see the effect. Text-only agents can't tell that the button is off-center, that the modal overlaps the footer, or that the page is blank because a CSS file failed to load. An agent that can see its own output catches these without a human in the loop.

Cursor /goal: from tasks to KPIs ​

Cursor's latest update is the clearest product signal yet that agentic coding has moved from "help me write code" to "own this outcome."

The /goal command does exactly what it sounds like. Open a conversation, type "/goal fix all flaky tests and make CI green," and leave. The agent works until the goal is met. Your instruction stops being a task and becomes a KPI.

Three supporting pieces make it work. Subscriptions let the agent attach to event sources: a PR, a Slack thread, a scheduled task. Something happens, the agent wakes up. Cloud agents auto-subscribe to the PRs they open and push them to merge, fixing CI failures and responding to bot comments without a human clicking anything.

Sub-agents now get isolated VMs, each with its own project copy and clean context. They can verify the main agent's changes in a fresh environment, or work on separate problems in parallel without stepping on each other.

Steering fixes the interrupt problem. Your follow-up messages queue until the agent's next tool call instead of breaking its flow. Previously, typing while the agent worked was like shouting at a waiter carrying soup.

The number Cursor published is the part to sit with: over 35% of internally merged PRs are now created by cloud agents. That was 30% in February 2026 and 0% eighteen months before that.

The reaction to /goal splits along predictable lines. Local-agent users point out that Cursor is catching up to what open-source harnesses have done for months, just with better product polish. Cloud-side users worry about the cost. The framing that keeps coming up: the programmer is becoming a project manager, setting KPIs and reviewing work instead of writing code.

Production numbers and the cost reality ​

Stampli used Codex and ChatGPT Work to compress weeks of launch production into days, cutting launch hours by 68%. A fixed deadline, design resources committed elsewhere, and an agent doing the work. No leaderboard involved.

Replit took a different angle: remove the cost barrier. Its new Free Mode, powered by GPT-5.6 Luna, lets anyone build software without worrying about token costs. The bet is that agentic coding's adoption bottleneck is economic, not technical.

Cursor's pricing shift points the same direction. Starting August 24, Auto moves to per-model billing. The flat per-million-token price goes away. Route to an expensive model, pay the expensive rate. Letting agents grind all night becomes a line item on the invoice.

SignalSourceNumberPractical meaning
Stampli launch hoursOpenAI case study68% reductionWeeks of launch work compressed into days
Cursor agent-created PRsCursor engineering35% of merged PRsOver a third of merges with no human touching the code
Cursor Auto pricingCursor changelogPer-model billing from Aug 24Grinding agents becomes a metered cost
Replit Free ModeOpenAI announcement$0 token costRemoves the cost barrier for first-time builders

Common pitfalls ​

Five mistakes keep showing up across the paper, the community threads, and the product changes.

  1. Blaming the model when the harness is the problem. The same Qwen 3.8 27B checkpoint produced meaningfully different results in PI Agent vs OpenCode. Before you swap models, test your agent loop.

  2. Running text-only agents. If the agent can't see its output, it's flying blind. Add a vision module, even a slow one offloaded to RAM.

  3. Treating /goal as fire-and-forget without wiring up event sources. An unattended agent needs something to react to. No CI subscription, no PR webhook, and the agent sits there waiting for a goal it can't observe.

  4. Ignoring context compression thresholds. Defaults like OpenCode's 67k compression point silently degrade long tasks. If your agent forgets earlier decisions, check when compression kicks in, not the model's context window spec.

  5. Letting zero-to-all agents free-build without architectural constraints. Repo0's 20 to 30 point gains come from explicit architecture state and convergence checks. Without structure, a generated repo becomes a monolith with extra steps.

One thing to remember: the gap between a mediocre agent and a great one is rarely the model. It's the loop. Context handling, vision, event wiring, and architectural state are where the 20 to 30 point improvements are hiding.

What to do next ​

Three takeaways, depending on where you sit.

If you're building a codegen agent, adopt an explicit architectural state like Repo0's Dual-DAG. Modularity-guided evolution beats free-form generation by 20 to 30 points on real repositories.

If you're running local agents, test your harness before your model, and add a vision module. The same checkpoint can look mediocre or excellent depending on the loop around it.

If you're adopting cloud agents for production work, wire up event subscriptions and use /goal for outcome-based tasks, but watch per-model billing. Agents that grind all night are powerful, and increasingly metered.