Skip to content

AI Promoted You to Reviewer. Your Reviewers Can't Fail.

#ai-coding-agents #verification #test-quality #developer-workflow #code-review

AI Promoted You to Reviewer. Your Reviewers Can't Fail. ​

Here's a number I can't stop thinking about: 11%.

That's the share of automated checks in one developer's repositories that can prove they can fail. He counted 204 conclusion-bearing guards across three repos, the tests that stand in for a human reviewer. Only 22 had ever been shown a known-bad input. The other 89% were green. Not because everything was fine. Because they were incapable of finding anything, and nobody could tell the difference.

It's one developer, three repos, a case series rather than a sample. But the shape survives any correction.

This is what happens when AI moves your job. You used to spend most of the day producing artifacts and a little of it verifying them. Now an agent produces most of the artifacts, and your job is verification. Which means your real codebase, the one your judgment actually ships through, is those 204 guards. And that codebase is held to a standard you would reject in application code. No test coverage. No review of the reviewer. Green as the default state, silence booked as success.

The model isn't the weak link. The unfalsifiable green checkmark is.

The green-and-blind check ​

Let me define the terms, because the distinction is doing all the work.

A conclusion-bearing guard is any test that reads source code, config, or system state and asserts a claim about it. Not "does this function return 4", but "no workflow downloads its cache over the network", "every page passes the same quarter filter", "this feature flag matches the deployed spec". These are the tests that stand in for a human reviewer.

A negative control feeds that guard a known-bad input and asserts it gets rejected for the expected reason.

The gap between them is the whole problem. "Found no violations" and "is incapable of finding violations" produce identical output. Only a negative control separates them.

The failure mode shows up everywhere once you look. A deploy gate that existed to catch a silent failure mode, a missing tool falling back to an empty result, called node -e to parse a health response. The deploy runner has no Node. Six consecutive deployments failed with exit 127. The check against missing tools failed on a missing tool, and nothing shipped for six hours. The step had been green in review because nobody had ever run it where it actually runs.

Or the harvester that collected data from public repositories and judged each run by exit code. One run wrote seven perfectly good records, then hit a non-fatal warning and exited non-zero. The machine booked its own completed work as "failed, retry later". Interrupted-with-partial-results had no representation, only success and failure.

Or the error classifier that looked for server errors with the pattern 50[024], anywhere in the output. It matched the "500" inside "4258 of 5000 quota points remaining" and classified a successful run as a server failure.

Three different systems. One shape: the check watched a messenger, an exit code, a pattern, a status, while the artifact that mattered told a different story.

The rule that survives: judge the artifact, not the messenger. Exit codes are messengers. Summaries are messengers. The agent's own "done" is a messenger. Green badges are messengers. The artifact is the diff, the file on disk, the served response body, the row in the database. When a messenger and an artifact disagree, the artifact is right. A check that only ever reads messengers should be treated as unverified, however green it is.

Key Numbers11%: conclusion-bearing guards with a negative control in one developer's three repos (22 of 204) 89%: guards that have never seen a known-bad input 4,465: human Kaggle trajectories in TraceML, vs 207 agent trajectories 40% vs 13%: ARC tasks solved by Narcissus vs raw LLM proposals

Quick Take: The bottleneck in AI-assisted development is not code generation. It's verification, and most verification layers can't fail.

Agents collapse into narrow loops ​

The same blindness shows up one level up, inside the agents themselves. TraceML, a new corpus from a research group, pairs human and agent work on the same Kaggle competitions under one version-level schema: 4,465 human trajectories across 134 competitions, seven of which were also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect.

Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold collapses into a narrow loop. Codex spends its steps re-weighting ensembles and tuning submissions. MLEvolve mutates its model in place. Neither pivots at the human rate, and neither reopens abandoned work.

A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores. But the effort profile stays agent-shaped. Instruction closes only the part of the gap that reduces to instructions.

The pattern is the same at both scales. An agent that never reopens abandoned work is an agent whose search space got pruned by its own first guesses. The algorithmic version shows up in program synthesis. Narcissus, a synthesizer for grammars rare in LLM training data, found that static guidance, which approximates LLM proposals into rule frequencies, prunes every rule the proposals miss, exactly when the proposals are wrong. Its fix is a regularization term that keeps every rule reachable, so wrong proposals delay the solution but cannot hide it. It solves 40% of ARC tasks where the raw proposals solve 13%, all without a single LLM call during search.

That's the deep lesson hiding in all three results: the agent's proposal should not become the boundary of your search. Humans reopen abandoned work. Narcissus keeps rules reachable. Good workflows keep alternatives alive in writing.

What actually works: written plans and proof over claims ​

Developers who get the most out of coding agents have converged on a workflow that looks like bureaucracy and isn't. Design, plan, execute. Before touching code, a conversation about the problem, two or three approaches with their downsides, which one gets picked and why. That becomes a short written design doc. The design turns into a task-by-task plan: which files each task touches, which test gets written first, which command verifies it, where the commit goes. Then the code.

The written plan is what turns "the agent did something weird" into "the agent went off script at step 4". Those are very different problems, and only one of them is hard to fix.

I keep a store of technical lessons rather than plan files. I counted it this week: 524 records, 257 of them (49%) have been overwritten at least once, 408 overwrites total. Half of everything I wrote down needed correcting. A plan that turned out to be the wrong approach sits in the same directory as the one that worked, with the same weight, and three months later both read as history. That's why I require the outcome note on updates and leave it optional on first write. The writer is usually the agent, mid-session, right after it diagnosed the thing. It has the reason in context, and the field costs it half a sentence.

The same principle shows up in autonomous game generation. godogen, a generator that builds Godot, Bevy, and Babylon.js games with Claude Code or Codex, runs on "proof over claims". The agent judges results from the running game, a live URL or a recorded clip, not from a clean compile. Visible defects drive the next iteration. That's the artifact rule applied to the agent itself.

The verification discipline is spreading to how people learn, too. The ai-engineering-from-scratch curriculum, 511 lessons on building AI systems end to end, requires evidence for every lesson: the command, the working directory, the exit code, the meaningful output, and the artifact you changed or produced. Same shape as a negative control. You can't claim you learned it if nothing downstream can tell the difference between shipped and working.

The dead-air problem ​

There's a third problem hiding in plain sight: the dead air. Agent runs take 5 to 20 minutes, and they don't fit the deep-focus blocks that used to define productive work.

The filter that works is the 30-second drop rule. If you can't drop the task in under 30 seconds when the agent finishes, it doesn't belong in the gap. Paperwork passes. A practice quiz passes. Slack fails, and it fails spectacularly. You open it for "30 seconds" and return 47 minutes later to find the agent finished, the coffee cold, and you the bottleneck.

The inversion that keeps coming up in the community: the best thing to do while AI codes is prepare the proof that its code shouldn't be trusted yet. Sharpen the acceptance criteria. Inspect adjacent tests. Write one negative control for the incoming diff. Then agent latency turns into verifier time without paying a context switch.

TaskPasses the 30-second rule?Why
PR review paperworkYesDrops instantly, state is external
Practice quizYesFits 5-minute chunks, nothing to rebuild
Negative control for the incoming diffYesCheckpointed on disk, turns wait into verifier time
Slack "quick" questionNoInterruptible in theory, never in practice
A meetingNoCan't drop in 30 seconds
Starting a second projectNoRestart cost lives in your head

What the community is saying ​

The comment threads on these posts converge from different directions. I ran the Slack experiment more times than I want to admit. I also hit the temporal version of the courier problem: the agent that fixed a subtle bug with me on Tuesday starts Wednesday knowing nothing, and I re-explain the project every morning like a colleague with nightly amnesia.

The sharpest exchange was about the router problem. When you route reviews across models, the router itself becomes a third reviewer, and a misrouted hard call is the new silent failure. Each tier only sees the traffic the router decided was its own, so a misroute never surfaces as a tier regression. It just becomes a quietly wrong cheap answer. The fix: pin an adversarial slice, known-hard-disguised-as-mechanical, and replay it through the classifier every release. Regression-test the reviewers, and regression-test the thing that decides who reviews.

I've been the punchline too. While building the feature this data comes from, my equivalence test failed by exactly 0.25, and the bug was in my test, not the code. Twice in one evening, I failed to test my own reviewer while writing this. That's not irony. That's the base rate, and it's why conventions beat discipline.

This gets worse as tools spread beyond engineers. loveholidays is using Codex to make software development accessible across the business. That's a productivity win, but it's also a verification problem with a bigger surface. Everyone becomes a builder. Everyone becomes a reviewer. Most of them have no training in what a green check can't tell you.

Common Pitfalls ​

  1. Trusting green checks that never saw a known-bad input. If a guard has no negative control, "found no violations" and "is incapable of finding violations" are indistinguishable. The fix is boring: wire one known-bad case through the live path. If the reviewer ever waves it through, the pipe is broken.

  2. Reading messengers instead of artifacts. Exit codes, the DOM, commit messages, the agent's "done". All messengers. The artifact is the response body, the file on disk, the diff. Verify at the surface your actual consumer reads, not the one that's easiest to reach. The browser was one console.log away; the response body needed curl. Three keystrokes of difference, and the check silently measured the wrong world.

  3. Describing your diagnosis instead of the behavior. Agents accept your premise too fast and are too agreeable about your ideas. If you tell the agent "fix this CSS bug", it will go looking for a CSS bug. If the real problem was route ordering, you get a CSS fix that covers the symptom. Describe the behavior, and ask for two or three options with their downsides before anything gets decided. If you don't ask for alternatives, they don't show up.

  4. Letting context live only in the conversation. Decisions that live only in a chat get lost. The file doesn't forget. Write the plan down, and write the outcome note when the plan turns out wrong. Require the outcome note on updates, leave it optional on first write.

  5. Skipping the smoke test on no-build projects. A page that renders beautifully and does nothing is the classic false positive. One missing import, dom.js vs dom_js.js, disabled every event listener on the page for eleven months. "It renders" is not a release test. Load the page, fail on any request over 400, exercise one interaction. Twenty lines.

The same applies to generated UI. In a head-to-head test of five design-to-code tools fed the same dashboard and the same prompt, the tool with the strongest visual redesign also had the cleanest code, but the correlation didn't hold across the board. One tool produced a polished interface and left more cleanup work than the others. Judge the code, not the screenshot.

One Thing to Remember ​

A green zero is the most dangerous answer a check can give. "Found no violations" and "is incapable of finding violations" produce identical output. Only a negative control separates them. And the same rule applies to the agent itself: proof over claims. Judge the artifact, not the messenger.

The Bottom Line ​

If you're building verification for agent-generated code, adopt negative controls as an admission rule, not a reaction. No checker gets into your pipeline until it has proven it can fail on a known-bad input. That's the one structural fix that beats the incident-first ordering problem.

If you're using a coding agent day to day, write the plan down before you let it execute, and read every diff. The written plan turns "the agent did something weird" into "the agent went off script at step 4", which is a much easier problem to fix. If you're not going to read the diffs, don't delegate.

One thing to watch: the router problem. As teams move to multi-model review pipelines, the classifier that decides who reviews is becoming the new silent failure point. Expect known-bad regression streams pinned to reviewer channels within the next year, because the alternative is "89% with a smaller invoice".