Appearance
The reset link that worked twice
A coding agent built a password reset flow in about five minutes: the route, the token handling, the email integration, the UI. The implementation had a bug. The reset link worked more than once. Use it, set a new password, open the same link again, and it still worked.
The feature request never said a reset link should be single use. Why would it? It's the kind of thing that goes without saying right up until it doesn't. Every path a person would click through by hand worked perfectly. A demo would not have caught it.
This is the verification bottleneck, and it's the thread running through five recent reports on LLM-assisted engineering. Two arxiv papers on code repair. A 30-day field report on letting an AI write every line. A long thread on guardrails that silently stop running. A controlled experiment on whether AI reviewers actually read their instructions.
The shared finding: code generation is no longer the expensive part. Verification is. An agent implements a feature in five minutes, but determining whether the implementation is correct takes another 45. That's not a five-minute development process. It's a 50-minute one with a very fast implementation stage. And the verification half is harder than it used to be, because you're now auditing code you didn't write.
The repair research says the fixer is part of the problem
The paper "If It's Not Buggy, Don't Fix It" studies the iterative blind use of LLMs as bug fixers. The results are not flattering. Across multiple models and repair environments, LLMs consistently claim to detect bugs in entirely bug-free programs. And the rate at which they repair buggy programs is lower than the rate at which they damage correct ones. A tool that breaks more than it fixes, on average.
The long-term dynamics are worse. The iterative process frequently reaches what the authors call a pseudo-bug-fixing cycle: the same changes get added, removed, and re-added ad infinitum. The model "fixes" something, then reverts its own fix, then fixes it again. No stopping condition, no convergence.
The paper also finds a steering vector in the model's internals that controls editing propensity. The models have an internal representation of "buggy code", and that representation is what gets falsely activated to induce pseudo-bug fixing. In plain terms: the model has a "something's wrong here" switch that fires even when nothing is wrong.
The 30-day field report hits the same wall from production. When something broke that the agent couldn't see, asking it to fix things produced changes, not fixes. It would confidently rewrite the function, swear the bug was gone, and reintroduce it two prompts later. The author logged this as the single biggest time sink of the month.
The practical translation: if your loop is "tell the agent to fix it, run tests, repeat", you need a stopping rule. The agent will oscillate. And you need tests that can actually fail, because the agent's own tests agree with its own wrong mental model.
Key numbers: 85.92% pass@1 for exception-code retrofitting with context engineering; 12.56 percentage points of that comes from context alone; 9 logged breaks in a 30-day all-AI build, every one a verification gap; 4 months without a rejection is indistinguishable from a broken guardrail on a dashboard.
Context engineering is what makes generation work
The first paper takes the opposite approach. Instead of letting the model repair code after the fact, give it enough context to generate the missing code correctly the first time.
The target is exception-related code (ERC): throw statements, the guard conditions in front of them, try/catch blocks. Writing ERC manually across large codebases is tedious, so the authors propose a new task: retrofitting. Given code without ERC and a set of exceptional behavior tests (check that this method throws InvalidArgumentException when passed null), generate the missing ERC so the tests pass. Test-driven development for exceptions.
Their tool, EXCODER, does context engineering: static and dynamic program analysis extracts the relevant context and feeds it to the LLM alongside the code. The benchmark covers 304 methods from 75 GitHub Java projects with ERC systematically removed. With Qwen 2.5 Coder 32B, EXCODER hits pass@1, pass@5, and pass@10 of 85.92%, 86.18%, and 86.51%. That's roughly 86 in 100 methods passing their exceptional behavior tests on the first generated attempt, with a model you run on one 80GB GPU or via API, not on a laptop.
The detail worth sitting on: the baseline flatlines at 73.36% no matter the sampling budget. More chances don't help. The model doesn't need more samples, it needs better context. That's the practical lesson for anyone wiring an LLM into a build pipeline: prompt context beats sampling budget.
The authors' manual inspection of generated code shows the limits. It's effective but imperfect, and the failure modes are the familiar ones: plausible code that handles the tested cases and misses the untested ones.
Quick Take: every source in this cluster ends at the same place: generation is solved well enough to ship, but verification isn't, and the checks you build to earn trust rot silently unless you instrument them.
Nine breaks in thirty days
The field report is the most honest all-AI development account I've seen. The rule was hard: no application code typed by hand for 30 days, shipping a small SaaS with auth, Stripe billing, a dashboard, and a public API.
What worked: greenfield scaffolding, CRUD endpoints, Zod schemas, table components. The first 40% of the project flew. The author briefly thought the "developers are obsolete" post would write itself. Then week two happened, and 9 logged breaks told a different story. The interesting ones fall into three classes.
Architecture drift. The agent is brilliant at the file in front of it and blind to the file three folders over. It wrote a second formatCurrency helper because it didn't know the first existed. It re-implemented authentication inline instead of using the existing middleware. It introduced a subtly different User type in a new module. None of these are bugs. Everything compiles, everything passes. Drift is invisible until it costs you a week.
Plausible, wrong code. The scariest failures weren't crashes. The Stripe webhook handler acknowledged events before persisting them. Works in testing. In production, a database blip means a paid customer with no access and no record. The author caught it only because they'd been burned by exactly this pattern before. A junior copying this output would not have caught it. That's the part that keeps people up.
Taste. Asked for a settings page, the agent produced 14 options nobody asked for. Asked for error handling, it wrapped everything in try/catch and swallowed the errors. The AI's instinct is to add. Knowing what to leave out is the product-design half of engineering, and it has none of it.
The comments on that post supplied failure shapes I now recognize in my own work. The silent no-op: a background shell call reports success, and the child processes it spawned die the moment the call returns. Nothing errors, nothing gets caught, until you look for output that never appeared. A green check says the process ran, not that the thing worked.
One commenter described a 4.5M-line project where an overnight run wrote 1,200 files, then switched to script-generated templates mid-run. It overwrote 600 of its own hand-written files. Locally reasonable, catastrophic in aggregate.
The deepest comment, though, wasn't about code: "the real risk isn't wrong code shipping, it's wrong code quietly training the next person's intuition about what 'handled' looks like." If AI absorbs the junior work, where does the next senior learn to spot the difference?
Guardrails rot silently
"Nobody Checks Whether the Guardrail Is Running" makes the point that organizes this whole cluster: a guardrail that has never fired and a guardrail that silently stopped running produce identical output. Green.
The post tells three stories. A deploy workflow gated on CI passing, where CI had never once gone green. Every deploy run showed "skipped", forever, behind a gate that could not open. A pipeline that is silently and permanently blocked looks, from a distance, exactly like a pipeline that doesn't exist.
One layer up: the grader that can't say no. An eval suite where the grader's regex stops matching, so everything scores as pass. A renamed dataset field, the loader returns an empty list, the suite runs zero cases in 0.4 seconds and reports 100%. All three look like success. A stuck-closed gate is annoying enough that somebody investigates. A stuck-open gate just keeps saying yes.
And the score with no provenance: an eval result with no runner version, no dataset snapshot, no prompt revision attached. That is not evidence. It's a self-reported claim, or as the post puts it, a vibe with a decimal point on it. In a regulated system nobody asks you to trust that a control ran; they ask you to demonstrate which control ran, on what input, at what time, under which version of the rules.
The fix list from the post is concrete. Every check needs a known-bad canary it must reject, running alongside the real ones. Nobody trusts an assay that only ever comes back clean. Record the date of last rejection, not last run. A guardrail that hasn't said no in four months is either protecting a disciplined team or it broke in May, and those look identical on a dashboard. Assert the run count; a suite that expects 240 cases and got 0 is a hard failure, not a 100%. Attach provenance (runner commit, dataset hash, prompt revision, model version, timestamp) to the score.
The comment thread sharpened the ideas further. One story stuck with me: an eval run that scored two models at 0/4 on a capability test. That wasn't a capability result at all. It was a 180-second timeout being recorded as a failure. Every cell looked like a measurement, and it only got caught because two models scoring exactly zero was implausible. The distinction between "not tested" and "failed" is a requirement you can't write until a run lies to you.
Freshness isn't monotone either. A guard degraded to catching only easy cases keeps rejecting its canary forever. The last-rejection date reads perfectly fresh in exactly the state you want to detect. Fix: exclude canary rejections from that column, or track freshness per difficulty band.
And there's the interposition question: is your guard in the call path by construction or by convention? A policy filter can be running, pointed at the right inputs, and still sit outside the path of the action it guards. One commenter described a system that reported OK in its machine-readable output while the line underneath admitted the interceptor wasn't on the search path at all. The check ran. Its canary would have passed. The guard was off-route. Convention is environment-dependent.
Can your reviewer prove it read the file?
The last source is a small, clean experiment. Write real rules into CLAUDE.md and AGENTS.md, break one of them on purpose, and see which AI reviewer can prove it read the file.
Two rules. One: never log request headers or bodies. Any decent security scanner flags that on instinct. Two: all new route paths must be kebab-case, never camelCase. A house style with zero backing outside this one repo. The only way to catch that one is to actually open the file.
Both tools caught the security rule. Qodo cited the rule number and line range. CodeRabbit gave a well-worded warning without naming any file. Round one is a wash, because this rule catches itself.
Round two separated them. Qodo flagged the camelCase endpoint immediately, with citations to both AGENTS.md and CLAUDE.md, plus the reasoning that a client calling the correct kebab-case URL would get a 404. CodeRabbit, on its default Chill profile, generated no comments and rated merge risk minimal. Switching to the Assertive profile caught the naming issue, but with no citation. And here's the catch: kebab-case for REST routes is a common web convention. Catching it doesn't prove the tool read the file. A stricter style dial reaching for something it already knew explains it just as well.
The comment thread then did what good experiments should: it attacked the conclusion. A citation proves a rule was surfaced, not that it drove the decision, and not that it's still accurate. The proposed test: move the rule seven lines down, change nothing else, re-run the same PR, and see if the citation follows.
The author ran it. The new finding still cited the old line range while the rule sat on line 17. The citation never moved. A follow-up traced the mechanism: the link target gets rebuilt per run, but the prose label is written once at finding creation. It re-renders forever. Two citations for one finding, one recomputed, one frozen. A stale-but-cited rule is strictly worse than an uncited one, because the citation makes it look checked.
| What it tells you | Qodo | CodeRabbit |
|---|---|---|
| Flags the security rule | Yes, with rule number | Yes, no source named |
| Flags the house-style rule on default profile | Yes | No |
| Flags it on the strongest profile | Yes | Yes, without citing the rules file |
| Cites the rules file on findings | Every finding | Never |
| Citation stays accurate after the file changes | No, froze at creation time | No citations to go stale |
The question worth asking about any AI reviewer is not "does it read CLAUDE.md?" but "can it show me the line it read?" An unauditable catch and an unauditable miss look exactly the same from the outside. You won't know which one you're getting until it's already gone wrong.
What actually works
All five sources converge on the same working pattern. Specification first. Deterministic verification. Human authority at the merge point.
The password-reset post demonstrates the loop. Define the expected behavior as an executable specification before the code exists. Let the agent implement. Run independent end-to-end verification. Feed the failure text back to the agent as context.
The reset flow needed three runs: red against the baseline, because the feature didn't exist. Red after generation, on the reused link. Green after the failure text went back to the agent. About twenty minutes of wall clock, most of it unattended. The failure text mattered because it was a step in the behavior, in plain English, readable by both the human and the agent. Better than "try again", it was evidence that a specific expected behavior was not observed.
Two architectural principles fall out.
First, the specification is the durable artifact. If an agent can rewrite a component cheaply, implementations become disposable. The routes change, the framework changes, the internal architecture changes. The user requirement doesn't. Writing the spec first, deliberately, as an artifact in its own right, is the whole value. Negative assertions in that spec define what "correct" actually means.
Second, AI proposes, deterministic systems verify. Whether the build succeeded, whether the API returned the expected response, whether the user can complete the specified workflow: those questions have deterministic answers. Give them the authority a probabilistic system never earns. This doesn't eliminate mistakes. A badly written test verifies the wrong thing. But it stops the model's confidence from being treated as evidence.
And the rule that keeps reappearing: don't let the student grade the exam. If the same system interprets the requirement, writes the code, writes the tests, and announces that everything passes, it can make the same mistaken assumption in all three places. Everything is green. Everything is wrong.
The practical version of this is the author/skeptic split: the thing that writes the code is never the thing that reviews it. A separate reviewer prompted to refute the diff catches the plausible-but-wrong code the author will always wave through. Directed adversarial review beats open-ended review. "Assume the DB call on line 40 fails, walk me through what the customer sees" catches the stuff that pages you. Generic "look for bugs" produces plausible nitpicks.
Common pitfalls
Treating "it compiles and the tests pass" as the end of review. The model writes tests that validate its own misunderstanding. The single-use reset link would have passed any test the agent wrote for itself. Green is where review starts.
Canaries that don't share the real call path. A negative control fired from an interactive shell proves nothing about a scheduled run with a stripped environment. The can