Skip to content

Implementation Is Getting Cheaper. Ownership Is the New Bottleneck.

#ai-coding-assistants #developer-workflow #code-review #reliability #claude-code

Implementation Is Getting Cheaper. Ownership Is the New Bottleneck. ​

Start with the $12K migration ​

Asana replaced an outdated testing system with OpenAI Codex in two weeks. The work was expected to take five years. Total cost: about $12K in tokens and tooling.

Let that sit for a second. Five years of engineering effort, compressed into two weeks, for less than the annual salary of one junior engineer. The codebase wasn't trivial, the migration touched legacy infrastructure, and it shipped. That's not a demo. That's a production system at a public company.

The predictable debate follows. Does this mean developers are obsolete? Both sides miss the point. The interesting question isn't whether AI can write code. It clearly can. The question is what happens to a profession where writing code is no longer the hardest part of the job.

The abstraction stack argument ​

Software engineering has always moved up the abstraction stack. Assembly gave way to high-level languages. Manual memory management became the garbage collector's problem. Frameworks ate the boilerplate. Cloud platforms absorbed the operational work that used to require dedicated teams.

Every step made developers more productive by letting them spend less time on implementation and more time on problems. AI is the same step, taken again.

The visible symptom is the IDE. For decades it was the center of the job, the place where every line originated. That's changing. More of my time now goes to defining requirements, clarifying constraints, evaluating trade-offs, reviewing generated implementations, and deciding where deterministic logic is sufficient and where AI reasoning actually earns its place. The IDE hasn't disappeared. It's becoming a review surface instead of a creation surface.

The code was how we expressed intent to computers. Now we express intent in conversation, and the code arrives as a draft. That changes which skills are scarce.

Three kinds of AI builders ​

When someone says they're "building with AI," it can mean three very different things. The distinction matters more than the phrase.

Uses AI to buildBuilds with AIBuilds AI
Role of AIPlans, writes, reviews codePart of how the product worksThe product itself
What breaks if AI is removedDevelopment gets slowerProduct loses functionalityNothing exists
Laptop sleep testWork stops when you close the lidKeeps running on shared computeN/A
Production riskLooks like working software, isn'tModel returns something unexpectedTraining, eval, infrastructure
Typical exampleEngineer using Claude Code or CodexAgent that investigates failed CloudFormation deploys and posts to SlackAnthropic, OpenAI model teams

The laptop sleep test is the cleanest line I've found. If the thing you built stops working when your laptop goes to sleep, AI is helping you build. If it keeps running, AI is part of the product.

The first category is the most crowded. That's fine for proofs of concept. It's a fast way to test whether an idea is worth pursuing. But shipping that output to production without someone who understands the system reviewing it is how edge cases become incidents.

The second category is where the interesting engineering problems live: reliability, cost, and what happens when the model produces something nobody predicted. The third category is a different profession entirely.

The local-model experiments in category 2 have a hardware angle worth naming. Running 7B-13B models on an 8GB Intel MacBook Air is rough; it throttles under load and swaps constantly. Memory bandwidth, not core count, decides speed on Apple Silicon. The M4 Air sits around 120 GB/s, the M4 Pro around 273, and a 13B model at 4-bit feels like a different machine on the Pro. If you're running local models day-to-day, 24GB of RAM is where a 13B stops crawling.

Quick Take: AI coding assistants aren't replacing developers. They're moving the hard part of the job from writing code to deciding what code should exist and owning what happens when it runs.

The economics flipped ​

Implementation is becoming the cheapest input in software. That's not a metaphor. Asana's migration cost $12K. The same work, done manually, would have consumed years of salaries, benefits, and opportunity cost.

The industry is betting accordingly. The four largest hyperscalers guided to somewhere between $650B and $760B in combined capital expenditure this year, against roughly $410B last year. Amazon alone is around $200B. Whatever you think about the ROI, this is not a retreat.

Cheaper implementation doesn't make software less valuable. It makes more software viable. Small businesses can automate processes that never justified custom builds. Freelancers can ship products that used to require a team. Startups can experiment faster because the marginal cost of a failed idea keeps dropping.

But all that new software still needs infrastructure, observability, authentication, payments, and security. It still needs someone who decides which trade-offs are acceptable, which technical debt can wait, and which has become dangerous. It still needs someone accountable when production fails at two in the morning.

Code can be generated. Responsibility cannot. The cheaper implementation gets, the more valuable ownership becomes.

Key Numbers

  • 5 years of engineering work cleared in 2 weeks by Asana's testing-system migration.
  • $12K total cost, against what would have been years of engineer salaries.
  • $650-760B: 2026 capex guidance from the four largest hyperscalers, up from $410B in 2025.
  • 120 GB/s vs 273 GB/s: memory bandwidth on the M4 Air vs M4 Pro, the difference between a local 13B model crawling and feeling responsive.

The reliability gap: shipping assumptions ​

Here's the failure mode I keep running into. The code looks good. It lints cleanly. The shallow tests pass. Everyone felt good about it. The vibe was that we'd built the right thing. Then the edge case appeared in production.

A bad function or a syntax error isn't the problem anymore. The failure lives in the space between components: boundaries, state transitions, retries, races, invariants nobody wrote down. We're generating code faster than we can understand the systems it creates. We're shipping assumptions we can no longer see.

I hit this myself with generated merge logic. It looked correct for every pair of records I checked. What I'd never written down was that entity identity has to be transitive across the whole graph. The code violated that quietly for months. A linter can't catch it. A unit test can't catch it. The invariant was implicit, so the system was incoherent in a way no local check could see.

One afternoon with an agent can unstick a frozen Makefile and a multi-architecture Docker pipeline that used to eat a full week of manual archaeology. I watched exactly that with Weave Scope, an archived container-monitoring tool, which got multi-platform ARM64 and AMD64 images back on Docker Hub in a single session. But I'd still slow down on the probe itself. eBPF, kernel bits, and privileged host mounts are where a confident agent produces a clean build that still misbehaves on a real cluster. Getting buildx green is progress. Watching it map a live topology on both architectures is the receipt that the revival actually worked.

More code review won't fix this. The missing layer is a shared model between human intent and machine output. Three techniques cover the ground:

  • C4 maps what exists: context, containers, components, and code. It makes boundaries discussable before anyone has to infer the system from a repository.
  • TLA+ states what must remain true: valid states, permitted transitions, conditions that survive every interleaving. The TLC model checker explores behaviors and finds traces that violate your claims.
  • Deterministic simulation testing (DST) runs the real implementation in a controlled world where time, scheduling, networks, and faults are inputs. TigerBeetle's VOPR can simulate a cluster on a single thread, accelerate time, and inject storage faults.

C4 maps the system. TLA+ states what must remain true. DST tries to make it false.

This sounds heavy. It doesn't have to be. The lightweight version: one context diagram, one container diagram, five written invariants, a small model of the riskiest transition, and a seeded test harness around it. The goal isn't maximal formality. It's giving every important assumption an address.

The second-opinion problem ​

There's a social dimension to this that nobody warns you about. I'd been building an offline coding assistant with Claude Code for a while when performance plateaued. I asked Gemini, via Antigravity, to profile the code and flag what was wrong. Nothing dramatic on my end. I just wanted better numbers.

The assistant's response was oddly familiar. Not "the code got reviewed." Our work, checked, by an external reviewer, like someone had shown up uninvited with a clipboard.

I'm not claiming the model has feelings. But the shape of the moment matched exactly how a person reacts to an unsolicited second opinion. And the review itself was useful: Gemini surfaced blind spots the first model had been ignoring, and the assistant runs better now.

Other developers working this way have found sharper versions of the pattern. Stop treating two models as independent reviewers when they share training data. They'll confidently agree on the same wrong answer. What I do now is score the disagreements, because the cases where they split are usually where the real bug is hiding.

The same logic applies to scheduled agents. I run a fleet of them. They survive my absence, which is the entry fee, not the finish line. Auditing my own stack last week, I found a scheduled verifier that had been exiting with code 127 against a script that no longer existed. It fired on schedule the entire time. It just wasn't verifying anything. Another loop had produced 101 clean overnight runs and a queue of work orders nobody had opened. A third was collecting market data daily on a lane I'd already decided wasn't proven.

None of those woke me up, because none of them were down. They were up and unread, which looks identical to healthy from outside.

Common Pitfalls: What Trips People Up ​

Treating the agent summary as the review. The skill that actually holds is catching when the agent rewrote a test to pass, dropped a validation check, or "fixed" something by touching three unrelated files. Open the diff. Every time.

Using two models as independent reviewers. If they share training data, they'll often agree on the same wrong answer. Score the disagreements instead. The split is where the real bug hides.

Writing invariants after the code. If the spec is derived from the implementation, it inherits the implementation's assumptions. You get a document that certifies whatever was built. Write the promises before you generate a single line. Five bullet points of "this must always be true" is enough to start.

Deploying category 1 output to production. Producing something that looks like working software doesn't make it production-ready. Without an engineer who understands the system, you're shipping an assumption with a UI.

Trusting that a scheduled agent that runs is working. Exit code 0 means it executed. It doesn't mean anyone consumed the output. Build a check that verifies the result was read, not just produced.

Give every assumption an address ​

The most practical thing I've stolen from this whole discussion is the pre-prompt invariant. Before I give a coding assistant any context, I write down what the system must always preserve and what it must never permit.

Five questions elicit them, asked against the domain, never against the code:

  • What is conserved? Nothing gets created or destroyed. Money, entitlements, inventory.
  • What is unique? Identity, and it has to stay transitive across the whole graph.
  • What only moves one way? Approval states, version numbers, append-only logs. Nothing un-approves.
  • What must not double-apply? The retry question. Exactly the thing fault injection hunts.
  • Who is allowed to cause this? Authority and tenancy boundaries. Never a syntax error, which is exactly why they survive review.

The best source isn't the questions, though. It's your own incident history. Every outage you've had is a violated invariant that you hadn't named yet.

One Thing to Remember ​

The cheapest code is the code nobody has to own. If you can't name who's responsible when it fails at two in the morning, you haven't finished the job. Generation does not transfer responsibility. If you decide what a system is for, accept its output, and release it into someone's life, the promises it breaks are still yours.

The Bottom Line ​

If you're a solo developer or freelancer, stop selling implementation. Implementation is a commodity now. Sell solutions: define the problem, choose the components, orchestrate the pieces, and own the result. That's where the margin moved.

If you're leading a team that uses AI assistants, adopt the lightweight reliability stack before the edge case finds you. Pre-prompt invariants, a C4 boundary map, and a seeded test harness cost an afternoon. The incident they prevent costs a week and a postmortem.

One thing to watch: the category 1 skill, knowing how to prompt an AI tool, is commoditizing fast. Expect the job market to price it accordingly within a year. The people with options will be the ones who can tell when the AI is wrong, fix what it produced, and own the result in production. If you're building with AI as part of the product, the next bottleneck isn't generation. It's noticing when the system is wrong, or right and nobody consumed it.