Appearance
Two headlines, one week
Between September 1 and September 3, 2026, the frontier shipped more capability than it usually does in a quarter. Anthropic released Claude Fable 5.1 on the 1st. Google and Meta followed with Gemini 3.8 Flash and Muse Spark 1.3 on the 2nd. Early on the 3rd, OpenAI launched GPT-6 Astra, after a two-week pause on some reinforcement learning runs because internal review flagged the model's cybersecurity ability as potentially "Critical."
Two results from that week rise above the launch posts. The first: GPT-6 Astra's agentic gains are the most visible jump OpenAI has posted in a single generation, concentrated in operating computers directly and sustaining long tasks. The second: a Claude-based system produced the first end-to-end machine-verified proof of Fermat's Last Theorem, an 11-day run that formalized more than 29,500 intermediate theorems in Lean. One is a benchmark story with an asterisk. The other is verifiable in a way benchmark scores never are.
The AGI talk was loud. Greg Brockman told reporters it was "not unreasonable" to feel the industry is now in the AGI era, and the Reddit thread that followed asked the right question: if these models are AGI, why do remote workers still have jobs? That gap between rhetoric and measured reliability is where the honest discussion lives. The measurable changes follow.
Access is staggered, which added its own friction. Only a small set of institutions got Astra on day one. Plus, Pro, and Business tiers get it over the coming days, and the API opens gradually. OpenAI offered paid users a credit reset for each day they wait, a rare preemptive apology.
What the numbers actually mean
OpenAI's benchmark suite shows Astra hitting ceilings in places that separated frontier models six months ago. Each score below is paired with what it means in practice.
| Benchmark | Score | Practical read |
|---|---|---|
| ARC-AGI-3 (with harness) | 99.9% | Tests that frontier models scored under 1% on in March now saturate. |
| ARC-AGI-3 (no harness) | ~60% | The unassisted number, per the release thread. Strong, but a 40-point gap. |
| ExploitBench | 100% | Turns known vulnerabilities into working exploits. This is why the safety pause happened. |
| FrontierMath Tier 4 v2 | 97.6% | Research-grade math reasoning is no longer a separator at the frontier. |
| ScreenSpot-Pro | 92.7% | Up from 76.9%. The model knows where to click roughly 19 times out of 20. |
| OSWorld 2.0 | 72.6% | Modest score gain, but average task time dropped from ~75 minutes to 40. |
| Terminal-Bench Science 0.1 | 64.6% | Up from 22.4%. Near-tripled on real scientific workflows. |
Two patterns stand out. The near-ceiling scores are all in narrow, high-stakes domains: novel reasoning, exploit construction, hard math. The big jumps are all in agentic work: science workflows, GUI operation, multi-step computer tasks. That split tells you where OpenAI actually spent its gains.
The harness asterisk
The 99.9% ARC-AGI-3 number comes with a harness. Without one, the same evaluation lands around 60%. That changes how you should read the headline. "With harness" means the model gets scaffolding that decomposes the task, tracks progress, and replans. That's a legitimate system, but it isn't the same thing as a naked model output, and the two numbers get conflated constantly.
Third-party composites reinforce the caution. Artificial Analysis's intelligence index puts Astra's highest reasoning tier at 61, identical to GPT-5.6 Sol, and fifth overall, below Claude Fable 5.1 at 66. So on broad reasoning, Astra didn't separate from its predecessor. The gains are real, specific, and concentrated. If your workload is long-horizon or GUI-driven, this is a big release. If you need breadth of reasoning, you won't feel the difference.
Quick Take: GPT-6 Astra didn't get broadly smarter. It got much better at long tasks and at operating software directly, and those are the numbers worth benchmarking yourself.
Computer use, judgment, and the thinning agent layer
Two Astra changes deserve separate attention. Computer Use lets the model read a screen via screenshots, decide the next action, convert it to clicks and keystrokes, then re-read the screen and loop. Judgment decides which small ambiguities the model resolves on its own and which it escalates to you. One handles execution, the other handles when to ask for permission. Together they attack the biggest obstacle to real agent work: supervision cost. Every escalation you don't have to make is a step toward a task that finishes without a human hovering.
The early test reports line up with the scores. I spent a week letting Astra build a Manhattan block in Unreal Engine; it kept adding buildings and street details day after day without losing the thread, though at its pace finishing all of New York would still take months. Another test pushed six projects at once, 3D models, animations, a complete game prototype, and hundreds of agents ran in parallel without visible degradation. The most striking one started with a house floor plan, built the 3D scene in Blender, ported it to Unreal, and produced a walkable space where the model handled the skybox, cinematic lighting, and camera angles after the first few prompts. Playco, the mobile game studio, reported 50% fewer manual fixes when prototyping three themed games from a single grey-box foundation.
The structural effect is that the agent layer gets thinner. Screen-reading, common tool use, and basic judgment are sinking into the base model. What stays valuable in vertical agents is domain knowledge, business process, permissions, and liability boundaries. If you build agents for finance or healthcare, the moat was never the click loop. It's everything around it.
I remember the earlier attempts at this, and they were rough. OpenAI's Atlas browser let a model drive a browser directly, but the latency was brutal and the results unreliable. The jump from "technically works in a demo" to "runs a week-long build in a professional tool" is the delta that matters this generation.
Fermat, formalized
The other headline was quieter and arguably bigger. Anthropic announced that a Claude-based system produced the first end-to-end machine-verified proof of Fermat's Last Theorem.
The backstory matters. Fermat wrote the conjecture in a book margin in 1637, claiming a beautiful proof that didn't fit in the space. Mathematicians chased it for 350 years. Andrew Wiles finally proved it in 1995 with a 129-page argument using mathematics that didn't exist in Fermat's century, and even his first public attempt had a fatal gap that took him a year to repair. Verifying proofs of this scale has always been the bottleneck. It takes experts months or years, and history is full of accepted results that later cracked.
Formal verification removes that bottleneck. Translate the proof into Lean, a language a computer can check, and every inference step gets mechanically verified. The catch: formalization was considered a multi-year engineering project on its own. Imperial College's Kevin Buzzard, who leads a formalization effort, produced an 86-page blueprint just for phase one.
Claude did it in 11 days, mostly autonomously.
Key numbers from the Fermat run
- 11 days of largely autonomous work
- 13 million lines of Lean code
- 30,300 theorems produced; 29,500 adopted into the final proof
- 6 billion tokens consumed across the entire campaign
- 5x the size of Mathlib, the largest math theorem library
- Largest Lean proof ever written
The 13 million lines and 29,500 auxiliary theorems are the point. The system didn't prove one theorem. It rebuilt a large chunk of formalized mathematics from the ground up, including areas like algebraic geometry and harmonic analysis that had never been formalized. The Lean compiler checked every step, and the root node eventually read PROVED. Even the model's internal log flagged the moment, which the researchers shared: "historical moment."
Note the token figure: 6 billion for the whole campaign, spread across the parallel agents, the dead ends, and the re-proving work. That's 6 billion, not 60. The Chinese-language reporting uses 60亿, which is 6 billion in English, and it's easy to inflate by 10x on conversion. The number is large but not astronomical, and the wall-clock time of 11 days is the impressive part.
Why a Lean proof is a stronger signal
Benchmarks ask a model to produce a right answer, and the scoring harness is part of the system. A Lean proof requires every single inference step to survive mechanical checking. No harness, no judgment call, no partial credit. That distinction is why I'd argue the Fermat result is the more meaningful capability signal of the week. You can argue about whether a harnessed 99.9% counts. You can't argue with the Lean kernel.
The human track record here is sobering. The Kepler conjecture took a review panel four years, ending in "99% certain." Perelman's Poincaré proof took the entire field years and three 300-page books to digest. Some theorems were accepted as true, built upon, and later found to have broken foundations. A machine-checked proof changes the economics of confidence. Verification that once required a decade of expert attention now runs in days.
And it isn't confined to Anthropic's internal setup. Three ordinary accounts on the Prove2Me platform formalized Vinogradov's three-primes theorem in three days. The tooling made frontier-scale verification accessible to people working with consumer AI subscriptions.
The orchestration lesson
The Fermat result almost didn't happen. Early runs with dozens of parallel Claude agents collapsed quickly; they lost track of each other, produced incompatible work, and coordination broke down. Researchers said the early failed attempts contributed only about 7% of the final code. Raw model capability was never the constraint. Orchestration was.
The team, led by Peng Tianyi, a Tsinghua Yao Class graduate with an MIT PhD who now splits time between Columbia and Anthropic, built a coordination platform called Prove2Me to fix it.
Three design choices mattered. A theorem DAG gave every agent a map of which dependency to prove next, which fought the models' memory loss on long runs. Separating statements from proofs sped up compilation and saved resources. A natural-language index let agents find and reuse each other's results instead of re-proving them.
What the community is saying about the result splits in instructive ways. The engineering crowd read it as proof that orchestration, not parameter count, is the new frontier: I found myself agreeing that a bare model, however large, wouldn't have made it through a week of coherent multi-agent math. Others joked that a 13-million-line proof is mathematical PUA, too long for any human to check, so you accept it and move on. Some mathematicians were more skeptical, asking what new mathematics this actually produced. The fair answer is that formalizing old theorems is groundwork: it creates the verified substrate you need before you can trust machine-generated new math.
Common pitfalls
People will trip on this week's news in predictable ways. Here are the ones I'd watch for.
Reading harnessed numbers as raw capability. The 99.9% ARC-AGI-3 requires a harness. Without it, it's roughly 60%. When you compare models, compare the same setup, or you're comparing scaffolding engineering, not intelligence.
Inflating the Fermat token figure. The campaign consumed 6 billion tokens, which is 60亿 in the original Chinese reporting. It's an easy 10x conversion error, and the inflated number changes the takeaway from "impressive orchestration" to "absurd compute." It was the orchestration.
Trusting either OpenAI's benchmarks or third-party composites alone. OpenAI's numbers show huge gains; Artificial Analysis shows Astra flat with Sol at 61. Both are correct, because the gains are concentrated in agentic and GUI tasks. Pick the evaluation that matches your workload, or you'll buy the wrong model.
Assuming OSWorld-level scores mean hands-off automation. 72.6% means one in four computer tasks still fails, and the 40-minute average is better but not fast. Judgment cuts escalations; it doesn't remove failure modes. Plan for a human review layer.
Ignoring the price change. The API went from $4/$20 per million tokens on Sol to $10/$50 on Astra, and Fast Mode doubles speed at double price. Long agent runs burn tokens quickly. A 40-minute OSWorld task at 2.5x the token price changes your unit economics.
The cost of the frontier
Astra's API pricing lands exactly where Anthropic's does, and that's the business story underneath the capability story. OpenAI needs to justify the premium on task completion, not on brand.
| Model | Input per M tokens | Output per M tokens | Notes |
|---|---|---|---|
| GPT-6 Astra | $10 | $50 | 2.5x GPT-5.6 Sol. Fast Mode: ~2x speed at 2x price. |
| Claude Fable 5.1 | $10 | $50 | Price parity with Astra. |
| GPT-5.6 Sol | ~$4 | ~$20 | Prior generation pricing. |
The revenue picture explains the pressure. OpenAI's Q2 revenue hit $6.7 billion, up 18% quarter over quarter, with an operating loss of $12.3 billion. Anthropic posted over $11.5 billion in Q2, more than doubled its revenue, and recorded a small operating profit: the first quarter Anthropic's revenue topped OpenAI's. Annualized run-rate estimates at the end of July put Anthropic above $65 billion and OpenAI around $40 billion, though the two companies count revenue differently. Enterprise spending has shifted hard toward agents. By June, Codex accounted for 64% of the output tokens from enterprise ChatGPT and Codex usage. Both companies have filed confidentially for IPOs. The compute bills, OpenAI's alone expected around $50 billion this year, are the reason the AGI era has a pricing page.
One thing to remember
The dependable signal from this week wasn't the AGI talk. It was that long-horizon reliability and machine-verifiable reasoning both took a visible step forward, and both required orchestration layers, Judgment for the agent work, Prove2Me for the math, rather than just a bigger model. Capability is becoming a systems property.
The Bottom Line
If you're building agent products, adopt GPT-6 Astra's Computer Use and Judgment despite the 2.5x API price, because task time on OSWorld dropped from ~75 to 40 minutes and ScreenSpot-Pro hit 92.7%. Budget for token burn and keep a human review pass.
If you're evaluating models for a specific workload, skip the headline scores and benchmark with your own long-horizon eval. The broad-reasoning composite shows Astra flat with Sol at 61, so pay for the jump only if your tasks look like its agentic gains.
If you're in research or any verification-heavy domain, start building on formal-verification tooling now. A multi-year formalization project became an 11-day run with Prove2Me, and within six months, machine-checked proofs will be the expected standard for AI-generated results in high-stakes math and safety work.