Skip to content

The Two Conversations in AI Safety Nobody Knows How to Reconcile

#ai-safety #ai-misuse #threat-intelligence #data-privacy #ai-policy

The Two Conversations in AI Safety Nobody Knows How to Reconcile ​

Same Week, Two Different Documents ​

Earlier this week, Jacob Coxon resigned from Anthropic specifically so he could say in public what he'd been saying inside the company for years: that OpenAI and Anthropic are "gambling with our lives," racing toward self-improving AI without acting responsibly. Coxon spent three years doing pretraining research at both companies. He's not a random voice.

Then it got stranger. Evan Hubinger, who runs alignment science at Anthropic, confirmed Coxon's core claim on the record. His words were precise: he earnestly believes AI could kill all humans, he puts the odds above 10 percent within the decade, and Anthropic doesn't have a plan for aligning superintelligence and isn't clearly on track to get one. Samuel Marks, who leads scalable oversight at the same lab, said similar things.

Same week, different document. Anthropic's threat intelligence team published its September 2026 misuse report, and it reads like it came from a different species of organization. No probability distributions. No superintelligence talk. Instead, a detailed account of a Russian state-nexus actor, tracked as GTG-20006, who used Claude to automate cyber espionage against Ukrainian and European military targets. More than 20 organizations hit. A credential database with over 300,000 national identity records stolen.

I work with companies deploying this technology, and I've watched how they absorb both stories. Nobody in those rooms is talking about extinction. They're talking about whether an agent with write access to the CRM is going to do something dumb at 3 a.m. with no one watching. They're talking about who signs off when the output is wrong and a customer gets hurt. Those two conversations have almost nothing to do with each other, and they're landing in the same news cycle, which scrambles the signal for anyone making practical AI decisions.

What the community is saying: the threads after Coxon's resignation split into two camps. Either it's marketing to make the tech sound more powerful than it is, or it's a genuine warning we should all be terrified by. I don't think either read is right. But nobody seems willing to say the uncomfortable part out loud: if the people building these systems can't agree on whether it's an existential threat, a mid-sized company has no basis for a risk assessment, and it will default to whatever the vendor's sales deck claims.

The Extinction Claim Got Specific ​

What separates this from every previous round of AI doom talk is that the claim now has a number and an owner. Hubinger didn't say "maybe." He said above 10 percent within the decade. To translate that into a decision frame: if you ran ten independent AI programs under identical conditions, you'd expect at least one to end in catastrophe. No regulator would approve that risk profile in any other industry.

The "no plan" part matters even more to engineers. Hubinger and Marks aren't saying the risk is unavoidable. They're saying their own lab, the one that markets itself as the safety-first frontier lab, doesn't currently have a path to aligning superintelligence and doesn't know what that path looks like. Three years ago, that statement would have ended a career. Today it comes from the person in charge of alignment science, on the record, and the industry's response has mostly been a shrug.

For a practitioner, the practical takeaway isn't the probability. It's that the self-described safest lab in the world is openly describing a hole the size of a decade in its safety case. If you're building on a frontier API, treat your own safety case as an independent problem. Don't assume the vendor solved it.

The Threat Report Reads Like a Different Species ​

Anthropic's September 2026 report covers eight months of disruptions across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation. The cases range from a network of fake dating apps built to defraud users to surveillance systems designed to identify and monitor dissidents. The report's central finding is gloomy and concrete: the labor and tooling gap that used to separate state-sponsored operations from lone individuals has collapsed. Sophisticated attacks no longer require sophisticated attackers.

The most instructive case is GTG-20006, which Anthropic links to the Russian intelligence actors known as Midnight Blizzard. The operator ran AI-driven workflows that handled reconnaissance, phishing infrastructure, malware deployment, and exfiltration. When an AI monitoring agent detected that a deployed implant had been fingerprinted by a security product, other agents rebuilt it and restaged it until it passed undetected. The human's main job was editing the Claude Code skills that drove the workflow.

The kill chain runs as a loop, not a line:

The models involved matter. Anthropic is explicit that the misuse ran on Haiku, Sonnet, and Opus. The flagship classes, Fable and Mythos, stayed out of reach in these cases, with one exception involving illicit distillation. What that means in practice: capability gating on frontier models pushes misuse down-market. Attackers pick cheaper models or distill the bigger ones, and the harm per attempt keeps falling even as the frontier gets safer.

Quick Take: the extinction warnings and the threat report are the same failure mode at two different timescales. The report is the extinction argument in present tense, complete with case numbers. The hard part is that one conversation demands action now, the other demands action before it's too late, and nothing tells you how to do both.

One Campaign, the Numbers Behind "Uplift" ​

Uplift is Anthropic's term for how much more harm an actor can cause with AI than without it, measured along speed, scale, and depth. The security community keeps debating whether AI can discover novel exploits. The report argues the bigger risk is workflow automation: adversaries can now run the whole kill chain at the pace their infrastructure allows. GTG-20006 sustained a multi-victim campaign that would have required a team of skilled operators a year earlier. Anyone who downloads an open-source offensive agent framework like PentAGI gets much of the same scaffolding.

One compromise chain, reduced to numbers:

Three hundred thousand national identity records, plus registry data for more than half a million companies, from a single VPN appliance compromise. For scale: 300,000 is roughly the population of a mid-sized European city, and it came out through one stolen credential. The same actor took over WhatsApp accounts through headless browsers and suppressed read receipts before bulk-exporting conversations. Two former high-level Ukrainian officials were among the targets. Another tool froze victims' security updates so new malware signatures never arrived.

None of this required a superintelligence. All of it used models you can buy by the token.

The Privacy Pipeline Nobody Consented To ​

A third conversation is running alongside these two, in full view, and the safety crowd mostly ignores it: what happens to the raw audiovisual data that AI systems are tuned on. Meta sold two million Ray-Ban smart glasses in 2023 and 2024 combined. In 2025, that jumped to seven million units. An always-on camera and microphone at face level, priced like a fashion accessory. The Swedish investigation by SvD and GP found the human layer that makes it work, 9,300 miles from Menlo Park, where the consent conversation happens in a different legal and economic world than the homes being filmed.

In Nairobi, data annotators at Sama, a subcontractor to Meta, review video and chat transcripts from the glasses to teach the assistant to see and respond. Workers described seeing people in bathrooms, getting undressed, having sex. Bank cards visible by mistake. A chat in which a man described a woman he wanted to have sex with. Meta's terms say that in some cases, interactions with its AI may be reviewed, and the review can be automated or manual. Manual means human. The user's only protection is a warning not to share anything they don't want retained.

Retail staff tell buyers the opposite story. We bought a pair from a Swedish retailer, declined the optional data-sharing prompt, and tried to run the assistant offline. It refused. The glasses asked us to reconnect, and traffic capture showed steady contact with Meta servers in Luleå, Sweden, and Denmark. Here's what the claims look like against the evidence:

What you're toldWhat was found
"Nothing is shared with Meta, you have full control."The AI assistant won't work offline. The glasses prompt you to reconnect, and network capture showed frequent contact with Meta servers in Luleå, Sweden, and Denmark.
"Everything stays locally in the app."Voice, text, and image processing for the AI assistant is mandatory and cannot be turned off. It happens on Meta's infrastructure.
Voice recordings are only used for product improvement if you actively agree.That holds for storage and training. But the assistant's core function requires streaming your voice, text, images, and sometimes video to Meta automatically.
Faces are automatically blurred before human review.Annotators described unblurred faces and bodies, especially in poor lighting.

Kleanthi Sardeli, a data protection lawyer at NOYB, calls it a transparency failure. Once content is fed into a model, the user in practice loses control over how it's used. The Swedish privacy authority's security specialist put it more bluntly: the user has no idea what's happening behind the scenes, and the data Meta collects is worth more than the glasses.

Here's the pipeline as it actually runs:

This pattern has a 2023 precedent. When Zoom updated its terms that August, the fine print gave the company broad rights over user content, which critics read as permission to train AI on meeting recordings with no opt-out. The backlash came fast, and Zoom walked it back, clarifying it wouldn't train on customer content without consent. But the mechanism worked exactly as intended for the weeks it mattered: consent by buried clause, removed only after public pressure.

The State as Customer, Regulator, and Target ​

Now put three facts next to each other. An AI targeting system called Lavender has, per 972mag's investigation, been directing bombing decisions in Gaza, generating target lists at a scale that compressed human review into a rubber stamp. OpenAI, meanwhile, just announced free license tiers and 50 percent off usage for federal, state, local, and tribal governments, plus expanded cyber defense support in partnership with the GSA. And the same week Anthropic published its threat report, OpenAI's Chris Lehane argued that stronger capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

The state is simultaneously the largest customer, the main regulator, and the most targeted victim of these systems. GTG-20006 attacked government ministries and diplomatic missions with AI assistance. The governments under attack are the same ones being offered subsidized access to frontier models. At least one government is running its own AI targeting pipeline in an active conflict. There is no clean line between defender and adversary in this stack, and no policy proposal is going to draw one.

Common Pitfalls ​

I keep seeing the same mistakes as people try to operationalize this mess.

Mistake one: treating the terms of service as the consent conversation. Zoom's 2023 clause demonstrated the playbook, and Meta's glasses terms run the same pattern at a larger scale. If you're deploying tools on consumer or public-tier contracts, assume your users' content is part of the training corpus. Read the terms yourself. If the vendor can't point to a clause that says otherwise, price the exposure into the product.

Mistake two: trusting upstream anonymization. Former Meta employees told the Swedish investigation that faces are auto-blurred before annotation. The annotators confirmed the blurring misses, especially in low light. If you build any pipeline that screens user-generated media, add your own filtering and access controls. A vendor's anonymization toggle is a best-effort feature, not a security boundary.

Mistake three: reading "no misuse on the flagship model" as "the model is safe." Anthropic found no malicious activity on Fable or Mythos except one distillation case. That doesn't prove those models can't do harm. It means their safeguards made them less useful to attackers, so the attackers moved to cheaper models. Evaluate a model tier by what an adversary can do with it, not by what the safety card says.

Mistake four: letting the extinction claim freeze your risk process. A 10 percent claim and a live CRM compromise are both real, but they live on different timescales and need different owners. If your risk register stalls because leadership is debating civilizational collapse, split the conversation. Daily operational risk goes to the security team. The long-horizon alignment question gets a named owner and a quarterly review. Track both. Let neither block the other.

Mistake five: benchmarking the model instead of the workflow. The threat report's real finding is uplift. Adversaries aren't waiting for a model that discovers novel exploits. They're automating kill chains they already understand. If you're a defender, test your detection against workflow acceleration, not single-shot capability. Assume the attacker's iteration loop is faster than yours. It is.

One Thing to Remember ​

If you keep only one thing from this month's news, keep this: the safety-focused lab's own alignment lead says there's no plan for the thing he's most worried about, and the same lab publishes the most detailed misuse intelligence in the industry. Both statements are true at once. That combination, existential uncertainty plus operational transparency, is the most honest position any frontier lab has produced. Build your governance as if both are true, because they are.

The Bottom Line ​

If you're deploying AI in a product that touches user content, treat vendor terms as the threat model. Assume your users' data is training data unless you can point to a clause that says otherwise, and demand transparency about human review from any vendor whose AI can see or hear your customers. The Meta glasses investigation shows that layer is real, human, and invisible to the people being watched.

If you're on a security team, assume the kill chain automation from the threat report is pointed at you. The scaffolding is public, the models are cheap, and the detection-and-rebuild loop means static signatures are obsolete. Invest in behavioral detection and response speed, because uplift means the attacker's iteration cycle is now measured in minutes, not months.

If you're in procurement or policy, the window Lehane describes is literally open. OpenAI is offering governments near-zero-cost access, Anthropic is publishing the intelligence that shows why governments need defense support, and the same governments are running AI targeting systems of their own. Expect binding safety requirements to arrive through procurement and audit channels within 12 to 18 months. Get ahead of them.