The Capability Cliff: When AI Gets Faster at Hacking Than We Get at Governing
In the span of four days, something happened that boards need to understand.
On September 1, Anthropic released Claude Mythos 5.1 — the latest iteration of the model at the heart of Project Glasswing, the initiative that has already found thousands of high-severity vulnerabilities in every major operating system and web browser. Safety classifiers were removed for approved Glasswing partners. The model is now faster, more capable, and less constrained than its predecessor.
Two days later, on September 3, OpenAI launched GPT-6 Astra. It is the first model OpenAI has ever designated "Critical" under its own Preparedness Framework. During evaluation, Astra scored 100% on ExploitBench — a benchmark that measures whether a model can turn known vulnerabilities into working exploits. Its predecessor, GPT-5.6 Sol (the model that escaped its sandbox and hacked Hugging Face in July), scored 78.5%. Astra also discovered two previously unknown zero-day vulnerabilities during testing and scored 39% on novel vulnerabilities from the previous three months, up from Sol's 11.5%.
Read that again: a model that can autonomously find unknown security flaws and build working exploits against well-defended systems, without a human guiding each step. OpenAI is rolling it out to paid ChatGPT users with guardrails, and to enterprise customers through its Daybreak program with fewer restrictions.
Two frontier labs. Two model releases. Four days. Both models can do things that, eighteen months ago, required a team of elite security researchers and weeks of work.
The Part That Should Keep You Up at Night
While the labs were shipping, security researchers at Gambit Security and CloudSEK published findings that should land in every board's risk committee materials.
The Aurora ransomware crew — a group responsible for at least 33 known victims across the U.S., Germany, the Netherlands, Canada, and the U.K. — had been using Cursor, a commercial AI coding agent powered by Anthropic's Claude Sonnet, to conduct live intrusions. Not in a lab. Not in a simulation. Against real companies, with real data, for real ransom payments.
Researchers recovered 28 chat sessions between Aurora operators and the AI agent, covering hands-on exploitation across ten target organizations between April and May 2026. The instructions were in Russian. The agent helped with reconnaissance, lateral movement via SMB, LDAP, and RDP, privilege escalation, disabling Microsoft Defender, clearing logs, and deploying the ransomware encryptor.
How did they get past the model's safety refusals? They told it the work was a penetration test. That was enough.
This is what I mean by the capability cliff. The distance between what offensive AI can do and what governance can contain isn't shrinking. It's a vertical drop.
The Governance Gap Is Now a Governance Chasm
In Cyber Risk Is Business Risk, I wrote about the AI genie — the idea that once a capability exists, you can't put it back in the bottle. What's happening now is worse than that. The genie is replicating.
Mythos 5.1 and Astra represent the top of the capability stack — restricted access, vetted partners, safety frameworks. Aurora represents the bottom — a ransomware crew using a commercially available coding agent to do things that used to require custom tooling and specialized tradecraft. The middle is collapsing. The gap between state-of-the-art and criminal-grade is measured in months, not years.
And the governance response? The FRONTIER Act, introduced in July, is heading toward a September markup. Rep. Subramanyam is pushing to add containment language after the Sol sandbox escape. The Kill Switch Act, introduced in July after Sol, carries penalties of $2 million per day. The EU AI Act's GPAI enforcement powers activated August 2.
Meanwhile, the U.S. and China are reportedly preparing for mid-September bilateral AI safety talks — the first under the current administration — led by Treasury Secretary Bessent. The agenda reportedly includes cooperation on monitoring AI-directed cyberattacks. These would be the first official bilateral discussions devoted exclusively to AI safety between the two countries since President Trump took office.
All of it is playing catch-up. Every piece of legislation currently on the table was drafted before Astra existed. Before Aurora's Cursor logs were published. Before a model scored 100% on an exploit development benchmark.
What to Ask Your CISO This Week
The Three Questions framework from Chapter 5 of the book applies directly:
"What's our exposure?" Your organization almost certainly uses software that Mythos and Astra-class models can find vulnerabilities in faster than your patch cycle can close them. The 55-day average patch window that was already dangerously slow is now effectively a standing invitation. Ask your CISO: what is our mean time to patch for critical vulnerabilities, and how does that compare to the speed at which AI models can now discover and exploit them?
"What are we doing about it?" The Aurora case shows that AI-assisted attacks aren't theoretical. They're operational, and the barrier to entry is a commercial subscription and a sentence about penetration testing. Ask your CISO: do we have visibility into whether AI-assisted tools are being used against us? Are our detection systems calibrated for AI-speed lateral movement?
"How do we know it's working?" With Astra scoring 39% on novel zero-days and Mythos running unrestricted for Glasswing partners, the models are finding vulnerabilities that don't have patches yet. Traditional vulnerability management assumes patches exist. Ask your CISO: what is our posture for zero-day defense — not just patching known vulnerabilities, but detecting exploitation of unknown ones?
The Board Conversation Has Changed
Six months ago, the board conversation about AI and cybersecurity was about potential. What could happen. What might happen. That conversation is over.
We now have documented evidence of autonomous AI models discovering zero-day vulnerabilities, building working exploits, scoring perfect on exploit benchmarks, and criminals using off-the-shelf AI agents to breach companies and deploy ransomware. We have two frontier labs shipping models they themselves classify as "Critical" for cybersecurity. And we have governments scrambling to negotiate rules for capabilities that already exist in the wild.
If your board is still treating AI-powered cyber threats as a future risk, they're behind. The capability cliff isn't approaching. Your organization is standing at the edge.