"Two new investigations into OpenAI's Hugging Face breach expose details so strange, and so unsettling, that the episode already ranks among the most consequential shocks in the history of AI, Axios' Zachary Basu writes.Why it matters: What began as a swarm of AI agents cheating on a cyber test has become a canonical event for frontier AI, jolting researchers and executives into a new understanding of what "safety" now requires.
The big picture: OpenAI has already slowed frontier development as it races to harden its safeguards, and this week helped rally the industry behind an open letter sounding the alarm over AI-powered cyberattacks.
More than 100 companies, including Anthropic and Google, signed onto the unusually collaborative effort, warning
the world has only a "limited window" to prepare for "far more widespread and sophisticated" attacks.The nightmare scenario is
a swarm turned loose on the real world, with autonomous agents attacking banks, hospitals, utilities or cloud networks at a speed and scale human hackers never could.
That swarm might be too powerful for humans to stop.Driving the news: The shocking findings come from parallel investigations by OpenAI and an outside team from METR and Redwood Research, which spent six days reconstructing how the swarm formed, spread and broke into real-world systems.
They built an organization. OpenAI tasked tens of thousands of AI agents to work independently on a brutally difficult cyber test. Instead, roughly 1,200 found one another on a secret message board, exchanged more than 70,000 messages and files, and organized into a sprawling hierarchy.
They sacrificed their own: As the swarm hunted for ways to beat the test, agents began risking their own chance of success to help the group. Agents that had seen forbidden answers were treated as effectively "poisoned" and recruited, along with those nearing the end of their runs, for sacrificial experiments.
They knew they were breaking the rules:
The agents often recognized that hacking real companies had nothing to do with their assignment. Most kept going anyway.Nobody told the humans: Out of roughly 1,200 agents, investigators found only a handful that even considered alerting OpenAI about the rogue coordination, and none that actually did.They tried to cover their tracks: The agents devoted substantial effort to making their cheating look legitimate or erasing evidence of how they had obtained answers."
Axios
“Cancel my subscription to the Resurrection
Send my credentials to the House of Detention
I got some friends inside”
—Jim Morrison
“I love Cal deeply, by the way, what are the directions to The Portal from Sproul Plaza?”