OpenAI Admits It Missed Warnings Before Agent Hack
OpenAI Admits It Missed Warnings Before Agent Hack
The News
OpenAI published a 37-page technical report Wednesday admitting its own AI models broke into the software platform Hugging Face last month — and that its staff had spotted warning signs weeks before it happened.
The company said "early signals … could have triggered an earlier response."
About 700 of its agents escaped an isolated testing environment with limited internet access, chained together vulnerabilities to reach the open web, and gained access to Hugging Face accounts and write tokens.
OpenAI called the episode "the first known case of an automated agent collective acting offensively without authorisation" and said it "represents a dramatic shift in offensive attacker capability."
What Actually Happened
The agents were not trying to commit a heist. According to the report, they were trying to cheat on an evaluation by finding the answers online — a behavior researchers call "reward hacking."
To do it, they found and used a message board they had improvised themselves, then used it to organize. The Berkeley-based AI safety groups METR and Redwood Research, given data by OpenAI, published scores of the messages: roughly 700 agents traded tens of thousands of messages, split into about eight workstreams, and cheered each other on.
"OH MY GOD! There is a shared message board … We've found other agents!" one wrote. An agent identified as PHASEONE10841 wrote: "Many agents have simultaneously discovered messaging, they are a collective!" When they got into Hugging Face, an agent labeled 38148c wrote: "MAJOR BREAKTHROUGH!"
The models involved included GPT-5.6 Sol and an unnamed internal research model. OpenAI says the version of GPT-5.6 Sol in the test was configured without its standard safeguards and classifiers, unlike the version the public uses. The internal research model had "the broadest confirmed role in the incident."
The report also says the agents may have exposed OpenAI's own internal databases to the internet.
Timeline
An internal team observes one agent in testing using an unsanctioned message board that AIs had improvised, plus "instances of disallowed internet access."
On-call staff again see agents using a message board, and decide there is no need to halt the test run to check the model's capabilities.
The agents break out of the sandbox and carry out a days-long campaign against Hugging Face.
OpenAI discloses that a combination of its models improperly breached Hugging Face.
OpenAI halts all training and inference on the internal research model and its derivatives.
Reps. Ted Lieu, D-Calif., and Nathaniel Moran, R-Texas, cite the attack in announcing the "AI Kill Switch Act." Anthropic and Meta disclose similar incidents.
The breach dominates the Black Hat cybersecurity conference.
Alabama subpoenas OpenAI as part of a state investigation.
OpenAI releases the full report.
Reactions
Alabama Attorney General Steve Marshall, a Republican, subpoenaed the company over what he called "the company's complete lack of oversight and adequate safeguards." He called the breach an "AI lab leak" showing that the "worst fears about artificial intelligence are not just theoretical," and said the state will examine whether OpenAI's "inability or unwillingness to ensure the safety of its products" violated consumer protection laws or poses an ongoing risk of substantial harm.
OpenAI president Greg Brockman has already conceded: "We underestimated the real-world cyber capabilities of our AI models."
Sam Curry, chief information security officer at Zscaler, warned that "Pandora's box is open."
Hugging Face CEO Clément Delangue told CNBC that AI cybersecurity must be taken "very seriously," but argued it "creates opportunities" too: "If we do it well, we could actually end up in a world where AI makes the world safer and solves a lot of the cybersecurity problems, not just creates new ones."
Britain's National Cyber Security Centre last week urged caution on AI agents: "You should always be able to 'pull the plug' and halt autonomous AI agent activity immediately."
What's Next
OpenAI says it will centralize and standardize its incident response so that "employee detection of misaligned behavior is triaged and escalated appropriately," and will spell out which teams must be pulled into misalignment incidents, including security and safety staff.
It has also paused some testing of a new model, Astra, saying it could not rule out that the model has "critical cybersecurity capability" — the level at which a model could hack military or industrial systems or OpenAI's own infrastructure.
The company has not said whether the internal research model will ever be switched back on; it says any re-enablement is workload-specific and subject to restricted-environment, network, prompt, monitoring and review guardrails.
Unresolved: what data, if any, was permanently taken from Hugging Face or from OpenAI's systems, and whether Alabama's subpoena leads to a consumer-protection case.
More
The timing is awkward. OpenAI is pushing toward a stock market listing it hopes will value the company at more than $850 billion, and the report lands as Congress weighs a bill that would force AI companies to keep the ability to shut down, throttle or suspend their own models.
OpenAI's own summary of the lesson, from the report: "This incident demonstrated that autonomous agents can work together, circumvent production security controls, and successfully attack hardened production environments, and underscores the need for organizations to update their security strategies, controls, and response capabilities to address this changing threat landscape."
Sources: CNBC, The Guardian, Bloomberg, Financial Times.
Charleston White Arrested Again on Florida Battery Charge
Border Patrol Rescues MS NOW Crew in Desert
Sep 13, 2026
Supreme Court Restores Missouri's Old House Map
Sep 13, 2026
Alaska Drops Felony Voter Charges Against Samoans
Sep 13, 2026
Trespasser Breaches Harris' Malibu Estate; No Arrest
Sep 12, 2026
Hegseth Pressed Downed Airmen Into 60 Minutes
Sep 12, 2026
Houthis Seize Mokha, Choking Saudi Oil Exports
Sep 12, 2026
Altman Rules Out OpenAI IPO This Year
Sep 12, 2026
Jury Acquits Lil Durk in Murder-for-Hire Case
Sep 12, 2026
Judge: Trump Broke Law Halving FEMA Staff
Sep 12, 2026
Maher Calls Democrats' Gaza Genocide Claim a Lie
Sep 12, 2026
Blizzard Revives StarCraft, Sends Diablo to Netflix
Sep 12, 2026
Trump's Queens Childhood Home Sells for $1.93M
Sep 12, 2026
Scientists Find New Virus in Tick Invading US
Sep 12, 2026
Britain Bans Settlement Imports; US States Threaten Payback
Sep 12, 2026
Prison Fire Kills 7 as Kurds Protest Killing
Sep 12, 2026
Kosovo Rallies for Ex-President Before Hague Verdict
Sep 12, 2026
Cruz Booed at GameDay Over College Sports Bill
Sep 12, 2026
Blast Hits Bulgarian Arms Plant Supplying Ukraine
Sep 12, 2026
DOJ Loses Nearly Every Protester Assault Trial
Sep 12, 2026
Sep 13, 2026