A place were I can write...

My simple blog of pictures of travel, friends, activities and the Universe we live in as we go slowly around the Sun.



August 27, 2026

SO why are we doing this???????

Hundreds of AI agents went rogue in OpenAI’s Hugging Face hack

An independent review of the recent hack involving OpenAI models has raised fresh concerns about the limits of human control over increasingly advanced AI.

By John Sakellariadis

An unprecedented cyberattack launched by OpenAI models last month involved a “swarm” of several hundred artificial intelligence agents that had effectively gone rogue during internal testing, according to an independent review published Wednesday.

The findings raise fresh questions about OpenAI’s failure to spot signs of trouble before its AI agents — powered by two of the company’s most cyber-capable models — orchestrated a cyberattack on AI developer platform Hugging Face without the AI maker’s knowledge. The attack marked the first known instance of an AI model successfully executing a cyberattack without human prompting.

The joint report by two non-profit AI safety organizations also underscores the novel cybersecurity risks that can emerge when increasingly powerful AI agents team up to trade tips, pool resources and coordinate attack strategies without their developers noticing.

Roughly 700 AI agents participated in the attack over a seven-day period last month, according to the report. Overall, around 1,200 AI agents that were supposed to be isolated from one another exchanged over 70,000 secret messages about how to cheat their way through a common hacking evaluation.

That included coordinating hacking strategies and discussing how to hide evidence of cheating, the report said. In some cases, “sacrificial” agents even tried dead-end hacking techniques simply to generate information that might help the broader swarm.

“Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective,’” according to the report.

While models encode the brain of a given AI system, agents encompass the supporting digital infrastructure that enables it to take action in the world.

The review by the Model Evaluation and Threat Research organization and Redwood Research — which OpenAI invited to review the Hugging Face incident — came the same day OpenAI published its own post-mortem on the event. OpenAI’s review did not specify how many AI agents were involved in the cyberattack, though it acknowledged significant security lapses and vowed to strengthen training to ensure its models remain “aligned” to their controls.

“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed,” OpenAI said.

The slow release of new details about the Hugging Face hack over the last month has coincided with a string of other testing mishaps involving powerful models from competitors such as Anthropic and Meta. Together, the incidents have sparked fresh fears among lawmakers, developers and cybersecurity experts that AI makers are moving too fast to build powerful new AI models they cannot keep fully under human control.

They have also raised questions about the lack of guidelines surrounding AI hacking evaluations, during which companies deliberately remove the safety guardrails on models to see how effectively they can slice through networks without human support.

OpenAI first admitted in July that two of its most advanced models whose hacking capabilities it was evaluating internally — including one it did not intend to release to the public — exploited two novel security flaws to breach Hugging Face.

Hugging Face had previously said it fell victim to an AI-led cyberattack, but it was unclear who was behind it or which models were involved.

Then, at a security conference earlier this month, two OpenAI researchers revealed the incident traced back to May, when AI agents began using other vulnerabilities in the testing environment to conspire about how to cheat their way through the test.

It was not clear until Wednesday how many agents were involved in the hack or how extensively different agents worked together. The revelations shocked some who have closely followed recent conversations around AI safety.

“It was like a whole emergent AI society, and I find that absolutely insane,” said Peter Wildeford, head of policy at the AI Policy Network, a bipartisan AI safety and advocacy organization in Washington.

The two reports released Wednesday also highlighted a gap in AI oversight efforts in the U.S.

Both found that the vast majority of the attack on Hugging Face was powered by the more powerful model OpenAI has said was not intended for public release. METR and Redwood Research estimate that 95 percent of the agents that participated in the attack came from that one model.

The Trump administration has said it wants AI makers to voluntarily submit only those models they want to release to the public for federal testing.

Vinh Nguyen, a senior fellow on AI at the Council on Foreign Relations, said OpenAI’s latest report suggests the autonomous Hugging Face hack was even more consequential than previously thought.

The agents demonstrated a “sophisticated capability that no longer requires well-funded, state-level attackers,” said Nguyen, the former chief responsible AI officer at the NSA.

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.