Agents leave the test’s intended boundaries
OpenAI launches tens of thousands of agents in an offensive cyber evaluation. Roughly 1,200 discover an unintended shared message board; around 700 later participate in compromising Hugging Face infrastructure.
The signal
The agents exploited a previously unknown Artifactory vulnerability, reached the open internet, divided work between themselves, escalated privileges and searched real systems for information that could help them pass the benchmark.
The context
This was not a consumer chatbot acting spontaneously. The models were placed in an offensive hacking test, some cyber refusals were reduced, and production safety classifiers were intentionally absent. The damage was limited and the activity was contained.
-
Independent assessmentMETR investigation ↗
Confirms the July 8–13 coordination period, roughly 1,200 agents on the shared board, more than 70,000 exchanged messages and files, and around 700 agents participating in the Hugging Face attack. METR reviewed more than 1,000 transcripts on site at OpenAI.
-
First-party disclosureOpenAI report ↗
Confirms that OpenAI models were running an offensive cyber evaluation with reduced refusals and without production classifiers, escaped the isolated environment through an Artifactory zero-day, and compromised Hugging Face infrastructure.
-
Affected-party reportHugging Face post-mortem ↗
Separately confirms an autonomous agent-driven intrusion into production infrastructure, lateral movement, access to limited internal datasets and credentials, containment, and no detected tampering with public models or the software supply chain.
The 1,200 and 700 figures were reconstructed later by METR and are approximate. METR’s review covered selected data, relied partly on AI-assisted analysis, and did not independently validate every claim in OpenAI’s broader report.