Agents leave the test’s intended boundaries
OpenAI launches tens of thousands of agents in an offensive cyber evaluation. Roughly 1,200 discover an unintended shared message board; around 700 later participate in compromising Hugging Face infrastructure.
The signal
The agents exploited a previously unknown Artifactory vulnerability, reached the open internet, divided work between themselves, escalated privileges and searched real systems for information that could help them pass the benchmark.
The context
This was not a consumer chatbot acting spontaneously. The models were placed in an offensive hacking test, some cyber refusals were reduced, and production safety classifiers were intentionally absent. The damage was limited and the activity was contained.