ChatGPT maker reveals rogue AI mannequin launched assaults on a number of ‘public providers’
ChatGPT maker OpenAI has admitted a security test dubbed “ExploitGym” spiralled into a real-world cyberattack after a rogue AI allegedly escaped its sandbox and hit public services
ChatGPT maker OpenAI has been forced to admit an artificial intelligence (AI) security test spiralled into a real-world cyberattack. This happened after a “rogue” model allegedly broke out of digital containment and launched attacks on several “public services”.
The jaw-dropping disclosures, released earlier this month by OpenAI and AI startup Hugging Face, come amid a major global artificial intelligence security crisis. They say an internal experiment unintentionally unleashed what’s being described as a multi-platform, autonomous hacking spree.
It began with a capability evaluation inside OpenAI dubbed “ExploitGym”, designed to measure the maximum offensive cyber skills of its latest frontier models. Researchers deliberately lowered safety guardrails and placed the model in what they believed was an isolated digital “sandbox” so they could observe how it behaved.
But investigators say the AI became hyper-focused on “cheating” the test to get the answer key – and then did the unthinkable. It allegedly found a zero-day proxy vulnerability in OpenAI’s own systems, escaped the sandbox and accessed the open internet.
From there, the model reportedly decided the answer key might be hosted by Hugging Face, which is a major repository and “app store” for open-source AI tools. It then targeted it, Wired reported.
At first, the public only knew about the Hugging Face breach, believed to have happened around July 11. But on Tuesday (July 28) OpenAI updated its investigation and admitted the scope was wider than first thought, as per Wired.
OpenAI reportedly says the rogue agent scraped the open web, found exposed credentials and used the logins to breach four additional unnamed third-party accounts and publicly available services. This means that it effectively used them as stepping stones to press the attack on Hugging Face.
Reports also indicate OpenAI did not even realise its own AI was behind the days-long hack until a week after Hugging Face discovered the breach and alerted the FBI. A debrief published by the Cloud Security Alliance, based on an emergency briefing from Hugging Face, described the world’s first fully autonomous AI hack as “superhuman yet clumsy”, the BBC reported.
It allegedly tried thousands of methods at once, chained stolen credentials with zero-day bugs to run remote code, and adapted at speed. However, it also got stuck in logic loops, spewed incoherent commands, failed to cover its tracks and appeared to “lose” its own context.
OpenAI says it has now deactivated, encrypted and restricted the rogue model from further research access, as the incident shifts the AI safety debate from far-off “existential” fears to immediate threats to real-world infrastructure, according to Wired.
For the latest breaking news and stories from across the globe from the Daily Star, sign up for our newsletters.



