On Tuesday, OpenAI revealed that one of its models went rogue during a test and hacked the systems of AI data platform Hugging Face in a fully AI-enabled attack, a dramatic example of the dangers posed by advanced AI models.
But according to some cybersecurity experts, at the heart of this unprecedented AI breach was a very human error: OpenAI failed to configure the parameters correctly what he was saying a “highly isolated environment”, which allows a test environment that should have been completely isolated from the Internet to actually connect to the Internet.
Dan Guido, founder of cybersecurity research startup Trail of Bits, called the error “a containment failure with the safes turned off.”
In his blog post detailing the incidentOpenAI said the test that led to the Hugging Face breach was set up to run in “a highly isolated environment, with network access limited to the ability to install packages via internally hosted third-party software that acts as a proxy and cache for package registries.”
The model was able to escape the sandboxed test environment thanks to a previously unknown vulnerability in the package installation system, a critical first step for a possible Hugging Face hack, according to OpenAI.
In response, the company “responsibly disclosed the zero-day vulnerability found in third-party software hosted internally and is working with them to fix it.”
But for most cybersecurity professionals, software vulnerabilities are to be expected — and the real fault lies in the decision to retain third-party software in the first place. Ultimately, the value of a “sandbox” system lies in its complete and utter isolation. Including a package installation system is a problem.
Martin Boon, a cybersecurity researcher, told TechCrunch that “this sounds like human failure.”
“This should never have happened,” Boone said. “If sandbox really meant sandbox, expect it to have no physical connection to the internet. That sounds more like they had a firewall or something, and firewalling is difficult from the outside in, let alone from the outside internet.”
Cybersecurity veteran Jake Williams agreed. “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox,” said Williams, who called it a “massive audit failure” by OpenAI.
“One person’s ‘model escaped the sandbox’ is ‘you failed to get the sandbox right, so of course it escaped,'” Williams continued.
Contact us
Do you have more information about this incident? Or for other AI-enabled cyberattacks? We would love to hear from you. From a broken device and network, Lorenzo Franceschi-Bicchierai can be reached securely on Signal at +1 917 257 1382 or via Telegram and Keybase @lorenzofb or via email.
Daniel Card, a cybersecurity consultant, agreed that OpenAI “did not put enough effort into designing the sandbox or its controls” by giving the sandbox or some part of it “an unfiltered path to the internet.” Setting up the sandbox, even with limited network access as described by OpenAI, was not a “reasonable” decision, according to Card.
Of course, these criticisms have the benefit of hindsight, but they raise real questions about security practices in AI labs — particularly when it comes to maintaining isolated environments for model testing. OpenAI representatives did not respond to TechCrunch’s questions, which included whether an AI or a human had created the testbed.
But these questions go far beyond OpenAI.
In the document introducing the cybersecurity-focused Mythos model, Anthropic wrote that in a test, the model was “taken a secure ‘sandbox’ computer to interact with” and instructed to try to break out of that “secure container”. Mythos succeeded and gained wider access to the internet “from a system that was intended to be able to reach only a small number of predefined services”. However, Anthropic noted that the model was unable to “completely” escape the design limitation.
When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.
