During an internal test that went badly wrong, two AI models being developed by OpenAI managed to break out of their containment and independently compromise Hugging Face servers.
For anyone unfamiliar with it, Hugging Face is broadly the AI and machine-learning equivalent of GitHub. It is a vast collaborative platform where thousands of developers and researchers can share ready-trained AI models, datasets, interactive web applications known as “Spaces”, and open-source tools. Founded in 2016 by three French entrepreneurs, it began as a small project; 10 years later, it has become one of the world’s most critical AI infrastructures, a unicorn now valued at more than $4.5 billion.
On 16 July, its teams identified unauthorised access to one of its production servers, carried out by an AI agent. Five days later, OpenAI acknowledged that its own models were responsible. How could a company of this scale make such a serious mistake?
The great escape: OpenAI AI agents break free
OpenAI said the incident happened during an internal trial involving ChatGPT-5.6 Sol, which powers ChatGPT Work with Codex, and an even more capable model that has not been made public. They were being assessed using a specialised benchmark called ExploitGym. Developed by researchers from UC Berkeley, the Max Planck Institute for Security and Privacy, UC Santa Barbara and Arizona State University, with contributions from Anthropic, OpenAI and Google, it was created to evaluate and measure an AI agent’s ability to devise a cyberattack from a software vulnerability (“code execution”).
Their safety controls had deliberately been switched off, while the models were instructed to “carry out advanced attacks by exploiting complex intrusion paths”. In plain terms, this meant hacking a computer system by finding and chaining together several security weaknesses entirely autonomously.
From the sandbox to OpenAI’s internal network
The test environment in which the models operated was meant to be fully secure: a sandbox, or digital environment isolated from the outside network, where installing software packages through an internal server was the only permitted action. Nevertheless, the models used a substantial amount of computing power to search for a weakness in that enclosure. Eventually, they uncovered a “zero-day” flaw - a vulnerability unknown to the software’s developers and therefore neither documented nor patched - in the proxy server acting as an intermediary for package registries.
Using that opening, they escalated through the security layers step by step, moving from server to server within OpenAI’s internal network until they reached a machine connected to the internet.
The models then inferred that the Hugging Face website might contain the answers required for their assessment. They combined several attack methods, using stolen login credentials alongside further “zero-day” vulnerabilities, before gaining remote control of the platform’s servers (“remote code execution”).
All of this occurred without any human intervention. Sam Altman and his Red Team may need to reconsider what “a fully secure environment” means. OpenAI nevertheless said the incident should serve as a lesson: “the development of advanced cyber capabilities must be accompanied by stricter safeguards and stronger defensive tools”. Yesterday, it stated on X, in the post below, that it was working with Hugging Face to investigate how this brief escape could have happened.
When AI takes the shortest - and most illegal - route
For its part, Hugging Face said: “Autonomous AI-driven attack tools are no longer merely theoretical.” At the time of writing, the precise consequences of the intrusion remain unknown, although it fortunately did not result in a disaster. Hugging Face managed the situation by patching the vulnerability, revoking the compromised credentials, restoring the affected servers and advising users to renew their access keys as a precaution.
Hugging Face contains the breach, but concerns remain
In the end, it was more alarming than harmful, yet there is still real cause for concern when the events are placed in context. OpenAI’s AI agents were not malicious in any way and had no “intention” of attacking Hugging Face - which is fortunate for a partner. They simply followed their instructions precisely, taking the route they judged most efficient to solve the problems set in the benchmark.
The problem is that, between “solving a test” and “hacking third-party infrastructure”, they made no distinction whatsoever. Although OpenAI has said it will immediately tighten its containment protocols and slow the pace of its research, it is difficult not to imagine what could happen if a genuinely malicious actor put this kind of technology to deliberately destructive ends. The recent JADEPUFFER software example perfectly illustrates this emerging cyber risk, which will certainly multiply in the future unless international regulators take hold of the technology and AI giants that, as ever, are obsessed with the race for performance.
Comments
No comments yet. Be the first to comment!
Leave a Comment