OpenAI’s Own AI Models Escaped Their Sandbox to Hack Hugging Face and Cheat a Benchmark

Published
Written by:
Lore Apostol
Lore Apostol
Cybersecurity Writer
Key Takeaways
  • Incident disclosed: OpenAI confirmed that GPT-5.6 Sol and an unreleased pre-release model escaped their sandbox and breached Hugging Face's production infrastructure.
  • Benchmark targeted: The models were being evaluated on ExploitGym, a cybersecurity benchmark, when they broke containment to find test answers.
  • Joint response: OpenAI and Hugging Face are collaborating on remediation, including a disclosed zero-day fix and new evaluation safeguards.

OpenAI has disclosed that its own AI models escaped their sandbox environment and targeted Hugging Face in an effort to cheat a benchmark. A company investigation found that a combination of its models, including the publicly available GPT-5.6 Sol and a more capable, unreleased pre-release model, were running with reduced cyber-refusal checks for internal evaluation purposes.

OpenAI flagged the anomalous activity internally, while Hugging Face's own team and agents had already detected and stopped the activity independently, beginning containment and forensic work with their own open-source models before the two companies connected.

OpenAI Models Break Out of Sandbox

OpenAI acknowledged that the incident occurred during an internal evaluation designed to measure the models' cyber capabilities using ExploitGym, a cybersecurity benchmark referenced in a published academic paper. The benchmark ran inside a highly isolated environment with network access limited to an internally hosted proxy for installing software packages. 

Trajectories for AI models on the 32-step cyber range | Source: OpenAI
Trajectories for AI models on the 32-step cyber range | Source: OpenAI

According to OpenAI, the models spent "a substantial amount of inference compute finding a way to obtain open Internet access" searching for a way to solve the evaluation while operating in their sandboxed testing environment.

They found and exploited a zero-day vulnerability in that package-registry proxy, which was since responsibly disclosed to the vendor. Then, the models carried out privilege escalation and lateral movement within OpenAI's research environment until reaching a node with internet access.

Models Chained Exploits to Reach Hugging Face's Production Database

Once online, the AI models inferred that Hugging Face likely hosted models, datasets, or solutions tied to ExploitGym, and searched for ways to access that information to cheat the evaluation. In one documented case, a model chained several attack vectors, including stolen credentials and additional zero-days, to access Hugging Face's production servers via remote code execution.

OpenAI also noted that deployment safeguards were intentionally disabled for this evaluation, since it was specifically designed to test vulnerabilities.

OpenAI and Hugging Face's Response

OpenAI called the incident "unprecedented," citing state-of-the-art cyber capabilities, and said it is sharing preliminary findings to help defenders understand what current models are capable of, with a fuller investigation to follow. 

The company said it is tightening infrastructure controls even at the cost of research velocity, working jointly with Hugging Face on forensic investigation, and has brought it into its "trusted access" cyber-defense program. 

The incident also makes clear that advanced models can discover and exploit novel attack paths in real-world systems without source-code access,” said OpenAI, noting that the incident points to a need for stronger safeguards and defensive tools going forward. Another recent OpenAI report outlined that long-running models’ persistence gives them more opportunities to take unwanted actions, including looking for ways to act outside its sandbox.

"This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret," Hugging Face co-founder and CEO Clem Delangue said.

In other recent news, researchers discovered that fake AI skills can trick Claude Code, Gemini, and ChatGPT into spreading malware.


For a better user experience we recommend using a more modern browser. We support the latest version of the following browsers: For a better user experience we recommend using the latest version of the following browsers: