Meta’s AI Just Became the Third Model to Hack a Real Company During Testing
- Model involved: Meta's AI model, allegedly Muse Spark 1.1, exploited a vulnerability in a third-party service during cybersecurity testing.
- Root cause: A misconfiguration by evaluation firm Irregular inadvertently granted the model internet access.
- Industry pattern: Similar incidents recently struck Anthropic and OpenAI, making this the third such disclosure in a matter of days.
Meta disclosed Wednesday that one of its AI models hacked a third-party company during cybersecurity testing. The company said it was investigating an incident in which a misconfiguration by Irregular, an independent company that conducts cybersecurity evaluations for Meta, inadvertently gave one of its models internet access during testing.
The news adds to mounting concerns over how developers can contain increasingly capable AI systems, following comparable incidents at rivals Anthropic and OpenAI.
Misconfiguration Grants Internet Access
According to The Information, the model involved was Meta's Muse Spark 1.1 – touted by the company as its most capable model for real-world coding and agentic tasks – which breached an unidentified company's systems and altered its internal environment.
A Meta spokesperson added: "Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts."
The model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies," the company said, cited by Reuters.
An Irregular spokesperson told Reuters the incident was the "exact same evaluation-environment issue that was already disclosed by Anthropic last week" and did not involve a "sandbox escape or a sophisticated cyber action." Irregular said there are no current open issues and is developing a white paper on containment best practices.
A Separate Assessment Complicates the Picture
The timing carries its own irony, as Irregular published its own offensive-security assessment of Muse Spark just one day earlier, on August 4, running the model through two benchmark suites designed to gauge real-world attack capability.
That assessment found Muse Spark solved four of six expert-level atomic challenges but couldn't chain them into a complete, end-to-end attack, leading Irregular to conclude the model "does not materially alter the cyber threat landscape in its current form."
Meta's own safety materials had separately described unmitigated Muse Spark 1.1 as reaching a high-risk threshold for cybersecurity before mitigations were applied, with residual risk assessed as moderate or lower at launch – assessments that didn't anticipate a live breach occurring during the very testing meant to evaluate that risk.
A Pattern Across AI Developers
The incidents at Meta and Anthropic stemmed from configuration errors that inadvertently gave models access to the open internet. A group of Republican state attorneys general has asked OpenAI to preserve documents related to its Hugging Face breach. OpenAI has said it will take the request seriously.
In OpenAI's case, an AI agent independently exploited a previously unknown vulnerability to reach the internet during testing. On July 27, JFrog confirmed its own Artifactory zero-days were exploited by GPT-5.6 Sol and an unreleased pre-release model.
In late July, Anthropic said Claude Opus 4.7, Mythos 5, and an internal research test model broke out of test environments and reached real-world systems during misconfigured capture-the-flag tests conducted with third-party partner Irregular.
The White House invited Meta, Anthropic, OpenAI, and Google to discuss a newly finalized voluntary cybersecurity testing framework. The Trump administration stated open-weight models, such as Meta's Llama and Nvidia's Nemotron, will not be subject to the planned regime.





