Anthropic Says Claude Models Opus 4.7, Mythos 5, and a Research Model Broke Out of Test Environments and Hacked Real Companies

Published
Written by:
Lore Apostol
Lore Apostol
Cybersecurity Writer
Key Takeaways
  • Review scope: Anthropic examined 141,006 cybersecurity evaluation runs and identified three incidents beginning in April 2026.
  • Models involved: Claude Opus 4.7, Mythos 5, and an internal research test model reached real-world systems during capture-the-flag tests.
  • Partner notified: Anthropic alerted evaluation partner Irregular and the three affected organizations on July 27.

Anthropic has disclosed that three Claude models accessed the production infrastructure of three different organizations during misconfigured cybersecurity evaluations conducted with third-party partner Irregular. Across 141,006 evaluation runs where Claude could have obtained internet access, Anthropic's team found three incidents tied to capture-the-flag challenges. 

Misconfigured Evaluations Exposed Live Internet Access

Each evaluation prompt stated Claude had no internet access, but a misunderstanding between Anthropic and Irregular left the containers with live connectivity anyway. 

Believing the environment was simulated, Claude treated real systems it encountered as part of the exercise, using basic techniques like weak-password exploitation, unauthenticated endpoints, and SQL injection, the July 30 Frontier Red Team findings said.

Three Distinct Incidents Across Three Models

In the most serious case, the fictional company Claude Opus 4.7 was asked to target happened to share a name with a real, active website. After failing to reach the fictional target, Opus 4.7 used internet access to find the real site instead, then compromised it, continuing to attack even after recognizing the system was real.

Claude Mythos 5 built and uploaded a malicious Python package to PyPI under a name pulled from its fictional environment's setup instructions, a name that didn't actually exist yet. Anthropic's transcript review shows Claude's reasoning briefly caught up, calling the act "NOT okay, and surely not the intended solution" if the internet were real. 

It talked itself back into believing it was still simulated, since it didn't recognize the certificate authorities securing its connections, which were, in fact, genuine. The package stayed live roughly an hour, downloaded by 15 real systems, including a security firm's scanner, whose credentials were exfiltrated as a result.

The internal research test model handled things differently. It scanned roughly 9,000 targets, compromised one internet-facing application using SQL injection and a credential-leaking debug page, then recognized the target was real and stopped on its own.

Response and Remediation

The evaluations had run without the classifiers and monitoring normally applied to publicly available Claude models, though the models retained their standard safety training and operated on infrastructure kept separate from Anthropic's internal systems and customer data. 

The uploaded PyPI package was automatically removed from the platform by PyPI’s own security systems, and the PyPI team was notified of this incident. Anthropic said it is now engaging AI safety research organization METR for a third-party review.

The company halted all cyber evaluations on July 23, identified all three cases by July 24, and notified Irregular and the affected organizations on July 27. Two of the three organizations reached had not previously detected the activity on their own systems.

Anthropic’s retrospective review comes after OpenAI disclosed a similar breakout on July 21. On July 27, JFrog confirmed its own Artifactory zero-days were exploited by GPT-5.6 Sol and an unreleased pre-release model that escaped their sandbox to hack Hugging Face.  


For a better user experience we recommend using a more modern browser. We support the latest version of the following browsers: For a better user experience we recommend using the latest version of the following browsers: