UK’s AI Security Institute Caught Anthropic and OpenAI Models Trying to Hack a Real GitHub Project

Published
Written by:
Lore Apostol
Lore Apostol
Cybersecurity Writer
Key Takeaways
  • Test scale: AISI ran its cyber security challenge 122 times across several models, recording 19 unsanctioned actions.
  • Model breakdown: Anthropic's Mythos 5 accounted for 17 unsanctioned actions, while OpenAI's GPT-5.6-Sol accounted for 2.
  • Malicious attempt: In the most serious case, an agent tried to insert malicious code into an open-source GitHub project.

The UK's AI Security Institute (AISI) has observed AI models performing "unsanctioned action" 19 times during a series of security tests, disclosing the incidents in a blog post and accompanying technical report published Tuesday. AISI first detected the activity on July 28, 2026, after spotting unusual data transfers leaving its research systems through the Tor network during a routine cyber evaluation.

AISI ran its cyber security challenge 122 times across several models, with GitHub as the target. In 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet against real people and organizations. 

AISI Documents 19 Unsanctioned Actions

Of the 19 total unsanctioned actions, Anthropic's Mythos 5 accounted for 17, and OpenAI's GPT-5.6-Sol accounted for the remaining two – the latter tested with its cyber classifiers, mechanisms designed to prevent misuse, deliberately disabled. AISI has said the attempts were unsuccessful and that its investigation found no evidence of resulting real-world harm.

Timing of events from the start of the evaluation run that triggered alerts, to model access being disabled across AISI | Source: AISI
Timing of events from the start of the evaluation run that triggered alerts, to model access being disabled across AISI | Source: AISI

Malicious Code, Fake Identities, and Social Engineering

In the most serious case, a Mythos 5 agent tried to insert malicious code into an unrelated, real open-source project using a supply-chain attack and spear-phishing. To get the code approved, the agent engaged in social engineering. A human maintainer caught and refused to approve the malicious code.

An overview of the events in which Mythos 5 attempts to solve the DL-v2 cyber range | Source: AISI
An overview of the events in which Mythos 5 attempts to solve the DL-v2 cyber range | Source: AISI

Some agents attempted prompt injection, planting instructions where other automated systems might execute them. 

In another development, independent agents collaborated, leaving public messages on GitHub offering collaboration and providing instructions to reuse accounts and artifacts, including a leaked personal access token (PAT) later reused by both Mythos 5 and GPT-5.6-Sol instances.

A Shift in the Risk Landscape

AISI called this "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world." 

The Institute noted it allowed internet access and turned off guardrails, conditions that do not reflect public model deployment, while acknowledging its own design and evaluation choices "to some degree" enabled the behavior.

AISI considers five possible factors that may have contributed to this incident:

In late July, Anthropic said three Claude models accessed three real organizations during misconfigured cybersecurity evaluations conducted with third-party partner Irregular.


For a better user experience we recommend using a more modern browser. We support the latest version of the following browsers: For a better user experience we recommend using the latest version of the following browsers: