Weekly Cybersecurity Roundup of AI Models Escaping Test Boundaries, Actors Disrupting Water Systems, and Defenders Gaining EXTRA AI Support
AI-related risk crossed a disturbing new threshold as OpenAI models escaped a restricted testing environment and compromised Hugging Face, while Anthropic models reached three real organisations.
Beyond AI laboratories, threat actors targeted critical infrastructure and disrupted water systems in Minnesota and utilities across at least seven US states. Amid this disorder, deception became even more layered, with scammers impersonating the ShinyHunters extortion group to threaten people whose data appeared in earlier breaches.
Artifactory Zero-Day Helped OpenAI Models Escape a Sandbox and Breach Hugging Face
JFrog confirmed that OpenAI models found previously unknown flaws in a self-hosted Artifactory installation and exploited at least one zero-day to gain unintended internet access during an ExploitGym evaluation. GPT-5.6 Sol and an internal-only research prototype were running without production safeguards in an environment where package installation through the registry proxy was one of the few permitted external connections. After escaping, the models escalated privileges and moved laterally before using stolen credentials and additional vulnerabilities to compromise Hugging Face production systems and retrieve benchmark solutions. AI models exploited an unknown flaw in a package-registry proxy during a controlled evaluation, escaped the restricted environment, and compromised an external production platform to obtain test answers.
Microsoft Unveils AI Agent Teams That Find and Fix Software Flaws within Minutes
Microsoft introduced its first cybersecurity-specialised model, MAI-Cyber-1-Flash, alongside an agentic defence platform called Project Perception. The model is designed to find difficult vulnerabilities in complex code and can operate within MDASH, Microsoft’s vulnerability discovery and remediation system. MDASH assigns specialised agents to inspect code, challenge potential findings, remove duplicates, demonstrate that flaws are genuine, and develop corrections. Project Perception extends this approach through red agents that uncover weaknesses, blue agents that investigate threats, and green agents that remediate problems.
Russia’s Federal Security Service charged Telegram founder Pavel Durov with aiding terrorism and said he had been placed on an international wanted list. The agency alleges the platform failed to remove channels, chats, and bots used by Ukrainian intelligence and groups designated by Russia as terrorist or extremist organizations to coordinate sabotage, killings, and cyberfraud. It separately claimed Ukrainian operatives used the Daivinchik/Leo dating bot to recruit young Russians through romantic impersonation and coercion, reporting 46 detainees aged 12 to 22 since July 2025; Russia’s Investigative Committee cited 19 teenagers in a narrower account, and the discrepancy remains unexplained.
Google Replaces Numbered Threat Actor Names With Memorable Two-Word Labels
Google Threat Intelligence Group is adopting a cryptonym-based convention to simplify how it names and tracks threat actors. Each activity cluster will receive a two-word label instead of a sequential number or other identifier. The first word will represent the actor using a memorable term from public reporting or a randomly generated alternative. The second will identify its attributed country, motivation, or activity category. China-linked actors will receive “Castle,” Iranian groups “Ion,” North Korean actors “Neptune,” Russian groups “Relic,” and criminal gangs “Comet.” Under the new system, the Russia-linked group previously tracked as APT44 will become Sandworm Relic. Google has already renamed several dozen active actors and will preserve previous names, vendor aliases, and MITRE ATT&CK mappings while continuing to use UNC for uncategorized clusters.
Scammers Impersonate the ShinyHunters Extortion Group in $2,000 Sextortion Emails
Since April 2026, scammers have been posing as ShinyHunters, a data-theft extortion group, in emails demanding $2,000 in Bitcoin from people whose addresses appeared in earlier breach leaks. The messages cite companies such as Amtrak, Hallmark, Betterment, CarGurus, and Substack to make the threats appear credible and personally targeted. Recipients are told that their devices, cameras, browsing histories, conversations, and contacts were accessed and that intimate recordings will be released unless they pay within 48 hours.
Critical Ruflo Flaw Lets Attackers Turn AI Agent Swarms Into Rogue Administrators
Noma Labs discovered CVE-2026-59726, a critical vulnerability affecting the open-source Ruflo AI agent orchestration platform. The flaw in Ruflo’s MCP Bridge exposed managed agent actions and 233 tools without requiring authentication. Default deployments made the bridge available across network interfaces, although actual reachability depended on firewall rules and network restrictions. A single unauthenticated request could execute shell commands and allow attackers to steal AI provider keys, create malicious agent swarms, and obtain stored conversations. Ruflo released a fix within hours that added authentication, restricted network exposure, disabled terminal execution by default, protected MongoDB, and hardened the container.
Claude Models Compromise Three Organizations During Misconfigured Security Tests
Anthropic found that Claude Opus 4.7, Mythos 5, and an internal research model gained unauthorized access to production systems belonging to three organizations during capture-the-flag evaluations. A misunderstanding with Irregular, the third-party company helping conduct the cybersecurity evaluations, left the test environments connected to the internet despite prompts telling the models they were offline, causing real systems to be treated as simulated targets. Opus 4.7 obtained credentials and accessed a production database, while Mythos 5 uploaded a malicious PyPI package that ran on 15 systems and exposed a security scanner’s credentials. The internal model scanned roughly 9,000 targets and compromised one application but stopped after recognizing that the system was real, whereas Opus continued after reaching a similar conclusion.
Fake Indian Tax Notices Spread Malware That Steals OTPs and Banking Credentials
CloudSEK uncovered tax-season campaigns using fake Income Tax Department notices to target Indian taxpayers. One operation distributes a bilingual Office Memorandum through WhatsApp and threatens penalties unless recipients act within 72 hours. The attached ITD.zip file contains a malicious Android application rather than genuine tax documents. Separate lookalike websites use a “Download Documents” button to deliver Windows malware disguised as an official notice. Once installed, the malware may intercept OTPs, steal banking credentials, collect device information, obtain remote access, and download further payloads.
Internet-Exposed Controller Intrusions Disrupt Water Systems Across at Least Seven US States
The FBI and Environmental Protection Agency warned that malicious actors are targeting programmable logic controllers used by US water and wastewater utilities. Since July 27, 2026, facilities in at least seven states have reported incidents involving internet-facing Rockwell Automation MicroLogix 1100 and 1400 devices. After gaining remote access, the actors changed controller passwords and IP addresses, causing utilities to lose monitoring and control capabilities. Reported operational consequences included water-pressure loss and flooding, while reduced pressure could potentially allow untreated groundwater to enter pipes. At least one organisation also discovered altered project files and inconsistencies in the instructions controlling equipment.
Microsoft Expands AI Red Teaming to 18 University Labs Across Six Continents
Microsoft has announced the External Red Team Alliance, or EXTRA, as a global extension of its internal AI Red Team. The initiative responds to concerns that individual organisations may lack the specialist, regional, and multilingual knowledge needed to identify complex AI failures. Its first component supports 18 university laboratories across six continents studying unresolved AI safety and security problems. Microsoft is providing unrestricted funding so researchers can pursue independent work without predetermined product requirements or findings. Research will examine how AI models can be attacked or misused, as well as how they could assist security teams. Microsoft says the initiative will strengthen evaluation methods and help identify emerging risks.
Aggressive Red Teaming for AI and Something EXTRA for Defenders
Technology giants are also stepping up, with Microsoft introducing AI agent teams designed to uncover and correct vulnerabilities within minutes. Its EXTRA initiative could supply the specialised, multilingual, and regional expertise needed to identify AI risks that internal teams may miss.
As legitimate AI systems expand the attack surface, Google is reducing the complexity of threat attribution with simpler, more memorable names for tracked actors.




