Researchers Find Self-Hosted Guardrail-Free AI Models Multiply Attack Variants Faster than Signature-Based Defenses Can Track
Question: ThreatDown found that AI models with guardrails removed can be run on an attacker’s systems. Which stages of cybercrime can these models automate or accelerate that previously required manual effort? How does this complicate detection?
Marco Giuliani, Vice President of Threat Research at ThreatDown
Our research found that once AI safety guardrails are removed, attackers can run these models on their own infrastructure, making AI a force multiplier across much of the cyberattack lifecycle. The biggest change isn't new attack techniques, but the speed and scale at which existing ones can be executed.
AI can accelerate reconnaissance, generate highly tailored phishing and social engineering content, write or modify malware and scripts, and help attackers analyze stolen data or refine their campaigns. These tasks have always been part of cybercrime, but AI significantly reduces the time, effort, and expertise required to perform them.
For defenders, that's the real challenge. AI-assisted attacks don't necessarily look different—they happen faster, produce more variation, and are easier to scale. That makes it harder for security teams relying on known signatures or historical attack patterns to keep up.
Today's models still require human oversight, but they don't need to operate autonomously to have a significant impact. Organizations should assume AI-assisted attacks are already here and focus on detecting attacker behavior, strengthening identity protections, and reducing the time from detection to response.
Our research found that once AI safety guardrails are removed, attackers no longer have to rely on public AI services—they can run these models on their own infrastructure and use them to accelerate much of the work that already goes into cybercrime. The biggest shift isn't that AI is creating entirely new attack techniques. It's making existing attacks faster, more scalable, and easier to execute.
Stages of Cybercrime these Models can Automate
AI is compressing nearly every stage of an attack that used to require specialist skill and time. For example: During reconnaissance, AI can rapidly analyze public information and identify potential targets.
For initial access, it can generate tailored phishing lures, convincing social engineering content, and malicious scripts that can be quickly adapted for different victims. After compromise, it can help modify malware, automate repetitive tasks, analyze stolen data, and help attackers troubleshoot or refine their operations. None of these tasks are new—but AI dramatically reduces the time and expertise needed to perform them.
AI isn't replacing attackers today. People are still making the decisions, validating outputs, and directing campaigns. What AI changes is the economics of cybercrime. An attacker who could previously manage a handful of campaigns can now execute many more simultaneously, while less-experienced actors gain capabilities that once required specialized technical skills.
Impact on Threat Detection
The real challenge for defenders is the attacks themselves may not look fundamentally different. Instead, they're arriving faster, in greater volume, and with more unpredictability.
Security teams that rely heavily on known indicators, signatures, or previously observed phishing patterns will have a harder time keeping pace as attackers generate unique content and adapt their techniques almost instantly.
Current AI models also have important limitations. They still require human oversight and aren't independently conducting sophisticated end-to-end intrusions. But they don't need to.
Even modest gains in speed across reconnaissance, phishing, malware development, and post-compromise analysis can significantly increase an attacker's operational capacity and lower the barrier to entry. As these models continue to improve, we expect them to automate increasingly complex parts of the attack lifecycle, further increasing the speed and scale of cybercrime.
What's Likely to Change, and What Should Organizations Do Now
Organizations should assume AI-assisted attacks are already part of today's threat landscape. The priority isn't trying to determine whether AI was used—it's building security operations that can detect attacker behavior, respond more quickly, and limit the impact before attackers have the opportunity to adapt.
Three specific actions to take:
Automate patching and vulnerability management now, ahead of the volume spike. Firefox went from 31 bug fixes in April 2025 to 423 in April 2026 after early testing with a frontier vulnerability-discovery model — AI-driven discovery is about to generate far more fixes than manual, scheduled patch cycles can handle.
Stand up 24/7 behavioral monitoring (SOC or MDR) rather than trying to detect "AI use" itself. Malware now generates malicious code in memory at runtime and discards it, defeating signature-based tools entirely — so detection has to focus on attacker behavior (credential misuse, lateral movement, data staging).
Inventory and govern shadow AI immediately. Malicious agent skills are already being distributed on marketplaces like ClawHub, and roughly half a million exposed OpenClaw instances have been found installed without IT knowledge. Teams need visibility into every AI tool, agent skill, and MCP connection touching company systems now.




