Artificial intelligence guardrails designed to prevent malicious hacking are increasingly obstructing the work of legitimate cybersecurity researchers, industry experts say. These restrictions, implemented by major AI developers like Anthropic and OpenAI, are intended to block harmful uses but are now creating unintended barriers for professionals tasked with identifying and fixing vulnerabilities.

For context: AI guardrails are built-in safety mechanisms that restrict certain prompts or outputs, such as requests to generate malicious code. While these measures aim to prevent abuse, they also limit the tools available to researchers who simulate cyberattacks to strengthen defenses.

Guardrails Create Roadblocks for Security Researchers

In June, the U.S. government imposed export controls on Anthropic’s AI models, Mythos and Fable, citing concerns over potential misuse. The restrictions were later lifted for Fable 5 in July, while Mythos 5 remains accessible only to vetted U.S. organizations. Anthropic and OpenAI both offer specialized programs—such as OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program—to grant approved researchers access to less restricted models.

However, these programs have faced criticism from offensive security professionals, whose work involves proactively probing systems for weaknesses. Mark Dowd, a veteran security researcher, told a cybersecurity podcast that AI companies’ arbitrary decisions on what constitutes “safe” security research are problematic. “It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not,” he said.

Chris Anley, chief scientist at NCC Group, explained that AI models are often used to test whether a software bug can be exploited—a critical step in confirming vulnerabilities. When guardrails block these prompts, they hinder both offensive and defensive security efforts. “The same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked,” Anley said, comparing AI to a hammer that can build or destroy.

Researchers Turn to Open-Source Alternatives

Frustrated by restrictions, some researchers have shifted to open-source AI models with no guardrails, including Chinese-developed tools like GLM. These models can be run locally, avoiding cloud-based data-sharing risks. Paolo Stagno, CTO of Crowdfense, said AI companies’ vetted programs treat users “like children who need babysitting,” pushing professionals toward unrestricted alternatives.

Giuseppe Cali, a security researcher, noted that while he uses AI for reverse engineering and tool-building, he avoids relying on it for discovering or weaponizing vulnerabilities. “I still want to own the actual bug discovery and weaponization myself,” he said, emphasizing that human expertise remains irreplaceable in offensive security.

Others, like an anonymous researcher at a smartphone-component manufacturer, reported that strict guardrails render AI tools nearly useless for vulnerability research. “If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the researcher said.

What Happens Next in AI Security Research

Experts warn that overly restrictive guardrails may push responsible researchers toward less regulated, foreign-owned AI systems. Chris Thompson, CEO of RemoteThreat, argued that AI labs should expand access for vetted professionals rather than tightening controls. “There’s this big wave of attacks that are going to happen at speed and scale like never before,” he said. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”

As AI continues to evolve, the balance between security and usability will remain a key challenge. Observers should watch for policy shifts from AI developers and governments, as well as the growing adoption of open-source alternatives in cybersecurity research.