AI Guardrails Are Hindering Offensive Cybersecurity Research, Experts Say

Cybersecurity researchers say AI guardrails from Anthropic and OpenAI are making legitimate vulnerability research harder, pushing some professionals toward unrestricted open-source models.

Jul 25, 2026 - 15:35
 0
AI Guardrails Are Hindering Offensive Cybersecurity Research, Experts Say
Image Credit: Chatgpt

Artificial intelligence companies have spent months introducing stricter cybersecurity guardrails and special access programs designed to prevent hackers from using powerful AI models for malicious purposes. While those restrictions are intended to reduce cyber risks, many offensive cybersecurity researchers argue they are also making it harder for legitimate defenders to identify and fix software vulnerabilities before attackers can exploit them.

The debate has intensified following recent restrictions placed on Anthropic's advanced AI models. In June, the U.S. government imposed export controls on Anthropic's Mythos and Fable models after concerns emerged that their cybersecurity safeguards could potentially be bypassed. Although the restrictions on Fable 5 were lifted on July 1 and Mythos 5 has since returned to vetted U.S. organisations as part of an ongoing government review, the episode renewed discussion about how powerful AI systems should be controlled.

Anthropic has promoted Mythos as a highly capable model intended only for carefully verified users operating under strict safeguards. The company also runs a Cyber Verification Program that grants approved researchers access to models with fewer cybersecurity restrictions. OpenAI offers a similar initiative through its Trusted Access for Cyber program.

Despite those programs, many security professionals believe the restrictions interfere with legitimate research. Their work often involves discovering previously unknown software flaws, known as zero-days, and developing proof-of-concept exploits so organisations can understand and fix security weaknesses before criminals discover them.

Mark Dowd, a longtime security researcher known for discovering and selling zero-day vulnerabilities to Western governments, recently questioned whether AI companies should determine what types of cybersecurity research are considered acceptable. Speaking on a cybersecurity podcast, Dowd argued that large technology companies are making arbitrary decisions about what is safe in security research.

Dowd acknowledged that his background may influence his perspective. His career has involved selling undisclosed software vulnerabilities and exploits to governments for intelligence operations rather than reporting them directly to software vendors. Because those vulnerabilities remain unpatched, they can carry significant value for intelligence agencies.

Other offensive security experts share similar concerns, even if their work differs from Dowd’s. Chris Anley, chief scientist at cybersecurity consultancy NCC Group, said AI has become an important tool for validating whether a suspected software flaw can actually be exploited.

According to Anley, asking an AI model to attempt an exploit is often a critical step in determining whether a vulnerability represents a real security risk. When a model refuses to answer because of cybersecurity guardrails, researchers lose an efficient way to verify issues that ultimately need to be fixed.

Anley said the problem is that the same prompt can serve both defensive and offensive purposes. A request to explain how vulnerable code might be exploited can help defenders understand how to patch a weakness, but it could also assist attackers. In his view, those two uses cannot easily be separated.

He compared AI tools to a hammer, arguing that while a hammer can be used as a weapon, it is also indispensable for building a house. Likewise, AI can simultaneously function as both an offensive and defensive cybersecurity tool.

When frontier AI models refuse to assist, Anley said researchers sometimes switch to open-source models that include few or no guardrails, allowing them to continue their work without restrictions.

Paolo Stagno, chief technology officer at vulnerability broker Crowdfense, also criticised the growing use of AI restrictions. He argued that companies often treat professional researchers as though they require constant supervision through verification programs and usage limitations.

Stagno said his team uses frontier AI models primarily for reverse engineering software rather than discovering vulnerabilities or developing exploits. For sensitive offensive research, however, they rely on locally hosted open-source models to avoid exposing confidential vulnerability information to cloud-based AI systems or future model training.

Independent researcher Giuseppe Cali takes a different approach. He said guardrails have little effect on his work because he does not use AI to discover or weaponise vulnerabilities. Instead, he relies on AI during the early stages of reverse engineering to better understand unfamiliar code and create supporting tools that improve productivity.

Cali said he still prefers to perform vulnerability discovery and exploit development himself, explaining that he enjoys the process too much to hand it over to an AI system, regardless of whether restrictions are relaxed.

One researcher employed by a smartphone component manufacturer, who requested anonymity because he was not authorised to speak publicly, said his employer does not participate in Anthropic’s Cyber Verification Program. As a result, he described the company’s models as largely unusable for vulnerability research because they frequently refuse to engage once they detect security-related requests.

Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, said another challenge is inconsistency. In his experience, the same prompts can produce different responses from one day to the next, even within Anthropic’s and OpenAI’s approved researcher programs.

Thompson said researchers often spend more time negotiating with AI models than analysing vulnerabilities because the systems sometimes over-sanitise responses or apply restrictions unpredictably.

Those limitations, he argued, are pushing legitimate researchers toward freely available open-source models, including Chinese-developed systems such as GLM, which can be downloaded, run locally, and used without verification requirements or cybersecurity restrictions.

According to Thompson, responsible researchers are increasingly moving away from U.S.-governed AI systems in favour of unrestricted alternatives, a trend he believes could ultimately weaken cybersecurity rather than strengthen it.

Instead of imposing tighter guardrails, Thompson said frontier AI companies should expand responsible access programs while holding users accountable for misuse. Otherwise, he warned, legitimate defenders could struggle to keep pace with increasingly sophisticated AI-powered cyberattacks.

He argued that the cybersecurity industry faces a future in which attacks will occur at unprecedented speed and scale, making advanced AI tools increasingly essential for defenders as well as attackers.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Angry Angry 0
Sad Sad 0
Wow Wow 0
Shivangi Yadav Shivangi Yadav reports on startups, technology policy, and other significant technology-focused developments in India for TechAmerica.Ai. She previously worked as a research intern at ORF.