AI Safety Measures Are Blunting Cyber Defences – Claril Noticias

Strict AI safety guardrails designed to block malicious hackers are currently hindering the vital work of legitimate offensive cybersecurity researchers and network defenders globally, industry experts warned this month.

The Fight Over Restricted AI Models

In June, the US government imposed export control restrictions on Anthropic’s highly publicised AI models, Mythos and Fable. The decision was prompted by reports demonstrating that users could bypass safety features designed to prevent the creation and execution of cyberattacks.

While the export controls on Fable 5 and Mythos 5 have since been lifted—with Fable 5 returning to general access on 1 July and Mythos 5 restricted to vetted US organisations—the incident highlights a growing friction. Anthropic has consistently marketed Mythos as a powerful capability requiring strict gatekeeping, a trend mirrored across the tech sector.

Vetted Programmes Fail to Solve the Friction

This gatekeeping is not unique to Anthropic. Both OpenAI and Anthropic have launched specialised pathways for security professionals, such as OpenAI’s Trusted Access for Cyber programme and Anthropic’s Cyber Verification Programme.

However, these safety measures face intense criticism from offensive security researchers whose primary role is to discover and exploit zero-day vulnerabilities before cybercriminals can abuse them.

Veteran security researcher Mark Dowd, who has spent decades unearthing and selling zero-days to Western governments rather than disclosing them to software vendors, expressed his concerns on a security podcast. Dowd stated that private AI corporations should not be making arbitrary decisions about what is deemed safe within the cybersecurity landscape.

var playerInstance_jwplayer_6a7a0661a8275 = jwplayer( “jwplayer_6a7a0661a8275” );
playerInstance_jwplayer_6a7a0661a8275.setup({
playlist: “https://cdn.jwplayer.com/v2/media/lv0GaEwB”,
});

The Dual-Use Dilemma of AI Security Tools

Legitimate offensive cyber specialists find themselves battling the same barriers meant for adversaries. Chris Anley, Chief Scientist at NCC Group, explained that asking AI models to exploit bugs is crucial for verifying vulnerabilities. When guardrails trigger outright refusals, defenders suffer.

“Fix this code” serves as a dual-use prompt: it helps secure software, but it also provides a blueprint for identifying critical flaws. Anley compared the technology to a hammer—an essential tool for construction that can also be used as a weapon. Consequently, when blocked by commercial systems, researchers often turn to open-source models that lack safety restrictions.

How Researchers Navigate Restrictive Guardrails

Paolo Stagno, Chief Technology Officer at Crowdfense—a firm specialising in acquiring zero-days for government agencies—echoed these frustrations. Stagno noted that major AI developers treat professional clients like children requiring constant supervision. To avoid leaking highly sensitive data to cloud-hosted models, Stagno’s team limits their use of frontier models to reverse engineering, choosing local, open-source models for actual vulnerability research.

Similarly, security researcher Giuseppe Cali utilises AI primarily for initial code comprehension and tool building rather than direct exploit development. Cali emphasised that even if guardrails were removed entirely, he prefers to retain ownership of the actual bug discovery and weaponisation process.

Conversely, an anonymous researcher at a smartphone-component manufacturer revealed that because their employer lacks access to Anthropic’s vetted programme, commercial AI tools are virtually useless for security audits, shutting down the moment any security-related activity is detected.

The Rise of Foreign AI Alternatives

Chris Thompson, Chief Executive of RemoteThreat, pointed out that guardrail behaviour remains highly erratic, even within vetted programmes. Security professionals frequently waste valuable time negotiating with models to bypass over-sanitisation rather than analysing vulnerabilities.

This friction is driving Western researchers towards unrestricted, downloadable Chinese open-source models such as GLM. Thompson warned that pushing responsible domestic researchers away from US-governed platforms to foreign-owned systems does more harm than good. He urged AI labs to expand access and hold bad actors accountable directly, warning that overly restrictive guardrails risk leaving defenders ill-equipped for an imminent wave of highly automated cyber threats.

By Claril

Leave a Reply

Your email address will not be published. Required fields are marked *