Why Anthropic’s New Fable AI Has Outraged Cyber Experts – Claril Noticias

Cybersecurity researchers worldwide are expressing intense frustration this week following Anthropic’s Tuesday launch of Fable—a restricted public version of its Mythos security model—due to over-sensitive safety guardrails that block even basic, harmless coding tasks.

Why Security Experts Are Frustrated by Fable

The newly released AI assistant has faced immediate backlash, with a number of cybersecurity researchers and professionals taking to social media and Reddit to air their complaints.

“[Fable] rejects any request that could be tangentially cyber related. Even innocuous tasks like reading a blog post,” stated Valentina “Chompie” Palmiotti, a prominent security researcher at IBM X-Force.

When users trigger these sensitive boundaries, Fable abruptly halts the conversation, displaying a warning that its safety measures have flagged the prompt for cybersecurity or biology-related issues.

The Reason Behind Anthropic’s Strict Guardrails

These stringent restrictions were implemented to prevent Fable from being weaponised to create malware or exploit software vulnerabilities—a long-standing concern for Anthropic. Similarly, the biological restrictions aim to mitigate risks associated with the development of biological weapons.

When Anthropic originally launched its more powerful Mythos model in April, access was strictly limited to a select group of organisations under “Project Glasswing,” designed to protect critical infrastructure. Last week, the company expanded Mythos access to hundreds of organisations across 15 countries. Fable was intended to bring a version of this power to the wider public, but the execution has left experts underwhelmed.

A Keyword-Based System Impeding Best Practices

Industry experts argue that the current restrictions are too blunt. Matt Suiche, a cybersecurity veteran and member of the technical staff at AI security startup Tolmo, explained that Fable struggles to differentiate between malicious intent and defensive engineering. “If you ask it to write secure code, it assumes it is cybersecurity-related work instead of software engineering best practices, and you get downgraded,” Suiche noted.

If Fable hits a guardrail, it automatically falls back to Claude Opus 4.8. According to Suiche, the detection mechanism appears to be heavily keyword-based, meaning any terminology associated with cybersecurity triggers an immediate block.

Other professionals have shared similar frustrations, with one researcher griped on X that the model refuses to perform basic code reviews.

However, Suiche suggests that a cautious approach is logical during early deployment. “It is understandable as we are still in the early days and they are still adapting their guardrails,” he said, adding that collaboration between frontier AI firms and modern cybersecurity startups will likely help refine these systems over time. “It’s better to catch more people than not enough when you do such a release and to relax the guardrails over time.”

Anthropic did not immediately respond to a request for comment regarding the feedback.

How Professionals Can Bypass Fable’s Restrictions

To avoid these disruptive guardrails, Anthropic requires legitimate cybersecurity practitioners to apply for its Cyber Verification Programme. Approved users are granted access with significantly fewer limitations. This approach mirrors OpenAI’s vetting system, which offers a similar Trusted Access for Cyber initiative for verified industry professionals.

By Claril

Leave a Reply

Your email address will not be published. Required fields are marked *