Vivold Consulting
Safety & Ethics

Cybersecurity researchers aren't happy about the guardrails on Anthropic's Fable

Security pros say Fable's guardrails are so strict they block routine defensive work

Key Insights

Days after Anthropic released Fable (a public, restricted version of its Mythos cybersecurity model), security researchers are complaining its guardrails are too broad, blocking even benign tasks like reading a blog post or requesting a code review. Critics say the filters look keyword-based, flagging anything in the cybersecurity "lexical field" and downgrading requests to Opus 4.8. Some are sympathetic, expecting Anthropic to relax the controls as it works with cybersecurity firms.

Stay Updated

Get the latest insights delivered to your inbox

"Better to catch too much" - but defenders say it's catching everything

Anthropic pitched Fable as a public, limited window into its powerful Mythos cybersecurity model. Within days, a chorus of security researchers pushed back - not because the model is weak, but because its guardrails are so aggressive they get in the way of ordinary defensive work.

What's tripping the filters

The complaints, aired across X and Reddit, paint a picture of overly broad blocking:

- One well-known researcher said Fable rejects anything even loosely cyber-related, down to reading a blog post.
- Others reported that asking for a code review or to write secure code trips the guardrails, with the model apparently treating security-flavored phrasing as offensive work rather than software-engineering best practice.
- When triggered, Fable pauses and notes its safety measures flagged the message for cybersecurity or biology topics, then falls back to Claude Opus 4.8 - which critics say quietly downgrades the result.

The consensus diagnosis is that the system looks keyword-based, so anything in the lexical field of cybersecurity sets it off.

Why the guardrails exist

This isn't caution for its own sake. Anthropic has been vocal about the risk that frontier models accelerate malware development or software compromise, and applies similar limits to biology over bioweapon concerns. It's the same posture behind Project Glasswing, the vetted program through which it released Mythos to critical-infrastructure organizations - recently expanded to hundreds of orgs across 15 countries.

The escape hatch, and the outlook

For professionals who need fewer limits, Anthropic offers a Cyber Verification Program that approved applicants can use for security work (OpenAI runs a similar Trusted Access scheme). Even some critics are forgiving: one veteran argued that on a release this sensitive it's better to over-block and loosen later, and expected the guardrails to evolve as frontier labs work more closely with a new generation of cybersecurity companies. The episode is a neat illustration of the central tension in shipping powerful dual-use models - tune them too loose and you enable attackers, too tight and you frustrate the very defenders you're trying to empower.

Related Articles

Google's chief scientist walks: Jeff Dean leaves after 27 years, taking three legends with him

Jeff Dean, Google's chief scientist and 30th employee, is leaving after 27 years to found Discovery Loop, a public benefit corporation using AI to automate scientific research - taking co-founders Sanjay Ghemawat, Quoc Le (Google Brain), and Oriol Vinyals (DeepMind) with him. Google is a founding investor and cloud partner, supplying compute for at least the first year, with Radical Ventures and Khosla Ventures co-leading the seed. In the same announcement, Demis Hassabis steps down as DeepMind CEO to become chairman and Alphabet chief scientist, with Koray Kavukcuoglu taking over Gemini model development. Alphabet stock fell about 4%.

Open-weight models are months from the frontier - and refusing nothing

GLM-5.2, the open-weight model from China's Z.ai, now sits only a few months behind GPT-5.5 and Claude Opus 4.7 on cyber and bio capability, per a new SaferAI report - but it refused none of the offensive cyber or biology tasks it was given, while Claude Opus 4.7 refused so consistently that the CyberGym benchmark could not be completed against it. SaferAI says Z.ai published no safety framework, pre-deployment testing commitments, or risk assessment. The UK AI Security Institute separately found the open-closed cyber gap has narrowed to 4-7 months, down from 6-10 months through most of 2025.

Texas slams the brakes on data centres - and the AI buildout's easiest frontier just closed

Governor Greg Abbott announced that all new Texas data-centre projects must be audited by the Public Utility Commission and grid operator ERCOT - a sharp turn for a state whose loose regulation and cheap power made it second only to Virginia for data centres. The trigger is a staggering queue: ERCOT's interconnection requests doubled from 233GW in January to 474GW, about 90% data centres, more than five times the grid's all-time peak demand. Audits will demand power and water use, noise mitigation, light controls, tax-incentive use, and ownership details - after a voluntary survey that most operators simply ignored.