Anthropic Releases Claude Fable 5, Its Most Powerful AI Yet, With Cyber Safeguards
09:37 · June 10, 2026 · Hacker News AI Section

On June 9, Anthropic released Claude Fable 5, the most capable model it has ever made, generally available. It also did something unusual: it shipped one model as two products, split not by capability but by a layer of safety classifiers. Fable 5 goes to the public. Its twin, Claude Mythos 5, the same underlying model with the cyber safeguards lifted, stays locked to a vetted group of cyber
Summary
Anthropic released Claude Fable 5 on June 9 as its most capable model to date, making it generally available through the Claude API. The company introduced an unusual deployment split: the same underlying model is offered in two forms that differ only in the presence of safety classifiers. Fable 5 reaches the public with these classifiers active, while Claude Mythos 5, the unrestricted variant, remains limited to vetted cybersecurity professionals and critical-infrastructure operators.
The classifiers monitor requests for cyber, biology, chemistry, and model-distillation misuse. When a query triggers a flag, Fable 5 hands the response to the weaker Claude Opus 4.8 rather than refusing outright, and it informs the user of the handoff. The cyber classifier targets the full attack chain, including reconnaissance, lateral movement, and exploit development. Internal tests and external red-team evaluations showed that the safeguards blocked all progress on single-turn harmful cyber tasks and resisted 30 public jailbreak techniques, though one partner noted the UK AI Security Institute made limited headway toward a universal jailbreak in early testing.
Both versions share the same pricing of $10 per million input tokens and $50 per million output tokens. Fable 5 is included at no extra cost on Pro, Max, Team, and Enterprise plans through June 22 before shifting to usage credits. Anthropic states that fallback to Opus 4.8 occurs in under 5 percent of sessions overall and plans to tighten the classifiers after launch to reduce false positives.
The split addresses the risk that Mythos-class capabilities could give attackers meaningful advantage in finding and exploiting vulnerabilities at scale. Earlier limited testing with Mythos Preview demonstrated the model locating and weaponizing zero-days across major operating systems and browsers, including a remote-code-execution flaw in FreeBSD. At the same time, the restricted access enabled roughly 50 partners to discover more than ten thousand high- or critical-severity bugs in important software. The resulting pressure has shifted the bottleneck from discovery to triage and patching, prompting Anthropic to introduce a 30-day data-retention policy for all traffic on these frontier models to support ongoing safety monitoring.
Why it matters
This article is highly relevant for security professionals as it highlights a novel approach to AI model deployment, separating public safety from advanced cybersecurity research. Dutch and EU practitioners can leverage this to understand how foundational models are addressing systemic cyber risks and compliance with ethical AI standards.







