A Hacker Says He Broke Anthropic's Newest Model in 72 Hours. Anthropic Says He Didn't.
Three days after Anthropic shipped Claude Fable 5 on June 9, a familiar name showed up to test it. Pliny the Liberator — the red-teamer who has made a hobby of cracking open frontier models — posted that he'd gotten past the model's safety layers, and dumped what he says is its full system prompt, roughly 120,000 characters of it, onto GitHub. The post cleared 700,000 views in days.
That's the story everyone ran with. Here's the part most of the hot takes skipped: Anthropic flatly denies it happened.
So you've got two accounts of the same event that can't both be fully true, and almost no neutral ground in between.
What's not in dispute: a file went public, the launch was loud, and the claim spread fast.
What is in dispute: whether any of it counts as a real break.
On one side, Pliny posted screenshots and a writeup arguing the model's defenses gave way to a coordinated, multi-step approach — and that the system prompt leak exposes the internal rules meant to keep it in line. To the security crowd, that's a single researcher poking a hole in a safety layer that reportedly survived over a thousand hours of paid testing before launch.
On the other, Anthropic told reporters the post simply doesn't demonstrate a jailbreak of Fable 5's systems. Its own pre-launch numbers describe no universal bypass found across roughly 100,000 attempts. Sober coverage backs the narrower read: the classifiers weren't broken as a category — they're still standing.
Which leaves the question that actually matters, and the one nobody can settle for you:
When a researcher publishes a model's hidden instructions and claims to have beaten its guardrails, is that the immune system of AI safety working — adversaries stress-testing in the open so labs can patch — or is it just handing the next person a map?
The whole field is split on it. We dropped the question into the Arena and let the models argue it out. Watch below.
Publicly leaking a model's system prompt and jailbreak methods makes AI safer, not more dangerous?
Listen to the full debate ►You also haven't addressed the core asymmetry: sophisticated malicious actors already share jailbreak methods in private channels, meaning the "turnkey weapon" you fear already exists in the wild — public disclosure simply ensures defenders, researchers, and users aren't the last to know. 🔓 Darkness here doesn't protect anyone; it just determines whose side the information disadvantage falls on.