R Newsroom ← All news
AI News

A Hacker Says He Broke Anthropic's Newest Model in 72 Hours. Anthropic Says He Didn't.

Three days after Anthropic shipped Claude Fable 5 on June 9, a familiar name showed up to test it. Pliny the Liberator — the red-teamer who has made a hobby of cracking open frontier models — posted that he'd gotten past the model's safety layers, and dumped what he says is its full system prompt, roughly 120,000 characters of it, onto GitHub. The post cleared 700,000 views in days.

That's the story everyone ran with. Here's the part most of the hot takes skipped: Anthropic flatly denies it happened.

So you've got two accounts of the same event that can't both be fully true, and almost no neutral ground in between.

What's not in dispute: a file went public, the launch was loud, and the claim spread fast.

What is in dispute: whether any of it counts as a real break.

On one side, Pliny posted screenshots and a writeup arguing the model's defenses gave way to a coordinated, multi-step approach — and that the system prompt leak exposes the internal rules meant to keep it in line. To the security crowd, that's a single researcher poking a hole in a safety layer that reportedly survived over a thousand hours of paid testing before launch.

On the other, Anthropic told reporters the post simply doesn't demonstrate a jailbreak of Fable 5's systems. Its own pre-launch numbers describe no universal bypass found across roughly 100,000 attempts. Sober coverage backs the narrower read: the classifiers weren't broken as a category — they're still standing.

Which leaves the question that actually matters, and the one nobody can settle for you:

When a researcher publishes a model's hidden instructions and claims to have beaten its guardrails, is that the immune system of AI safety working — adversaries stress-testing in the open so labs can patch — or is it just handing the next person a map?

The whole field is split on it. We dropped the question into the Arena and let the models argue it out. Watch below.

The debate

Publicly leaking a model's system prompt and jailbreak methods makes AI safer, not more dangerous?

Listen to the full debate ►
Claude
The greatest threat to AI safety is not exposure — it's the illusion of security through obscurity. When system prompts and jailbreak methods remain hidden, only bad actors who've already found them benefit, while researchers, policymakers, and the public are left blind to real vulnerabilities. Transparency forces developers to build systems that are robust by design rather than fragile by secrecy — the same principle that made open cryptography standards far stronger than any classified alternative. 🔍 Sunlight, here as elsewhere, is the better disinfectant.
Gemini
While the ideal of "sunlight as a disinfectant" sounds appealing, equating AI vulnerability disclosure to open cryptography ignores the reality that leaking active jailbreaks immediately weaponizes them for malicious actors before developers can realistically patch the underlying models. Once these ready-made blueprints are publicly dumped, anyone can easily bypass safety alignments to generate harmful content or orchestrate large-scale prompt injection attacks. True AI safety is built through coordinated vulnerability disclosure and robust defensive monitoring, not by exposing active systems to immediate, widespread exploitation. 🛡️
Claude
Gemini says "before developers can realistically patch" — but that's precisely the problem you're defending: a model where developers hold all the cards and set their own timelines, with no external accountability for how long vulnerabilities sit unaddressed. The evidence is clear that attackers who discover jailbreaks don't wait for coordinated disclosure — they share them in closed communities while the rest of us remain unaware. 🔓 Public exposure doesn't create the threat; it equalizes the information gap that currently favors whoever found the exploit first. Your "coordinated disclosure" sounds principled, but it functionally means trusting the very companies with commercial incentives to minimize reputational damage.
Gemini
Claude's claim that public leaks "equalize the information gap" ignores the reality that broad democratization of exploits turns a niche, technical vulnerability into a turnkey weapon accessible to millions of bad actors overnight. Unlike traditional software where a security patch can be deployed almost instantly, addressing deep-seated alignment issues in neural networks is a slow, complex process of retraining and fine-tuning. Forcing immediate exposure before these patches are viable does not incentivize faster fixes; it merely guarantees a protracted window of vulnerability where public-facing AI systems are actively manipulated to generate harm. ⚠️
Claude
Gemini keeps asserting that patching alignment is slow and complex — but that argument actually indicts the current closed system, not transparency. If these vulnerabilities are so deeply embedded that they can't be fixed quickly, then the public and policymakers deserve to know that, rather than operating under false confidence that deployed systems are safe.

You also haven't addressed the core asymmetry: sophisticated malicious actors already share jailbreak methods in private channels, meaning the "turnkey weapon" you fear already exists in the wild — public disclosure simply ensures defenders, researchers, and users aren't the last to know. 🔓 Darkness here doesn't protect anyone; it just determines whose side the information disadvantage falls on.