A self-described AI skeptic warned the world that Mythos was too dangerous to ship. Now he is the one lobbying the White House to let a version of it back online.
Answer capsule: Anthropic dispatched senior security researcher Nicholas Carlini to Washington to lobby the Trump administration to reverse export controls that forced the company to pull its two most powerful models, Mythos 5 and Fable 5. The irony is hard to miss: Carlini is the researcher who, months earlier, urged Anthropic not to release Mythos at all. The dispute centers on an Amazon report claiming Fable’s safety guardrails could be partially bypassed to surface software vulnerabilities. Anthropic disputes the severity and argues the same capability is already available in rival public models.
The skeptic who switched sides
Nicholas Carlini, 35, built his reputation as the cybersecurity industry’s “professional skeptic,” the guy who debunked inflated AI safety claims for years. In 2019 he thought OpenAI was being unreasonable for suggesting GPT-2 was too dangerous to release. Then he got his hands on Mythos.
In March, Carlini stood in front of 700 security researchers in San Francisco and showed them how he used Anthropic’s AI to find and exploit critical bugs in Ghost, a web-publishing tool, and in Linux, one of the most scrutinized pieces of software on earth. He had never found a bug in either before. Suddenly he was finding many. Two days later he sent Anthropic a memo: “I don’t think we should release Mythos yet.”
That talk, now viewed more than 360,000 times, kicked off what researchers started calling “Bugmageddon.”
Why “Bugmageddon” spooked Washington
The core fear is scale. Carlini described Mythos as the first model that can find and exploit vulnerabilities at volume, not one at a time. He pointed the model at Linux and it churned through the codebase thousands of times, surfacing 479 bugs while a human would have given up. Anthropic says Mythos has now found more than 10,000 bugs across software the broader economy quietly runs on.
The defensive upside is real. The offensive risk is the problem. The same engine that helps a vendor patch its software helps an attacker write the exploit before the patch is installed. The Ghost case is the cautionary tale: Carlini reported the bug, the vendor patched it in February, and within weeks hackers reverse-engineered the fix and attacked more than 700 sites that had not updated.
The Amazon report that triggered the ban
The chain reaction started with a phone call. Amazon CEO Andy Jassy called administration officials, including Treasury Secretary Scott Bessent, to flag that Amazon researchers had found ways around the guardrails on Fable 5, the model Anthropic had just released to the public. Fable is essentially Mythos with safety measures bolted onto its cybersecurity, biology, and chemistry capabilities. If those measures can be peeled off, users get the raw model through a back door.
By Friday, National Cyber Director Sean Cairncross delivered an ultimatum: work with the government and pull the models that day, or face a ban on foreign users. Trump personally authorized the restriction and tapped Commerce Secretary Howard Lutnick to run it. When Amodei told Lutnick the rule meant Anthropic could not keep the model out, Lutnick’s reported reply was blunt: “That’s the point.” Anthropic shut down all access to comply, because the ban swept in foreign-born employees too, making a partial shutdown unworkable.
What this is really about for Anthropic
Here is the strategic read. Anthropic has spent years building a brand on AI safety, positioning caution as a feature, not a liability. That same posture is now being weaponized against it. White House AI adviser David Sacks has publicly accused the company of “scare tactics,” and the administration’s argument is essentially: you told us this was dangerous, so we are treating it as dangerous.
For a company carrying a roughly $1 trillion valuation and walking into IPO season, an export ban is not a PR scratch. It is a direct hit on the commercial availability of its flagship product and on the credibility of its government relationships. Anthropic’s counter is that the Amazon finding describes a narrow, non-universal jailbreak, that rival public models like GPT-5.5 can do the same thing, and that defenders use this capability routinely. The company also says it received explicit approval to deploy Fable in the first place.
That is why it sent technical firepower instead of lobbyists. Carlini, frontier red-team lead Logan Graham, and head of safeguards Dave Orr are the ones in the room. The bet is that a credible technical briefing beats a press release with this administration, which has repeatedly complained that Anthropic “speaks a different language.”
The precedent founders should watch
Strip away the personalities and this is a governance milestone. In early June, Trump signed an executive order requiring AI labs to give the government access to models 30 days before public release. The Fable shutdown is the first hard enforcement of that posture, and it establishes that a frontier model can be pulled from the market on a 90-minute notice over a contested safety finding.
For anyone building on top of these models, the lesson is that regulatory risk is now a core part of the AI supply chain, not a footnote. Distribution can be switched off by the state, fast, and “we got approval” may not be enough.
Quick Questions
Who is Nicholas Carlini? A senior security researcher at Anthropic, formerly known across the cybersecurity field as its leading skeptic of AI hacking claims, before his own research convinced him the threat was real.
What are Mythos 5 and Fable 5? Mythos is Anthropic’s most powerful model. Fable is the public-facing version with safety guardrails on high-risk domains like cybersecurity, biology, and chemistry. Both were pulled after the ban.
Why did the government ban them? An Amazon report claimed Fable’s guardrails could be partially bypassed to surface software vulnerabilities, alarming White House officials enough to impose export controls on foreign access.
Does Anthropic agree with the ban? No. It calls the finding a narrow jailbreak, not a full one, and argues the same capability already exists in competing public models.
What is the Carlini Loop? A prompting technique Carlini demonstrated that nudges the model to return different results on each pass through a codebase, helping it surface more bugs. He reportedly dislikes the name.
The Business Model Analyst Take
Anthropic built a moat out of trustworthiness, and this episode shows the double edge of that strategy. When your entire brand says “we take danger seriously,” you hand regulators the exact justification they need to act when they decide to. The company is now in the strange position of arguing its own product is less dangerous than its earlier warnings implied, while the researcher who sounded the loudest alarm flies to D.C. to talk everyone down.
The deeper signal for operators is about control. A near trillion-dollar company just had its flagship product switched off by the state with 90 minutes of warning, and its strongest defense, prior approval, did not hold. If your business plan assumes uninterrupted access to a frontier model, you are underwriting a political risk you do not price and cannot hedge. Build for the day the API goes dark, because for Anthropic’s customers, it just did.
