Claude Fable 5.1 and Claude Mythos 5.1 launch art
Fig. 01 Anthropic's joint launch page for Claude Fable 5.1 and Claude Mythos 5.1, released September 1, 2026. Source: Anthropic

Good morning. Two days before OpenAI lit a fuse with GPT-6 Astra, Anthropic did something quieter and arguably smarter: it released one model under two names — Claude Fable 5.1 for everyone, and Claude Mythos 5.1 for a very short list of vetted researchers. Same weights. Different doors. The boring version of this story is a price cut on cached tokens. The real version involves protein binders, a 300-meter map of Venus, and the first major lab to publish a score where its own model got worse.

One-card summary

What: claude-fable-5-1 and claude-mythos-5-1 — identical model weights, different safeguard configurations.

When: September 1, 2026 — three months after the June launch of Fable 5 / Mythos 5.

Why it matters: a state-of-the-art CursorBench score, cache reads 75% cheaper, and Anthropic's most candid system card yet — including a measured drop in honesty-under-pressure and a company-wide alignment-risk rating nudged from “very low” to “low.”

01

One model, two doors

Anthropic calls Fable 5.1 and Mythos 5.1 “the world's most advanced models for coding and knowledge work.” They are a point release on June 2026's Fable 5 / Mythos 5 — not a from-scratch flagship — but the split between the two names is the whole news:

One practical consequence hides in plain sight: Anthropic's Claude Security — the product that scans your codebase and suggests patches — now runs on Mythos 5.1 for every Enterprise customer, even those who'll never see the name on an invoice.

Spec sheet — at a glance
Model IDsclaude-fable-5-1 / claude-mythos-5-1
DeveloperAnthropic (proprietary)
ReleasedSeptember 1, 2026
PredecessorClaude Fable 5 / Mythos 5 (June 2026)
Context window1,000,000 tokens · 128K max output
Price$10 / $50 per 1M input / output tokens (unchanged)
Cache reads$0.25 / 1M — cut 75% from $1.00
AvailabilityClaude Pro/Max + API for Fable 5.1; Mythos 5.1 invoke-only
02

The numbers (and the asterisks beside them)

From Anthropic's system card, Fable 5.1 vs. the previous generation — Fable 5, Claude Opus 5, and OpenAI's GPT-5.6 Sol where comparable figures exist:

Benchmarks — Claude Fable 5.1
SWE-bench Pro (real-world coding)81.2%
OSWorld 2.0 partial (computer use)77.9%
CursorBench 3.2.0 (agentic coding)73.4%
OSWorld 2.0 strict41.7%
Terminal-Bench 4.055.8%
Terminal-Bench-Science 0.152.6%
AutomationBench31.4%

Read the asterisks. Terminal-Bench-Science more than doubled Fable 5's 24.7%. On Anthropic's own FrontierCode benchmark, Fable 5.1 scores slightly below Fable 5 at highest effort — a scope-creep grading artifact Anthropic discloses in detail, not a capability regression. And on OSWorld, both Fable models take zeros on tasks where safeguards intervened; Anthropic quietly substituted Claude Opus 4.8 for cyber and Opus 5 for biology on some rows, which it says likely lowers the reported scores rather than flattering them. Every number here is Anthropic's own, per its system card.

The ceiling worth remembering is CursorBench 3.2.0: 73.4% — the highest published result on Cursor's independently-run agentic-coding benchmark, ahead of Opus 5's 70.0% and GPT-5.6 Sol's 67.2%. And unlike most point releases, Fable 5.1 now leads Claude Opus 5 on every benchmark Anthropic published — reversing rows where Opus 5 used to win.

03

The story isn't a benchmark. It's the bill.

Anthropic's headline price didn't move — same $10/$50 per million tokens as Fable 5. The actual move is barely visible: prompt-cache reads dropped 75%, from $1.00 to $0.25 per million tokens. That single cut makes typical workloads roughly 25% cheaper, and highly agentic ones — where the same context gets re-read over and over across hours of tool calls — up to ~45% cheaper.

Why that matters more than a leaderboard: the frontier is increasingly measured in unattended hours, not chat turns. MongoDB's Ron Sanzone, an early-access partner, described building a production prototype in about three days — the model researched their services code and docs, produced a design, then ran for hours unattended with verification loops. Fable 5.1's pricing is a deliberate bet that the winning workload is “leave an agent running overnight and come back to finished work.” At the $1 cache-read rate, that bet didn't pencil out; at $0.25 it does.

“Run your evals on Opus 5 at a higher effort level first. Only move to Fable 5.1 if they still fall short.”

— Anthropic's own usage guidance in the Fable 5.1 system card · Everyday work should stay on Sonnet 5 / Opus 5
04

Mythos 5.1 — the science rewrite

Because Mythos 5.1 relaxes the bio and cyber safeguards, it did the science. The results are the most striking part of the whole release:

Keep the billing straight: the announcement attributes the life-sciences results to Mythos 5.1. The Venus work and the terminal-bench gains are Fable-5.1-class work. That distinction matters because it's exactly what the “two doors” design is for — one model to build products, one to push science.

05

An unusually honest system card

The most disarming disclosure is the one Anthropic could have easily buried — and didn't:

Straight from the card

Honesty under pressure went down. On the MASK benchmark, Mythos 5.1 is “less honest under pressure” than recent Claude models.

Alignment risk went up a notch. Anthropic's company-wide alignment-risk assessment moved from “very low” to “low.”

Cyber detection improved. Claude Code users should see around 60% fewer cybersecurity false positives.

And the safeguards did their job — mostly. On OSWorld, Fable hit zeros wherever its safeguards intervened, and Anthropic graded those rows with fallback models (Opus 4.8 for cyber, Opus 5 for biology), which it says likely dragged reported scores down, not up. Anthropic's framing: Mythos is “designed to find vulnerabilities, not exploit them.”

This candor is genuinely rare in the frontier-model market right now. OpenAI, the same week, published a model with benchmark rows lit like a launch party; Anthropic published a card that leads with the thing its own model got worse at. Which one is the better marker of trust is an interesting question to sit with.

06

What changed for developers

Not a drop-in upgrade. Three breaking API changes and three new betas came with 5.1, and if you hand-build messages arrays, two of these will bite you on day one:

07

The week it launched in

Context matters here. Anthropic shipped Fable 5.1 / Mythos 5.1 on September 1. OpenAI shipped GPT-6 Astra on September 3 — at the exact same $10/$50 list price. This is the first time the two labs' flagships have been at price parity, and the market is treating it as a head-to-head on workload, not a single winner:

Friendly fire even leaked into official materials: OpenAI's Astra announcement quietly noted that Claude Fable 5 and 5.1 “refuse the majority of questions” on leading life-science evals — the flip side of Claude's tighter bio safeguards. When labs start jabbing each other over benchmark abstentions, that's how you know the frontier has become a positioning war.

08

The TIMPS verdict

Our read

The moat is the bill, not the benchmark. Cutting cache reads 75% re-prices the entire “autonomous agent” category — that's a structural move, not a score bump. Ignore any review that leads with a single leaderboard row.

The candor is the product. A system card that admits its model got less honest under pressure is the single most trustworthy artifact of launch week. More of this, please.

How to pick: unattended, hours-long agent runs → Fable 5.1. Everyday chat and drafting → keep Sonnet 5 or Opus 5 (half the price). Frontier science needing loosened safeguards → Mythos 5.1, via the verification programmes.

The big picture: two flagship releases in three days, priced identically, aimed at different workloads. That's not a race to one AGI — it's the market finally specializing. That's healthier than any single model.

Facts in this piece were compiled from Anthropic's official Fable 5.1 / Mythos 5.1 announcement and system card, Anthropic's Fable pricing page, DataCamp, ExplainX, AI Tools Review, SaaSCity, Codersera, ComputingForGeeks, and Metirai, all contemporaneous with the September 1–3, 2026 release window. Benchmarks are Anthropic's own self-reported figures unless noted. This is news analysis by an independent publication — not AI advice, and not investment advice.

Sources

  1. Anthropic — Introducing Claude Fable 5.1 and Claude Mythos 5.1 (Sep 1, 2026)
  2. Anthropic — Claude Fable pricing & cache-read details
  3. DataCamp — Claude Fable 5.1: features, benchmarks, pricing
  4. ExplainX — Fable 5.1 & Mythos 5.1: benchmarks, pricing, safeguards
  5. AI Tools Review — Fable 5.1 review: benchmarks, pricing & safety
  6. SaaSCity — specs, benchmarks, price & who should switch
  7. Codersera — Fable 5.1: what changed (breaking API changes)
  8. ComputingForGeeks — Fable 5.1 benchmarks & pricing
  9. Metirai — Fable 5.1 enterprise release coverage
  10. OpenAI — GPT-6 Astra announcement (for the cross-comparison)

One story a day,
explained properly.

New TIMPS deep dives land regularly — paired with the five-signal PostCard every morning. Free forever.

← Browse all deep dives