← Back to blog

Claude Fable 5.1 & Claude Mythos 5.1 Benchmarks Explained

Sep 2, 2026·10 min·By Nicolas Zeeb
Model Comparisons
Claude Fable 5.1 & Claude Mythos 5.1 Benchmarks Explained

Claude Fable 5.1 & Claude Mythos 5.1 Benchmarks Explained

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, 2026, three months after Fable 5. The framing is incremental. The name has a .1 in it, the sticker price is unchanged at $10 and $50 per million tokens, and Anthropic calls the pair "the world's most advanced models for coding and knowledge work." The table underneath the announcement is less modest: Fable 5.1 more than doubles Fable 5 on agentic scientific research, nearly doubles it on business workflows, and finishes ahead of Opus 5 on every category Anthropic published.

Same model, two doors. Fable 5.1 is generally available today as claude-fable-5-1 on the Claude API, plus AWS, Google Cloud, and Microsoft Azure. Mythos 5.1 is the identical model with lighter safeguards, restricted to vetted organizations through the Cyber Verification Program and the Life Sciences Verification Program, currently US-only. Here's the full walk through the table, in the order Anthropic published it.

1. Scientific research: Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 tests whether a model can run a scientific investigation end to end: plan it, execute it in a terminal, and check its own work. Fable 5.1 lands at 52.6%. Fable 5 scored 24.7% on Anthropic's setup, Opus 5 29.0%, and GPT-5.6 Sol 22.4%.

Terminal-Bench-Science 0.1
Accuracy, % (higher is better)
Fable 5.1
52.6
Opus 5
29.0
Fable 5
24.7
GPT-5.6 Sol
22.4
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table. Standard error is ±3.5 to 4.5 pts per model.

More than doubled in a .1 release. Anthropic's Felix Rieseberg said the same thing within the hour on Hacker News: the score "more than doubled Fable 5's Terminal-Bench-Science," and he called it meaningful. The standard error is ±3.5 to 4.5 points per model, so the doubling survives any reasonable error bar. One methodology note: the public leaderboard, running 3 trials per task in a Claude Code harness, has Opus 5 at 30.0% and Fable 5 at 21.4%; Anthropic says its setup reproduces both within noise at 29.0% and 24.7%.

2. Terminal work: Terminal-Bench 4.0

Terminal-Bench 4.0 measures long-horizon terminal work, and it's the one row where Mythos 5.1 appears: 60.9% against Fable 5.1's 55.8%, Opus 5's 52.3%, Fable 5's 42.0%, and GPT-5.6 Sol's 37.3%.

Terminal-Bench 4.0
Accuracy, % (higher is better)
Mythos 5.1
60.9
Fable 5.1
55.8
Opus 5
52.3
Fable 5
42.0
GPT-5.6 Sol
37.3
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table. Fable 5.1 and Mythos 5.1 are the same underlying model; the gap reflects tasks where earlier, less precise cyber safeguards intervened.

The two 5.1 variants are the same model. The five-point gap is tasks where Fable 5.1's cyber safeguards stepped in and Mythos 5.1's didn't, and Anthropic expects the more precise safeguards shipping today to shrink it. For anyone buying the generally available model, 55.8% is the honest number: an 8-point lead over Opus 5 on the row closest to predicting real terminal-agent work.

3. Knowledge work: GDPval-AA v2

GDPval-AA v2 is Anthropic's knowledge-work benchmark, scored in points. Fable 5.1: 1853. Opus 5: 1824. Fable 5: 1723. GPT-5.6 Sol: 1711.

GDPval-AA v2
Rubric score, points (higher is better)
Fable 5.1
1853
Opus 5
1824
Fable 5
1723
GPT-5.6 Sol
1711
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table.

The gap over Opus 5 is small, 29 points, about 1.6%. The more useful read is the lineage: Fable 5.1's jump over Fable 5 (130 points) is bigger than the jump Opus 5 made over Fable 5 (101). The Opus-tier deficit on professional knowledge work didn't just close, it flipped, inside one release cycle.

4. Computer use: OSWorld 2.0

OSWorld 2.0, scored two ways on the benchmark authors' August 2026 task release. Partial pass: Fable 5.1 77.9%, Opus 5 75.4%, Fable 5 72.9%. Strict pass: 41.7%, 39.6%, 36.1%. No GPT-5.6 Sol score.

OSWorld 2.0 (August 2026 task release)
Pass rate, % (higher is better)
PARTIAL PASS
Fable 5.1
77.9
Opus 5
75.4
Fable 5
72.9
STRICT PASS
Fable 5.1
41.7
Opus 5
39.6
Fable 5
36.1
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table. Task files differ from earlier OSWorld 2.0 releases, so these numbers aren't comparable to previously published OSWorld results, which is why no competitor score is shown.

Two pieces of fine print matter here. First, the task files differ from earlier OSWorld 2.0 releases, so none of these numbers compare to previously published OSWorld results. Second, both Anthropic models scored a zero wherever their safeguards intervened. Strict-mode 41.7% is the number that says computer use still has a long way to go. The five-point lead over Opus 5 is the number that says Fable 5.1 leads anyway.

5. Reasoning: Humanity's Last Exam

Humanity's Last Exam, scored without tools and with tools. No tools: Fable 5.1 60.9%, Fable 5 57.8%, Opus 5 56.6%. With tools: 65.0%, 63.8%, 63.6%. Again no Sol score.

Humanity's Last Exam
Pass rate, % (higher is better)
NO TOOLS
Fable 5.1
60.9
Fable 5
57.8
Opus 5
56.6
WITH TOOLS
Fable 5.1
65.0
Fable 5
63.8
Opus 5
63.6
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table.

The with-tools cluster is tight: 1.4 points separate all three models. That makes the no-tools gap the informative one. Raw reasoning with no crutch, Fable 5.1 is 3 points clear of both predecessors, and the tools gap narrows rather than widens the story.

6. Business workflows: AutomationBench

AutomationBench is business workflows, and it's the biggest relative jump in the table: 31.4% for Fable 5.1 against 17.1% for Fable 5, nearly double, with Opus 5 at 26.9% and GPT-5.6 Sol at 19.6%.

AutomationBench
Pass rate, % (higher is better)
Fable 5.1
31.4
Opus 5
26.9
GPT-5.6 Sol
19.6
Fable 5
17.1
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table. Fable 5 scored a zero wherever its safeguards intervened on this benchmark.

Two honest footnotes. Fable 5 scored a zero wherever its safeguards intervened on this benchmark, so part of the doubling is subtraction on the old model's side. And 31.4% is still 31.4%: autonomous business workflows remain unsolved for everyone, Opus 5 included.

7. Coding: CursorBench 3.2.0

CursorBench 3.2.0, agentic coding: Fable 5.1 73.4%, Fable 5 70.5%, Opus 5 70.0%, GPT-5.6 Sol 67.2%.

CursorBench 3.2.0
Accuracy, % (higher is better)
Fable 5.1
73.4
Fable 5
70.5
Opus 5
70.0
GPT-5.6 Sol
67.2
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, benchmark table.

SpaceXAI's director of ML Sualeh Asif called it the most capable model they've run on CursorBench 3.2, scoring 73.4% at max effort, and singled out self-verification: it takes difficult coding tasks from start to finish. The 3-point lead over Opus 5 is the smallest gap in the table, which is its own signal. Coding leadership is contested now, and the .1's edge is that it leads while using less of your budget.

The root-cause stories are where the qualitative jump shows. Investment firm Millennium had a crash that hit roughly once in a million runs, unexplained by their engineers and every model they tried, including Fable 5, for four to five years. Fable 5.1 disassembled an external vendor library, matched it against the core dump, and traced the crash to a bug in that library.

The zeros behind the table

Fable 5.1 was evaluated with its production safeguards enabled. Where safeguards intervened, both Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 did on AutomationBench. Elsewhere, safeguarded cybersecurity tasks were completed by Opus 4.8 and biology tasks by Opus 5, which Anthropic flags as likely depressing both models' scores; OfficeChai's read of the table was the same. Treat the affected rows as likely floors.

The safeguards themselves got more precise in the same release. Cyber safeguards now block 60% fewer false positives, and Fable 5.1 can be used to discover software vulnerabilities, though still not to develop exploits. Penetration testing, exploit generation, and binary-based vulnerability scanning still route to Opus models. Biology safeguards fire 85% less often on benign medical and elementary biology questions, with research-grade life sciences work moving to Mythos 5.1 through the new Life Sciences Verification Program, built with the US government. One more compliance note: as a model released after August 2, 2026, Fable 5.1's outputs carry Anthropic's EU AI Act watermark, invisible without their detection API.

Cheaper at the same sticker price

Pricing is unchanged: $10 per million input tokens, $50 per million output tokens. The cut is underneath. Cache reads now cost $0.25 per million tokens, 75% less than Fable 5, and cache reads are most of the bill in context-heavy agentic work.

Indexed cost of Fable usage
Indexed cost (Fable 5 = 100), four weeks of August 2026 usage, default effort
TYPICAL WORKLOAD
Fable 5
100
Fable 5.1
75
HIGHLY AGENTIC WORKLOAD
Fable 5
100
Fable 5.1
55
Source: Introducing Claude Fable 5.1 and Claude Mythos 5.1, indexed cost of Fable usage across Claude Enterprise, Claude Code, and the API.

That works out to roughly 25% cheaper for typical workloads and up to roughly 45% cheaper for highly agentic ones. The market noticed immediately: Walden Yan at Cognition said "with the new cache read pricing a Fable-class model is finally economical for the workloads we'd kept on Opus, starting with code review," and Cognition is moving Devin's Opus 5 traffic to Fable 5.1 on launch day. Effort settings matter too. Defaults are High in Claude Code, Medium in Claude Cowork and on Claude.ai, and Anthropic claims Low and Medium effort already match Fable 5's full-effort results at substantially lower cost.

Science is the real headline

The scientific-research section is where Anthropic spent the announcement, and the results are the kind that don't fit a leaderboard.

Mythos 5.1 designed protein binders against 12 targets using open-source design and folding tools, then two external organizations tested them in the lab. Hit rate: nearly 50%, against a 10-15% norm in protein design today. On three targets drawn from Adaptyv Bio's public design competitions, its designs came in ahead of the best submitted entries.

Fable 5.1 trained the neural network behind a new elevation map of a third of Venus, built from Magellan radar data more than 30 years old. It resolves features down to 2-3 kilometers versus the previous 10-20, shows heights up to 25% more accurately, and is being released under a Creative Commons license ahead of NASA's VERITAS and ESA's EnVision missions.

Mythos 5.1 also wrote custom GPU kernels that sped up seven open-source genomics and protein models by up to 2.5x on an NVIDIA H100, with identical outputs, cutting estimated genome-wide analysis costs by 30-60%. That's normally weeks of performance-engineering work; the model did it in days from public source code alone, and Anthropic plans to open-source the optimizations.

Read the two tiers together and the positioning is clear. Fable 5.1 is the generally available workhorse. Mythos 5.1, through its trusted-access programs and Claude Security, is the research instrument. The science results are how Anthropic is answering the question of what comes after coding benchmarks saturate.

What the early reads say

Every has been testing for about a week and titled their vibe check "Anthropic Is So Back (Again)". Kieran Klaassen's verdict: "this model is Fable for everyone," a Fable you can use in the loop instead of only for long hauls. Katie Parrott: "my trust issues with Claude are starting to heal." Dan Shipper's first message to the team was "I can actually understand what it's saying." On their Slack assistant it used less than half the tokens of Opus 5 with comparable results.

The counter-evidence is real too. Asked for 1,000 words it wrote 1,288. At Extra-high effort it sometimes kept working after being interrupted. On Every's tests the one job where it trailed was the X post, where GPT-5.6 Sol led every Anthropic model. And on Hacker News, where the launch thread passed 400 points in its first hour, the underrated upgrade per Anthropic's own Felix Rieseberg is prose: it "sounds a lot less stereotypically like other Claude models" and follows style instructions more reliably.

Takeaways

  • The .1 label undersells it. Fable 5.1 leads Opus 5 on all seven benchmark rows Anthropic published, and more than doubled Fable 5 on Terminal-Bench-Science (52.6% vs 24.7%).
  • GPT-5.6 Sol appears on five of seven rows and trails on all five. On computer use and reasoning, no Sol score at all.
  • The safeguard zeros mean the OSWorld and AutomationBench numbers likely understate raw capability. The Mythos 5.1 gap on Terminal-Bench 4.0 (60.9% vs 55.8%) is the same artifact, and today's safeguard update is meant to shrink it.
  • The real price cut is cache reads at $0.25 per million tokens: about 25% off typical workloads, up to 45% off agentic ones, sticker unchanged at $10/$50 per million.
  • The strategic story is science: a near-50% protein-binder hit rate, a 2-3 km Venus map, 2.5x GPU kernels. Anthropic is pitching Mythos-class models as research instruments with coding as the entry point.

If you want Fable 5.1's long-horizon behavior pointed at your actual work instead of a benchmark, that's the Vellum thesis: an assistant that holds context across Mac, iOS, web, Slack, and Telegram, running the model you choose through custom LLM credentials.

Hatch your assistant →

Similar Articles

The Personal AI you were promised

GET STARTED