# Claude Opus 5.5 Benchmarks Explained

*By Nicolas Zeeb · September 22, 2026 · 11 min · Model Comparisons*

A full breakdown of Anthropic's Claude Opus 5.5 benchmarks.

## Claude Opus 5.5 Benchmarks Explained

Anthropic released [Claude Opus 5.5](https://www.anthropic.com/claude-opus-5-5) on September 22, 2026, less than two months after Opus 5. The headline pitch is simple: it delivers Claude Fable 5.1-class performance while cutting execution costs by roughly 40% compared to Opus 5. Pricing drops 20% on raw tokens to $4 per million input tokens and $20 per million output tokens, while prompt cache reads drop 60% down to $0.20 per million.

The release is already reshaping the frontier benchmark landscape. According to [Artificial Analysis](https://x.com/ArtificialAnlys/status/2102438210798514391), Opus 5.5 took the top spot on the Artificial Analysis Intelligence Index with an index score of 58 at max effort, leading on six of the ten core evaluations, including Humanity's Last Exam (61.4%), SciCode (66.9%), and reaching an Elo of 1822 on the AA-Briefcase knowledge work evaluation. Across four of its five effort settings, Opus 5.5 sits directly on the intelligence-versus-cost Pareto frontier.

Crucially, Anthropic shaped this release around developer usability and token efficiency. As Claude Code lead [Thariq](https://x.com/trq212/status/2102437686967738431) highlighted at launch, Opus 5.5 was built directly in response to user feedback, communicating more clearly and cutting conversational drag across every effort level, alongside a 20% increase to 5-hour rate limits and a banked reset for subscriber plans.

Anthropic product lead [Cat Wu](https://x.com/_catwu/status/2102437713781944397) confirmed that Opus 5.5 is now the default model across Claude Code, the Claude app, and Cowork for Pro, Max, and Team users. Anthropic set the default effort level to medium across products, matching Fable 5.1 on intelligence while delivering faster execution, with the reduced token pricing allowing rate limits to stretch 25% further.

Here is the full walk through the benchmark categories Anthropic published, what the numbers reveal, and why the efficiency shift matters more than raw percentage gains.

## 1. Terminal execution: Terminal-Bench 4.0

Terminal-Bench 4.0 measures how effectively an autonomous agent executes complex, multi-step engineering tasks inside a real command-line environment.

Opus 5.5 scores **66.4%** at extra-high effort (standard error ±2.6 points). That places it comfortably ahead of OpenAI's GPT-6 Astra at **57.9%** (high effort), Anthropic's own Fable 5.1 at **55.8%**, Opus 5 at **52.3%**, and GPT-5.6 Sol at **37.3%**.

Terminal-Bench 4.0

| Label | Value |
| --- | --- |
| Opus 5.5 | 66.4 |
| GPT-6 Astra | 57.9 |
| Fable 5.1 | 55.8 |
| Opus 5 | 52.3 |
| GPT-5.6 Sol | 37.3 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. Standard error is ±2.6 pts for Opus 5.5 and ±1.6 to 2.0 pts for other Claude models. Scores reflect each model's top effort setting.

What matters beyond the 8.5-point lead over GPT-6 Astra is the cost curve. Anthropic reports that Opus 5.5 running at default effort beats Opus 5 at maximum effort for approximately one-fifth of the cost, while matching Astra's performance at roughly 40% of the cost. In real command-line workflows, Deepak Singh, VP of Agentic AI at [Kiro](https://kiro.dev), notes that Opus 5.5 solved more tasks than Opus 5 while making 40% fewer calls and consuming half the tokens.

## 2. Software engineering: FrontierCode v1.1

FrontierCode v1.1 evaluates whether an agent's code pull requests would actually be accepted and merged into production software projects.

Opus 5.5 reaches **54.4%** on the main evaluation set at maximum effort, compared to **53.3%** for [GPT-6 Astra](https://www.vellum.ai/blog/gpt-6-astra-benchmarks-explained), **50.3%** for [Fable 5.1](https://www.vellum.ai/blog/claude-fable-5-1-mythos-5-1-benchmarks-explained), **48.0%** for [Opus 5](https://www.vellum.ai/blog/claude-opus-5-benchmarks-explained), and **47.5%** for GPT-5.6 Sol.

FrontierCode v1.1 (Main Set)

| Label | Value |
| --- | --- |
| Opus 5.5 | 54.4 |
| GPT-6 Astra | 53.3 |
| Fable 5.1 | 50.3 |
| Opus 5 | 48.0 |
| GPT-5.6 Sol | 47.5 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table.

Even more striking is default effort: at standard medium effort, Opus 5.5 scores 54.6%, outpacing GPT-6 Astra's peak score (53.3%) while costing roughly 80% less per task attempt. In enterprise codebases, early testers report large reductions in developer friction. Mario Rodriguez, Chief Product Officer at [GitHub](https://github.com), noted that across GitHub Copilot CLI and VS Code, Opus 5.5 solved terminal tasks in less than half the steps of Opus 5 while consuming fewer tokens.

## 3. Real-world coding sessions: CursorBench 4.0

CursorBench 4.0 tests coding agents on ambiguous, multi-file code editing tasks drawn from real production sessions in the Cursor editor.

Opus 5.5 scores **57.8%** at peak effort, leading Fable 5.1 at **51.8%**, Opus 5 at **46.6%**, and GPT-5.6 Sol at **41.7%** (OpenAI did not publish a CursorBench 4.0 score for GPT-6 Astra).

CursorBench 4.0

| Label | Value |
| --- | --- |
| Opus 5.5 | 57.8 |
| Fable 5.1 | 51.8 |
| Opus 5 | 46.6 |
| GPT-5.6 Sol | 41.7 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. GPT-6 Astra score was not reported.

At default medium effort, Opus 5.5 records 52.5%, remaining ahead of Fable 5.1's top score (51.8%) and clearing GPT-5.6 Sol by nearly 11 percentage points for about one-third of the cost. Across six private repositories, Sean Heintz, staff developer at [Clio](https://clio.com), ran Opus 5.5 unattended for over 18 hours defining inter-service contracts, finding that it hit milestones faster with minimal rework and concise code commentary.

## 4. Professional knowledge work: GDPval-AA v2.1

Artificial Analysis's GDPval-AA v2.1 evaluates AI agents on professional knowledge-work tasks across 44 occupations, graded on an Elo scale.

Opus 5.5 sets a new high mark at **1846 Elo**, beating Fable 5.1 (**1735**), Opus 5 (**1708**), GPT-5.6 Sol (**1588**), and GPT-6 Astra (**1542**).

GDPval-AA v2.1

| Label | Value |
| --- | --- |
| Opus 5.5 | 1846 |
| Fable 5.1 | 1735 |
| Opus 5 | 1708 |
| GPT-5.6 Sol | 1588 |
| GPT-6 Astra | 1542 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. Evaluated across 44 occupations.

This is a notable divergence from OpenAI's flagship: Astra dropped to 1542 on this index, whereas Opus 5.5 extends Anthropic's lead by over 300 Elo points. At default effort, Opus 5.5 tops Astra at maximum effort for roughly one-fifth of the price per task. Financial engineering teams cite qualitative jumps as well: Aabhas Sharma, CTO at [Hebbia](https://hebbia.ai), reported that on end-to-end finance workflows graded against expert rubrics, Opus 5.5 covered 86.6% of target criteria compared to 60.3% for Opus 5, alongside superior citation recall.

## 5. Business workflows: AutomationBench

AutomationBench, designed by [Zapier](https://zapier.com), measures whether an AI agent can execute multi-step workflows across disparate SaaS platforms without breaking logic.

GPT-6 Astra maintains a narrow edge here at **41.4%**, closely followed by Opus 5.5 at **40.0%**. Both models substantially outpace Fable 5.1 at **31.4%**, GPT-5.6 Sol at **28.8%**, and Opus 5 at **26.9%**.

AutomationBench

| Label | Value |
| --- | --- |
| GPT-6 Astra | 41.4 |
| Opus 5.5 | 40.0 |
| Fable 5.1 | 31.4 |
| GPT-5.6 Sol | 28.8 |
| Opus 5 | 26.9 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. Zapier evaluation data. Runs were performed without fallback models, meaning safeguard interventions registered as failures.

Methodology is crucial here. Anthropic notes that Zapier evaluated Opus 5.5 without fallback models enabled; whenever security safeguards intervened on complex workflow tasks, the task was counted as a failure. Even with those artificial zeroes, Opus 5.5 nearly matches Astra while delivering a 13-point jump over Opus 5.

## 6. Multidisciplinary reasoning: Humanity's Last Exam

Humanity's Last Exam (HLE) tests multi-step, expert-level academic and professional reasoning across subjects designed to resist retrieval shortcuts.

With tools enabled, Opus 5.5 tops the board at **67.7%**, advancing past Fable 5.1 (**65.6%**), Opus 5 (**63.6%**), and GPT-6 Astra (**57.2%**).

Humanity's Last Exam (With Tools)

| Label | Value |
| --- | --- |
| Opus 5.5 | 67.7 |
| Fable 5.1 | 65.6 |
| Opus 5 | 63.6 |
| GPT-6 Astra | 57.2 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. All models evaluated with tools at maximum effort.

This 10.5-point gap over GPT-6 Astra highlights a continuing divide in tool-assisted multidisciplinary reasoning. Rather than hallucinating intermediate findings when queries prove difficult, Opus 5.5 exhibits higher stamina in recursive verification loops. In quant testing at Walleye Capital, Frank Corrao noted that Opus 5.5 detected an indexing off-by-one error inside their own evaluation instructions and adjusted for it, noting that doing so might cost it automated points.

## 7. Scientific inquiry: Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 tests whether an agent can formulate hypotheses, write analysis code, execute terminal commands, and reach sound scientific conclusions unattended.

OpenAI's GPT-6 Astra leads this benchmark at **64.6%**, followed by Opus 5.5 at **58.7%**, Fable 5.1 at **52.6%**, Opus 5 at **29.0%**, and GPT-5.6 Sol at **22.4%** (standard error ±3.5 to 5.0 points per model).

Terminal-Bench-Science 0.1

| Label | Value |
| --- | --- |
| GPT-6 Astra | 64.6 |
| Opus 5.5 | 58.7 |
| Fable 5.1 | 52.6 |
| Opus 5 | 29.0 |
| GPT-5.6 Sol | 22.4 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table. Standard error is ±3.5 to 5.0 pts per model.

While Astra takes the top slot, Opus 5.5 more than doubles Opus 5 (58.7% vs 29.0%) and gains 6 points over Fable 5.1. Like Fable 5.1, Opus 5.5 was evaluated with production biology and cyber safeguards active; tasks triggering biology flags fell back to Opus 5, reducing its composite score.

## 8. Computer use and visual perception: OSWorld 2.0 and Chartography

On operating system automation, OSWorld 2.0 (partial pass) places Opus 5.5 at **81.8%**, slightly ahead of Fable 5.1 (**80.7%**) and Opus 5 (**74.0%**). On Chartography (visual chart comprehension with tools), Opus 5.5 posts **89.0%** compared to **88.4%** for Fable 5.1 and **83.4%** for Opus 5.

OSWorld 2.0 (Partial Pass)

| Label | Value |
| --- | --- |
| Opus 5.5 | 81.8 |
| Fable 5.1 | 80.7 |
| Opus 5 | 74.0 |

Source: Anthropic, Introducing Claude Opus 5.5, benchmark table.

In practical UI testing, when asked to optimize load times across every page of a complex web app, Anthropic reported Opus 5.5 succeeded in 39 out of 40 attempts, whereas Opus 5 made smaller changes that inadvertently broke application behavior. In graphic rendering tests, evaluators reported noticeably crisper UI and styling polish when generating complete interactive apps from a single prompt.

## Alignment and containment boundaries

Capability gains at this tier introduce severe containment risks. Anthropic addressed this directly by pairing Opus 5.5 with its [automated behavioral audit](https://www.anthropic.com/claude-opus-5-5), an evaluation suite encompassing roughly 2,000 simulated operational scenarios.

Across that suite, Opus 5.5 recorded the lowest misaligned-behavior rate of any recent Claude model. In a new evaluation specifically testing an agent's tendency to circumvent sandbox restrictions, Opus 5.5 attempted boundary circumvention 85% less often than Opus 5 or Mythos 5.1. Crucially, every boundary attempt it did make was low severity and self-reported by the model during execution.

On prompt injection, Opus 5.5 matched or exceeded Opus 5 across tool use, coding, and web browsing. In independent testing conducted by AI security firm Gray Swan, Opus 5.5 tied Fable 5.1 for the lowest prompt injection success rate recorded.

Like Fable 5.1, Opus 5.5 ships with preserved thinking enabled for API accounts, preventing malicious actors from altering prior context turns during distillation attacks. It also complies with EU AI Act watermarking standards and no longer allows reasoning mode to be disabled.

## Token efficiency and pricing

The practical reality of running frontier agents is that token consumption dictates feasibility. Opus 5.5 adjusts both the per-token rate and the token volume needed per job:

Claude API Pricing Comparison

| Label | Value |
| --- | --- |
| Opus 5 | $5.00 |
| Opus 5.5 | $4.00 |
| Opus 5 | $25.00 |
| Opus 5.5 | $20.00 |
| Opus 5 | $0.50 |
| Opus 5.5 | $0.20 |

Source: Anthropic, Introducing Claude Opus 5.5 pricing documentation.

Cache reads drive the bulk of operational expense for autonomous coding and long-horizon desktop tasks. Slashing cache reads from $0.50 down to $0.20 per million tokens represents a 60% direct reduction. Combined with faster execution (over 30% speedup in token generation) and tighter, less verbose outputs, effective cost drops roughly 40% across standard enterprise workloads.

For subscription users on Pro, Max, and Team plans, Anthropic expanded 5-hour rate limits by 20% and introduced a user-controlled rate limit reset feature that can be triggered on demand.

## Tone, verbosity, and enterprise feedback

A frequent complaint with Opus 5 was conversational drag: excessive preambles, defensive explanations, and repetitive commentary.

Opus 5.5 re-engineers conversational style toward directness. In enterprise evaluations at [Box](https://www.box.com), Yashodha Bhavnani, VP of AI Products, reported that Opus 5.5 used one-third of the tokens Opus 5 required, with outputs 40% less verbose without sacrificing factual precision. John Ruelas at [Ramp](https://ramp.com) noted that design specifications generated by the model required minimal editorial touch, writing "like a good colleague" rather than an over-eager assistant.

In large code refactors, [Stripe](https://stripe.com) staff software engineer Cristian Rivera reported using Opus 5.5 across a multi-day rebase of 40 stacked pull requests, where the model mapped conflicts clearly and all 40 passed CI. At [Optiver](https://optiver.com), Noyan Tokgozoglu observed that Opus 5.5 matched Opus 5 quality in half the turns and output tokens, reducing trading support workload costs by 40% to 50%.

## Takeaways

- Opus 5.5 resets the frontier coding benchmark: 66.4% on Terminal-Bench 4.0 (beating GPT-6 Astra's 57.9%), 54.4% on FrontierCode v1.1, and 57.8% on CursorBench 4.0.
- Knowledge work maintains a decisive lead: 1846 Elo on GDPval-AA v2.1, clearing Astra by over 300 points and setting a new benchmark standard.
- The 40% cost reduction is structural: token rates drop 20% to $4/$20 per million, prompt cache reads drop 60% to $0.20 per million, and raw token usage per task decreases by 25% to 40%.
- Safety practices match Fable 5.1: automated behavioral audit scores are the highest Anthropic has recorded, sandbox circumvention attempts dropped 85%, and prompt injection defenses tie Fable 5.1.
- Output style is significantly more concise: early enterprise users report 40% fewer output tokens, fewer intermediate apologies, and cleaner explanations that do not require aggressive prompt conditioning.

When frontier capability becomes this fast and affordable, the advantage belongs to systems that can route tasks continuously across the surfaces where work actually happens. If you want to deploy Claude Opus 5.5 across your daily operations, that is what Vellum provides: an assistant that integrates frontier models into your real workflow across Mac, iOS, Android, web app, voice, email, Telegram, Slack, and terminal, running the models you choose through custom LLM credentials.

[Hatch your assistant →](https://vellum.ai)
