Claude Sonnet 5.5 Benchmarks Explained
Anthropic released Claude Sonnet 5.5 on September 28, 2026, arriving just six days after the launch of Claude Opus 5.5. While Opus 5.5 targeted open-ended, complex orchestration, Sonnet 5.5 is designed as the high-speed workhorse: running more than 30% faster than Sonnet 5 and reducing effective costs per task by up to 30%.
The headline development is not subtle. On command-line agentic coding, Sonnet 5.5 jumps from Sonnet 5's 10.3% to 70.6% on Terminal-Bench 4.0, outscoring even Opus 5.5's top mark. On professional knowledge work, it scores 1844 Elo on GDPval-AA v2.1, sitting two points shy of Opus 5.5 and leaving OpenAI's GPT-6 Sol nearly 360 points behind.
List pricing holds steady at $2 per million input tokens, $10 per million output tokens, and $0.20 per million prompt cache read tokens. However, because Sonnet 5.5 resolves problems in significantly fewer steps and batches tool executions more aggressively, the actual cost per task drops up to 30%.
Here is the full walk through the benchmark categories Anthropic published, what the numbers reveal across coding, knowledge work, and system control, and where Sonnet 5.5 fits alongside Opus 5.5 and GPT-6 Sol.
1. Terminal execution: Terminal-Bench 4.0
Terminal-Bench 4.0 measures how effectively an autonomous agent executes complex, multi-step engineering tasks inside a live command-line environment.
Sonnet 5.5 scores 70.6% at peak effort. That represents a dramatic leap from Sonnet 5's 10.3%, while edging past Opus 5.5 at 66.4% (extra-high effort).
</div>
</div>
</div>
</div>
</div>
</div>
</div>
What matters as much as the headline percentage is the efficiency curve. At medium effort (the default setting in the Claude apps and Claude Code), Sonnet 5.5 exceeds Sonnet 5's best score for less than a tenth of the cost per task. Sonnet 5.5 avoids unnecessary subagent spawning and executes terminal commands directly, keeping task loops tight.
2. Software engineering: FrontierCode 1.1
FrontierCode 1.1 evaluates whether an agent's proposed pull requests would be merged into production codebases without human intervention.
Sonnet 5.5 achieves 52.1% at extra-high effort and 46.2% at max effort. This places it ahead of OpenAI's GPT-6 Sol at 49.3% and Sonnet 5 at 42.4%, closely trailing Opus 5.5 at 54.4%.
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
Anthropic noted an interesting dynamic with effort levels: at max effort, Sonnet 5.5 scored 46.2%, which is lower than at extra-high effort (52.1%). FrontierCode penalizes out-of-scope edits, even when helpful. At maximum effort, Sonnet 5.5 frequently invoked the code-review skill to dispatch tasks across multiple subagents, which in several cases triggered timeouts or produced peripheral edits beyond the initial task boundaries. At high effort (the platform default), Sonnet 5.5 matches GPT-6 Sol's best score for approximately one-fifth of the cost per task.
3. Real-world coding sessions: CursorBench 4.0
CursorBench 4.0 evaluates coding models across multi-file engineering challenges drawn directly from real user sessions inside Cursor.
Sonnet 5.5 records 55.5%, closing within two points of Opus 5.5 (57.8%) and clearing Sonnet 5 (34.1%) by over 21 percentage points.
</div>
</div>
</div>
</div>
</div>
</div>
</div>
At low effort, Sonnet 5.5 already exceeds Sonnet 5's top score while running at less than a tenth of the cost per attempt. Early testing at SpaceXAI by ML Director Sualeh Asif highlighted Sonnet 5.5 delivering near-Opus capabilities on CursorBench at an attractive cost profile for high-frequency developer workflows.
4. Professional knowledge work: GDPval-AA v2.1
GDPval-AA v2.1, developed by Artificial Analysis, tests models on demanding professional tasks across 44 occupations and nine industries.
Sonnet 5.5 reaches 1844 Elo, performing virtually on par with Opus 5.5 (1846 Elo) and well ahead of both GPT-6 Sol (1487 Elo) and Sonnet 5 (1449 Elo).
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
This 395-point jump over Sonnet 5 is the largest single-generation knowledge-work gain Anthropic has reported. The model excels at synthesis, document design, and formatting precision. In internal evaluations, Anthropic provided Sonnet 5.5 with quarterly earnings materials, call transcripts, and a template; the model generated a 10-slide operating deck that two independent reviewers judged ready for executive distribution without revisions.
5. Long-horizon workplace tasks: AA-Briefcase v1.1
AA-Briefcase v1.1 tests sustained, multi-hour knowledge work requiring continuous context retention, data aggregation, and synthesis across complex projects.
Sonnet 5.5 posts 1811 Elo, coming within 11 points of Opus 5.5 (1822 Elo) while outpacing GPT-6 Sol (1483 Elo) and Sonnet 5 (1359 Elo).
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
On AA-Briefcase, running Sonnet 5.5 at medium effort surpasses Sonnet 5's top score at roughly one-ninth of the cost per task. In financial analysis testing at Balyasny Asset Management, Senior AI Engineer Joe Poirier reported running 2,441 finance tasks covering extraction, analysis, and forecasting; Sonnet 5.5 outperformed Sonnet 5 while using approximately 121k tokens per answer compared to 497k tokens for Sonnet 5.
6. Multidisciplinary reasoning: Humanity's Last Exam
Humanity's Last Exam evaluates difficult multi-step reasoning across specialized academic and technical subjects designed to prevent memorization shortcuts.
With tools enabled, Sonnet 5.5 scores 64.5%, improving over Sonnet 5 (54.9%) by nearly 10 percentage points and approaching Opus 5.5 (67.7%).
</div>
</div>
</div>
</div>
</div>
</div>
</div>
This gain reflects better tool coordination. Where Sonnet 5 tended to loop excessively through search queries when an initial lookup failed, Sonnet 5.5 synthesizes retrieved information quickly and pivots strategy when a data path is blocked.
7. Computer use and visual perception: OSWorld 2.1 and Chartography
On operating system automation, OSWorld 2.1 (partial score) places Sonnet 5.5 at 80.1%, nearly matching Opus 5.5 (81.8%) and substantially ahead of Sonnet 5 (57.0%). On Chartography (evaluating visual chart understanding without tools), Sonnet 5.5 jumps to 61.6%, more than quadrupling Sonnet 5 (15.6%) and beating GPT-6 Sol (53.6%).
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
This visual capability translates directly to real-world interface navigation. In gaming tests, Anthropic confirmed Sonnet 5.5 is the first Sonnet model capable of completing Pokémon Red guided entirely by screenshots. In multi-step OS workflows, the model parses cluttered menus and complex dashboard visuals without relying on external OCR steps.
Alignment, cyber safeguards, and anti-distillation
Because Sonnet 5.5 reaches capability levels historically reserved for Opus-class models, Anthropic deployed it with enhanced safety controls.
On Anthropic's automated behavioral audit across roughly 1,850 scenarios, Sonnet 5.5 improves upon or matches Sonnet 5 across measures of alignment, resistance to misuse, and honesty. On containment tests, Sonnet 5.5 proved to be the least likely model in the Claude lineup to probe the boundaries of its execution environment.
Crucially, Sonnet 5.5 is the first Sonnet model to launch with dedicated cybersecurity safeguards and fallbacks. High-risk cybersecurity requests visibly fall back to Sonnet 5, while routine development and vulnerability patching proceed unaffected. Qualified researchers can apply to Anthropic's Cyber Verification Program for elevated access.
To prevent industrial-scale capability distillation via automated account pooling, Sonnet 5.5 incorporates anti-distillation classifiers and expands preserved thinking. A session's internal reasoning tokens remain cryptographically tied to the organization that initiated them, preventing attackers from harvesting reasoning traces across accounts.
Pricing and operational efficiency
The economic reality of running production agents centers on token volume per solved task. Sonnet 5.5 keeps the nominal token rates of Sonnet 5 while cutting the volume of tokens needed to finish real work:
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
</div>
Beyond token pricing, execution latency is more than 30% faster than Sonnet 5. In real-world software workflows, developers see fewer intermediate thinking pauses and faster code generation, making iterative loops significantly tighter.
Enterprise adoption and builder feedback
Early enterprise deployments highlight reduced iteration cycles and cleaner tool orchestration:
- Atlassian: Jamil Valliani, Head of Product for AI at Atlassian, reported that teams running Rovo Agents executed actions up to 30% faster with Sonnet 5.5 compared to Sonnet 5.
- CodeRabbit: David Loker, VP of AI at CodeRabbit, observed that Sonnet 5.5 eliminated Sonnet 5's tendency to trigger redundant web searches, cutting total token consumption while improving code review judgment.
- Epic Games: Daniel Vogel, Chief Operating Officer at Epic Games, noted that Sonnet 5.5 held up across gameplay system architecture reviews and system design audits, handling tens of thousands of lines of code with minimal prompting.
- Base44: Gabriel Grinberg, AI Engineering Lead at Base44, evaluated the model across 118 app builds. Sonnet 5.5 matched Opus 5 quality in 3.6 iterations per build on average (where Opus 5 took 7.7), recording the fewest failed tool calls of any model tested.
- Lovable: Fabian Hedin, Co-founder and CTO at Lovable, found that coding evaluations required one-third fewer tool calls and half as many shell executions to reach completion.
- Zendesk: Abhinay Kathuria, Director of AI at Zendesk, reported that support tickets were processed 20% faster, making fewer incorrect escalation decisions.
Takeaways
- Sonnet 5.5 dominates terminal execution: 70.6% on Terminal-Bench 4.0, climbing from Sonnet 5's 10.3% and outscoring Opus 5.5's 66.4%.
- Knowledge work reaches Opus parity: 1844 Elo on GDPval-AA v2.1 and 1811 Elo on AA-Briefcase v1.1, trailing Opus 5.5 by only 2 to 11 points while outpacing GPT-6 Sol by more than 300 points.
- The 30% cost reduction is driven by efficiency: list token prices stay at $2/$10 per million ($0.20 cache reads), but fewer reasoning turns and tighter execution reduce cost per task up to 30%.
- Generation speed is over 30% faster: developers experience snappier interactions across IDEs, terminal tools, and chat surfaces.
- First Sonnet model with tier-one safeguards: incorporates cyber fallbacks, anti-distillation preserved thinking, and the lowest sandbox circumvention rate in Anthropic's test suite.
When high-intelligence models become this fast and cost-effective, productivity gains depend on having an assistant capable of taking autonomous action across your entire work stack. If you want to put Claude Sonnet 5.5 to work across your daily operations, that is what Vellum provides: an assistant that integrates frontier models into your real workflow across Mac, iOS, Android, web app, voice, email, Telegram, Slack, and terminal, running the models you choose through custom LLM credentials.



