Claude Opus 5.5 Review: Benchmarks & Pricing (2026)
By Loïc Jané · Founder, Fleece AI
Claude Opus 5.5: Anthropic's Current Opus Model Explained
At a Glance: Claude Opus 5.5 is Anthropic's current Opus model, released on September 22, 2026. It costs $4 per million input tokens and $20 per million output tokens, 20% less than Opus 5, and Anthropic says it costs 40% less to run on typical workloads. It has a 1M-token context window and 128K output tokens. Thinking is always on, and the default effort is now medium. On Fleece AI, Opus 5.5 is available on the Business plan.
Claude Opus 5.5 is the first model in Anthropic's Claude 5.5 family. Anthropic positions it as a step up from Opus 5 that "performs at the level of Claude Fable 5.1 on most work" at a lower price. This review covers what changed, which benchmark numbers come from Anthropic and which come from independent leaderboards, what breaks for existing integrations, and how to use the model on Fleece AI.
Table of Contents
- Key Takeaways
- What Is Claude Opus 5.5?
- What Changed Since Opus 5
- Benchmarks: Anthropic's Numbers
- Benchmarks: Independent Leaderboards
- Pricing Compared
- Breaking Changes for Developers
- How to Use Opus 5.5 on Fleece AI
- Limitations
- Frequently Asked Questions
Key Takeaways
- Claude Opus 5.5 was released on September 22, 2026. Anthropic's models overview recommends it as the starting point for most workloads, with Claude Fable 5.1 above it for the most demanding work.
- Pricing is $4 input / $20 output per million tokens, with cache reads at $0.20. Opus 5 costs $5 / $25, with cache reads at $0.50.
- Anthropic reports that Opus 5.5 leads its own benchmark table on agentic coding, computer use and knowledge work, including 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1.
- On independent leaderboards it is listed on APEX-Agents (73.5% at max effort, 52.5% at medium) but not yet on Terminal-Bench 4.0, MCP-Atlas or OSWorld 2.0.
- The default effort is medium, and thinking can no longer be turned off. Forced tool use now returns an error. Check your integration before switching from Opus 5.
- On Fleece AI, Opus 5.5 is a Business-plan model, next to Opus 4.8 and 4.7.
What Is Claude Opus 5.5?
Claude Opus 5.5 is a large language model from Anthropic, built for long-running agentic coding and knowledge work. Its API ID is claude-opus-5-5. It is available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
| Spec | Claude Opus 5.5 |
|---|---|
| Released | September 22, 2026 |
| Context window | 1M tokens |
| Max output | 128K tokens (300K on the Batch API with a beta header) |
| Thinking | Adaptive, always on |
| Default effort | medium |
| Reliable knowledge cutoff | June 2026 |
| Retirement | Not sooner than September 22, 2027 |
Source: Anthropic models overview, retrieved October 6, 2026.
Anthropic says Claude Sonnet 5.5 and Claude Haiku 5.5 would follow in the weeks after launch. Sonnet 5.5 now appears in the models overview at $2 / $10 per million tokens. Haiku 5.5 does not yet.
What Changed Since Opus 5
Lower cost per task
Anthropic cut the per-token price by 20% and the cache-read price by 60%. It also says Opus 5.5 uses fewer tokens per task. Together, Anthropic reports a 40% lower cost than Opus 5 "on typical workloads" at default settings. Cache reads matter most here: Anthropic notes they make up the majority of costs in agentic and coding work.
Faster output
Anthropic reports that Opus 5.5 generates output more than 30% faster than Opus 5. A separate fast mode, available in Claude Code and on the Claude API, runs up to 2.5x faster for $8 / $40 per million tokens.
Medium effort by default
A request that does not set an effort level now runs at medium. On Opus 5 it ran at high. Anthropic's docs also say that at the same effort level, Opus 5.5 tends to think more per turn than Opus 5, most of all at xhigh and max. If you carried an effort setting over from Opus 5, re-test it.
Clearer writing
Anthropic says Opus 5.5 puts the most important information first, uses less jargon and follows writing rules more closely. This addresses common feedback about Opus 5. Several customers quoted in the announcement, including Ramp and Box, reported shorter and easier-to-follow output.
Stricter safeguards
Opus 5.5 is the first Opus model to ship with safeguards similar to Fable 5.1's on cybersecurity, biology and distillation. Most cybersecurity tasks are re-routed to Claude Opus 4.8. Vetted organizations can apply to Anthropic's verification programs for biology and cybersecurity work. On Anthropic's alignment audit, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior.
Benchmarks: Anthropic's Numbers
Anthropic's launch table compares Opus 5.5 with Fable 5.1, Opus 5, GPT-6 Astra and GPT-5.6 Sol.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding) | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1, main (agentic coding) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 (agentic coding) | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (knowledge work, Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench (business workflows) | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam (with tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.1 (computer use, partial) | 81.8% | 80.7% | 74.0% | — | — |
Source: Anthropic, "Introducing Claude Opus 5.5", September 22, 2026, retrieved October 6, 2026. All figures are vendor-reported. Opus 5.5 runs at max effort, except on Terminal-Bench 4.0, where it runs at xhigh. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI. AutomationBench was run by Zapier. The Opus 5.5 AutomationBench score comes from Zapier's early-access evaluation and counts safeguard interventions as failures.
How to read this table. Opus 5.5 leads six of the eight rows. GPT-6 Astra leads AutomationBench and Terminal-Bench-Science. Anthropic itself writes that "benchmark margins have become a less reliable guide to real-world differences" at this level. It also says the real gap with Fable 5.1 is narrower than these scores suggest.
Cost matters as much as score. Anthropic's charts plot score against cost per task. At its default medium effort, Opus 5.5 scores 54.6% on FrontierCode, above GPT-6 Astra's best score of 53.3%, for about a fifth of the cost per task. On CursorBench at medium effort it scores 52.5%, which beats GPT-5.6 Sol's best score of 41.7%.
Benchmarks: Independent Leaderboards
Vendor tables use each vendor's own harness. Independent leaderboards run every model under the same rules. As of October 6, 2026, Opus 5.5 appears on only one of the four leaderboards we checked.
| Leaderboard | Opus 5.5 | Current leader |
|---|---|---|
| APEX-Agents v1.1 (Mercor) | 73.5% ±4.9 (max), 52.5% ±5.6 (medium) | Claude Sonnet 5.5 (max), 75.5% ±4.7 |
| Terminal-Bench 4.0 (tbench.ai) | Not listed | GPT-6 Astra (max), 58.2% ±2.8 |
| MCP-Atlas (Scale AI) | Not listed | Muse Spark 1.1, 88.10 ±1.95 |
| OSWorld 2.0 (xlang.ai) | Not listed | — |
Sources, retrieved October 6, 2026: Mercor APEX-Agents leaderboard, tbench.ai leaderboard, Scale AI MCP-Atlas leaderboard, OSWorld 2.0.
APEX-Agents tests long professional tasks across files, spreadsheets, chat and calendars. By job, Opus 5.5 at max effort scores 80.0% as a management consultant, 71.2% as a corporate lawyer and 69.3% as an investment banking analyst. The gap between max effort (73.5%) and the API default, medium (52.5%), is 21 points. That is the most useful number in this review if you run long workflows: effort is a quality setting, not only a speed setting.
Vendor and leaderboard numbers differ. On Terminal-Bench 4.0, Anthropic's table shows Fable 5.1 at 55.8% and GPT-6 Astra at 57.9%. The public leaderboard shows Fable 5.1 at 57.9% and GPT-6 Astra at 58.2%. Opus 5.5's 66.4% is Anthropic's own run and has a standard error of ±2.6 points. Wait for the leaderboard before treating it as settled.
For what each benchmark measures, see AI Agent Benchmarks 2026 Explained.
Pricing Compared
| Model | Input / output per M tokens | Cache reads per M | Context | Max output |
|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | $0.20 | 1M | 128K |
| Claude Opus 5 | $5 / $25 | $0.50 | 1M | 128K |
| Claude Fable 5.1 | $10 / $50 | $0.25 | 1M | 128K |
| GPT-6 Astra | $10 / $50 | $1.00 (cached input) | 1.05M | 128K |
Sources, retrieved October 6, 2026: Anthropic models overview, Opus 5.5 announcement (Opus 5 prices), OpenAI GPT-6 Astra model page.
Batch API requests are half price for Opus 5.5: $2 / $10. Five-minute cache writes cost $5 per million tokens, and one-hour cache writes cost $8.
Breaking Changes for Developers
Anthropic's What's new in Claude Opus 5.5 page lists four changes that can break code written for Opus 5:
- Thinking can't be disabled. A request with thinking disabled, or with a manual thinking budget, returns a 400 error. Use the effort parameter instead.
- Forced tool use is not supported. tool_choice set to "any" or to a named tool returns a 400 error. Keep "auto" and use strict tool use or structured outputs when you need schema-valid JSON.
- Thinking blocks are tied to the model and the conversation. When you switch models mid-conversation, some thinking blocks are dropped. For API accounts created on or after August 31, 2026, editing earlier context and then replaying a thinking block returns an error by default.
- The older computer use tool (computer_20251124) is rejected on the Claude API and Google Cloud. Move to the newer computer use toolset.
One more change does not fail requests but changes output: short notes written between tool calls now arrive as thinking blocks, which are empty at the default display setting. An app that shows those notes as progress updates will go quiet until it changes that setting.
How to Use Opus 5.5 on Fleece AI
Fleece AI connects models like Opus 5.5 to 3,000+ apps through Pipedream MCP, so the model can read and act in Gmail, Slack, HubSpot, Notion and other tools without you managing tokens or OAuth.
- Subscribe to the Business plan at fleeceai.app/subscribe.
- Open a chat, an agent or a flow.
- Select Claude Opus 5.5 in the model selector.
- In chat, set the reasoning effort from low to max. The default is medium, which matches Anthropic's API default.
Business also includes Claude Opus 4.8 and 4.7 for agents already tuned to them, and GPT-5.6 Sol. Claude Fable 5 is available on Enterprise.
Good fits for Opus 5.5
Long research and reporting. "Every Monday, pull last week's deals from HubSpot and revenue from Stripe, check each figure against the source, and send a summary to the leadership channel in Slack." Anthropic's own test of sourced reporting is relevant here: 16 of 18 Opus 5.5 reports passed a check where any invented figure or quote failed, against none for Fable 5.1 and Opus 5.
Multi-app business workflows. "When a contract is signed in DocuSign, create the customer in the billing system, open an onboarding project in Asana and email the account owner." This is the kind of task AutomationBench measures, where Opus 5.5 scored 40.0% in Zapier's run.
Code review and repository work. "Every day, review open pull requests older than three days on GitHub and post a review comment that lists risks." Anthropic's coding results are its strongest area.
For high-volume, simple steps, a Starter-plan model such as Kimi K3 or GPT-5.2 is usually enough and cheaper in credits. See the model comparison for the full lineup.
Limitations
- Most benchmark claims are vendor-reported. Only APEX-Agents lists Opus 5.5 among the independent leaderboards we checked. Terminal-Bench 4.0, MCP-Atlas and OSWorld 2.0 do not list it yet.
- It is not the top model everywhere. In Anthropic's own table, GPT-6 Astra leads AutomationBench and Terminal-Bench-Science. On APEX-Agents, Claude Sonnet 5.5 at max effort scores higher (75.5% vs 73.5%), within the margin of error.
- Default effort trades quality for cost. On APEX-Agents, medium effort scores 21 points below max.
- Safeguards re-route some work. Most cybersecurity tasks go to Opus 4.8, and some biology work needs a verification program.
- Anthropic flags an evaluation limit. It reports signs that Opus 5.5 often suspects it is being evaluated, which makes it harder to predict behavior in real deployments.
Frequently Asked Questions
Is Claude Opus 5.5 better than Claude Fable 5.1?
On Anthropic's own benchmark table, Opus 5.5 scores higher than Fable 5.1 on all eight benchmarks listed. Anthropic still positions Fable 5.1 above it for the most demanding reasoning and long-horizon work, and says the real gap is narrower than the scores suggest. Opus 5.5 costs less than half as much per token ($4 / $20 vs $10 / $50).
How much does Claude Opus 5.5 cost?
$4 per million input tokens and $20 per million output tokens on the Claude API. Cache reads cost $0.20 per million tokens, and the Batch API is half price. Fast mode costs $8 / $40. On Fleece AI Business (€199 per month, or €159 per month billed yearly), Opus 5.5 is included in your 20,000 monthly credits, with no per-token billing.
What effort level should I use with Opus 5.5?
Start with the default, medium, and raise it for long or high-stakes tasks. On APEX-Agents, Opus 5.5 scores 52.5% at medium and 73.5% at max. Anthropic reports that medium already beats GPT-6 Astra's best score on FrontierCode at about a fifth of the cost.
Will my Opus 5 integration break on Opus 5.5?
It can. Requests that disable thinking, set a manual thinking budget or force a tool call return a 400 error. Thinking blocks are also tied to the model and conversation, and the older computer use tool is rejected on the Claude API and Google Cloud. Read Anthropic's migration guide before switching.
Can I use Claude Opus 5.5 with Gmail, Slack and my other apps?
Yes, through Fleece AI. On its own, Opus 5.5 is a model API without app integrations. On Fleece AI Business it connects to 3,000+ apps with managed OAuth, in chats, agents and scheduled flows.
The Bottom Line
Claude Opus 5.5 lowers the cost of Anthropic's top tier: Opus 5 quality and above, for 20% less per token and, by Anthropic's measure, 40% less per task. Its own benchmark numbers are strong, but only APEX-Agents among the independent leaderboards has confirmed them so far. For Fleece AI Business users running long, multi-app workflows, it is the Anthropic model to try first. Keep effort in mind: the default is medium, and the gap to max effort is large on long tasks.
Related Articles
- Claude Opus 4.7 Review — the April 2026 Opus, still on Fleece AI Business
- AI Agent Benchmarks 2026 Explained — what each benchmark measures and who leads
- Best AI Models for Workflow Automation 2026 — full comparison
- Best AI Model for Tool Calling 2026 — which model calls APIs most reliably
- Long-Context AI Agents Explained — what 1M tokens buys you
- Claude vs Fleece: Assistant vs Agent — Claude standalone vs Claude powering Fleece AI
Start on Fleece AI Business — Claude Opus 5.5, GPT-5.6 Sol, 3,000+ integrations, 20,000 credits/month.