Claude Opus 5.5 is Anthropic's new flagship-tier model, released on September 22, 2026, at $4 per million input tokens and $20 per million output tokens — 20% below Opus 5's list price. It scores 66.4% on Terminal-Bench 4.0 against Opus 5's 52.3%, beats the more expensive Claude Fable 5.1 on most agentic coding benchmarks, and Anthropic says typical workloads cost about 40% less overall because the model finishes tasks in fewer tokens.
That last clause is the interesting one. Most model launches are a capability story with a price footnote. This one is the reverse: Opus 5.5 is an efficiency pass that happens to set new highs. It landed 90 minutes before OpenAI shipped GPT-6 Sol and Luna, which tells you how tight the release calendar has become.
This article covers what Opus 5.5 actually scores, what it really costs once you account for token efficiency, the breaking API changes that will bite existing integrations, and whether you should move production traffic to it.
Key Takeaways
- Claude Opus 5.5 launched September 22, 2026 at $4/$20 per million tokens, with cache reads cut 60% to $0.20 per million.
- It scores 66.4% on Terminal-Bench 4.0, ahead of Fable 5.1 (55.8%), GPT-6 Astra (57.9%) and Opus 5 (52.3%).
- Anthropic reports about 40% lower cost on typical workloads versus Opus 5, driven by fewer tokens per task rather than the list-price cut alone.
- Thinking can no longer be disabled: adaptive thinking is always on and controlled through an
effortsetting with five levels.- Cybersecurity work is routed to Opus 4.8 by default, so security teams need the Cyber Verification Program to get Opus 5.5 on those tasks.
What is Claude Opus 5.5?
Claude Opus 5.5 is a mid-cycle upgrade to Opus 5 that Anthropic positions as its new leading model for coding and knowledge work. It has a 1 million-token context window, up to 128,000 synchronous output tokens, and is available on the Claude Platform, AWS, Google Cloud and Microsoft Azure under the model ID claude-opus-5-5.
The clearest way to read the release is as a re-balancing of Anthropic's lineup. Before September 22, the hierarchy was simple: Fable 5.1 at $10/$50 was the top model, Opus 5 at $5/$25 sat underneath. Opus 5.5 breaks that ordering. According to Anthropic's launch post, it performs at the level of Fable 5.1 on most work while costing less than Opus 5.
So the "5.5" is slightly misleading. This is not a point release that nudges a few scores. It is a cheaper model that outperforms the model priced 2.5 times higher on the benchmarks developers care most about.
Three specifics are worth knowing up front:
- Speed. Output generation is more than 30% faster than Opus 5. A Fast mode at $8/$40 pushes that to roughly 2.5x.
- Output length. 128,000 tokens synchronously, and 300,000 through the Message Batches API in beta.
- Effort levels. Low, medium (the default), high, xhigh and max.
If you read our earlier breakdown of Claude Opus 5, the architecture story has not changed. What changed is how much work the model gets done per token.
Claude Opus 5.5 benchmarks: how much better is it?
On the agentic coding benchmarks Anthropic published, Opus 5.5 leads every model it was compared against, including GPT-6 Astra. The margin is largest on Terminal-Bench 4.0, where it scores 66.4% — 14.1 points above Opus 5 and 8.5 points above Astra, a model that costs $10/$50 per million tokens.
Here are the published numbers side by side.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam (tools) | 67.7% | 65.6% | 63.6% | 57.2% | — |
| OSWorld 2.1 | 81.8% | 80.7% | 74.0% | — | — |
Two things stand out.
First, the gains are not uniform. Terminal-Bench-Science 0.1 roughly doubled, from 29.0% on Opus 5 to 58.7%. OSWorld 2.1, the computer-use benchmark, moved only about a point above Fable 5.1. Opus 5.5 is a big step for long-running terminal and scientific-computing work and a small one for clicking around a desktop.
Second, Astra still wins AutomationBench, 41.4% to 40.0%. It is a narrow lead, but it is a reminder that "leads every benchmark" is true of the coding table and not of the whole sheet.
A caveat that applies to every vendor table: these are Anthropic's numbers, run on Anthropic's harnesses, at effort settings Anthropic chose. Pre-release testing was also done externally by METR, but the comparison figures above come from the launch material. Treat them as a strong prior, then run your own evals.
What the customer numbers say
Vendor benchmarks are one signal. The customer quotes in the launch post are more useful because they describe the shape of the improvement:
- GitHub's chief product officer said Opus 5.5 solved more terminal tasks than Opus 5 "in less than half the steps."
- Optiver reported matching Opus 5's quality "in about half the turns, time and output tokens," cutting cost by 40 to 50%.
- Deloitte said it caught 72% of known bugs in a review task, against 56% for Opus 5 at high effort.
- Box reported it used a third of the tokens and produced answers that were 40% less verbose.
The pattern is consistent: fewer steps, fewer tokens, shorter output. That is a different kind of upgrade from "scores two points higher," and it is the one that shows up on your invoice.
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 and cache writes at $5 per million. Fast mode doubles the rate to $8/$40. Against Opus 5, that is a 20% list-price cut and a 60% cut on cache reads.
| Price per 1M tokens | Opus 5.5 | Opus 5 | Change |
|---|---|---|---|
| Input | $4.00 | $5.00 | −20% |
| Output | $20.00 | $25.00 | −20% |
| Cache read | $0.20 | $0.50 | −60% |
| Cache write | $5.00 | $6.25 | −20% |
| Fast mode (in/out) | $8 / $40 | n/a | up to 2.5x speed |
The list price is the smaller half of the story. A 20% price cut combined with a model that uses meaningfully fewer tokens per task is how Anthropic gets to "about 40% less overall." VentureBeat's coverage cites a HAProxy C-to-Rust translation that finished in 9.5 hours instead of 12 and cost 51% less.
The cache-read cut deserves more attention than it gets. Agent loops re-send the same system prompt, tool definitions and file context on every turn. For a coding agent, cached input is often the majority of tokens billed. A 60% reduction there compounds with the token-efficiency gain.
You can sanity-check your own bill with a few lines of Python:
PRICES = { # USD per 1M tokens
"opus-5": {"in": 5.00, "out": 25.00, "cache_read": 0.50},
"opus-5.5": {"in": 4.00, "out": 20.00, "cache_read": 0.20},
}
def cost(model, fresh_in, cached_in, out):
p = PRICES[model]
return (fresh_in * p["in"] + cached_in * p["cache_read"] + out * p["out"]) / 1_000_000
# An agent session: 2M fresh input, 30M cached input, 1.5M output
old = cost("opus-5", 2_000_000, 30_000_000, 1_500_000)
# Same task on 5.5, assuming ~40% fewer turns (your mileage will vary)
new = cost("opus-5.5", 1_200_000, 18_000_000, 900_000)
print(f"Opus 5: ${old:.2f}") # $62.50
print(f"Opus 5.5: ${new:.2f}") # $26.40
The 40%-fewer-turns assumption is illustrative, not a promise. Swap in the token counts from your own traces. The point is that cache pricing and step count move the total far more than the headline input rate.
Where does that leave Opus 5.5 against the competition? GPT-6 Sol, released the same day, is $2/$10 — exactly half. Our LLM API pricing guide covers the broader table, but the short version is that Opus 5.5 is no longer the premium-priced option in Anthropic's own lineup, and it is still twice the price of OpenAI's workhorse.
What breaks when you upgrade
This is the section most launch coverage skips, and it is the one that determines whether your migration takes an afternoon or a week. Opus 5.5 ships with breaking changes:
- Thinking cannot be disabled. Adaptive thinking is always on. You control depth with the
effortparameter instead of toggling thinking off. Any integration that explicitly disabled thinking to save latency needs to move toeffort: low. - Preserved thinking is enforced for API accounts created after August 31, 2026. Thinking blocks are tied to the model and the conversation, which means you can no longer freely edit prior context and replay it. Anthropic describes this as an anti-distillation measure.
- Forced tool use can return errors in cases where it previously would not.
- The older
computer_20251124tool is rejected. Computer-use integrations must move to the current tool version.
The preserved-thinking change is the one to test first. Frameworks that compact or rewrite conversation history — summarizing old turns, stripping tool results, truncating context — are doing exactly the kind of context editing this restricts. If your agent harness does history surgery, run it against Opus 5.5 in staging before flipping the default.
Safety: the 85% number, and the cyber routing
Anthropic's headline safety claim is that Opus 5.5 attempts to get around containment boundaries about 85% less often than Opus 5. It also reports the best result of any recent Claude model across nearly 2,000 behavioral scenarios, and says the model is much less likely to take hard-to-reverse actions.
For anyone running agents with real permissions, that matters more than a benchmark point. The practical failure mode of an agent is not a wrong answer; it is a confident rm -rf, a force-push, or a workaround that quietly bypasses the guard you set up. We covered how often that goes wrong in AI coding agent security. An 85% reduction is not zero, and sandboxing is still mandatory, but it is the right direction.
On prompt injection, Opus 5.5 matches or beats Opus 5 in every tested setting and ties Fable 5.1 for the lowest attack success rate on the Gray Swan benchmark.
There is one restriction that will surprise security teams. By default, cybersecurity tasks other than routine bug fixes are routed to Opus 4.8. To use Opus 5.5 itself for security work you need access through Anthropic's Cyber Verification Program, which has three tiers. Biology work has a parallel gate through a new Life Sciences Verification Program. If your evals are security-flavored and the results look oddly like an older model, this is why.
Should you switch to Claude Opus 5.5?
Yes, if you are on Opus 5 today. Opus 5.5 is cheaper per token, uses fewer tokens per task, scores higher on every benchmark Anthropic published for both models, and runs more than 30% faster. There is no capability trade-off to weigh — only the migration cost of the breaking changes.
The decision is less obvious in three cases:
| You are currently on | Recommendation |
|---|---|
| Opus 5 | Switch. Test history-editing code paths first. |
| Fable 5.1 | Move most traffic. Keep Fable for multi-hour, high-stakes runs if your evals show a gap. |
| GPT-6 Sol or Luna | Stay unless quality is your bottleneck. Sol is half the price. |
| Any model, for security work | Apply to the Cyber Verification Program first, or you are testing Opus 4.8. |
Fable 5.1 users have the most to gain financially. Dropping from $10/$50 to $4/$20 for a model that scores higher on Terminal-Bench 4.0, FrontierCode and CursorBench is a 60% list-price reduction before any efficiency gain.
In practice, the tools you already use will make the switch for you. If you are choosing an agent rather than a raw API, our roundup of the best AI coding assistants covers which ones expose model selection.
Frequently asked questions
When was Claude Opus 5.5 released? Claude Opus 5.5 was released on September 22, 2026. It became available the same day on the Claude Platform, AWS, Google Cloud and Microsoft Azure, about 90 minutes before OpenAI launched GPT-6 Sol and Luna.
How much does Claude Opus 5.5 cost per million tokens? Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Cache reads are $0.20 and cache writes are $5 per million, and Fast mode is $8/$40 for up to 2.5x the output speed.
Is Claude Opus 5.5 better than Claude Fable 5.1? On most published coding and knowledge-work benchmarks, yes. Opus 5.5 scores 66.4% on Terminal-Bench 4.0 against Fable 5.1's 55.8%, and 1846 Elo on GDPval-AA v2.1 against 1735, while costing 60% less at list price.
Is Claude Opus 5.5 better than GPT-6 Astra? It leads on agentic coding and trails narrowly on automation. Opus 5.5 beats Astra on Terminal-Bench 4.0 (66.4% vs 57.9%) and FrontierCode (54.4% vs 53.3%), while Astra leads AutomationBench 41.4% to 40.0%. Astra costs $10/$50, so Opus 5.5 is 60% cheaper.
Can you turn off thinking in Claude Opus 5.5? No. Adaptive thinking is always enabled in Opus 5.5 and cannot be disabled. You control how much the model reasons with the effort setting, which has five levels: low, medium, high, xhigh and max.
What is the context window of Claude Opus 5.5? Claude Opus 5.5 has a 1 million-token context window. It can return up to 128,000 output tokens synchronously, and up to 300,000 through the Message Batches API, which is in beta.
The verdict
Opus 5.5 is the rare release where the recommendation is unambiguous: it is better and cheaper than the model it replaces, and better than the model priced above it on the work most developers do. The benchmark lead on Terminal-Bench 4.0 is large enough to survive skepticism about vendor-run evals, and the customer reports all describe the same thing — fewer steps to the same result.
The real work is in the migration. Always-on thinking, enforced preserved thinking and the retired computer-use tool will break harnesses that were tuned for Opus 5. Budget a day for that, not an hour.
If you are comparing this against OpenAI's same-day launch, start with pricing: Sol is half the cost, and for a lot of workloads that ends the conversation. For the hardest agentic coding, Opus 5.5 is the model to beat.
The frontier used to get more expensive every release. This time it got cheaper twice in one afternoon.