September 22, 2026 · 9 min read
Opus 5.5 beats Fable 5.1 at 40% of the price. Read the benchmark table twice.
Anthropic released Opus 5.5 today, September 22, 2026. Two months ago I wrote that Opus 5’s real story was the API changes, not the benchmarks. This time the benchmarks deserve a closer look, partly because they are big and partly because the first independent run already disagrees with one of them.
One thing first, since I went looking for it: there is no Sonnet 5.5 yet. The announcement says Sonnet 5.5 and Haiku 5.5 will follow “in the following weeks” with similar improvements. So this post is about Opus 5.5 alone, and Sonnet 5 is still the current Sonnet.
Quick picks
| If you are… | Pick | Why |
|---|---|---|
| Running long agentic coding sessions | Opus 5.5 | Leads Terminal-Bench 4.0 and FrontierCode in every table, including the independent one |
| Paying Fable 5.1 prices for coding work | Opus 5.5 | Higher scores at $4/$20 against Fable’s $10/$50 |
| Doing agentic science or SaaS workflow automation | Test GPT-6 Astra too | Astra leads Terminal-Bench-Science and AutomationBench |
Relying on thinking: disabled or forced tool use | Stay on Opus 5 until you migrate | Both now return a 400 |
| Waiting for a cheaper mid-tier upgrade | Sonnet 5 for now | Sonnet 5.5 is announced, not released |
Anthropic’s table
Here is the full comparison from the launch page:
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | n/a | 41.7% |
| GDPval-AA v2.1 (Elo) | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity’s Last Exam (tools) | 67.7% | 65.6% | 63.6% | 57.2% | n/a |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 (partial) | 81.8% | 80.7% | 74.0% | n/a | n/a |
| Chartography (tools) | 89.0% | 88.4% | 83.4% | n/a | n/a |
The headline is the first row. Terminal-Bench measures an agent working in a real shell on multi-step tasks, which is the closest public benchmark to what I actually do with these models all day. A jump from 52.3% to 66.4% in two months, and a 10-point lead over Fable 5.1, would be the largest single-release gain on that benchmark I have seen.
GDPval-AA is the other big one. It scores professional deliverables (reports, spreadsheets, analyses) head to head and reports an Elo rating. A 300-point gap over GPT-6 Astra means Opus 5.5 wins most direct comparisons, not just a few.
Two rows go the other way, and Anthropic printed them anyway, which I appreciate. Astra leads AutomationBench (business workflows across SaaS tools) by 1.4 points, a gap too small to matter. It leads Terminal-Bench-Science by about 6 points, which does matter if your agents do scientific computing.
The independent run is smaller
Artificial Analysis ran its own evaluations the same day. Their summary: Opus 5.5 at max effort scores 58 on their Intelligence Index, the highest they have measured “by several points,” and it leads six of their ten component evaluations.
But their Terminal-Bench 4.0 score is 59.6%, not 66.4%. That is still 11 points over Opus 5, and it ties GPT-6 Astra at xhigh effort. What it does not show is a 10-point lead. Their Humanity’s Last Exam number is 61.4% against Anthropic’s 67.7%, although that gap is at least partly because the two runs use different tool setups.
This is normal. Vendors tune the harness, the effort level, and the timeout budget, and independent evaluators use one standard harness for every model. Neither number is wrong. They measure different things: Anthropic’s shows what the model can do in a setup built for it, and Artificial Analysis shows how it compares with everything else under the same conditions. If you are choosing between Opus 5.5 and Astra for terminal agents, “tied, at a lower price” is the honest summary. That is still a good result.
Less code, not always better code
Sonar ran its code-quality evaluation across 4,444 Java tasks from HumanEval, MBPP, and ComplexCodeEval and compared Opus 5.5 against Opus 5. This is the most interesting data from launch day, because it measures what the code looks like, not just whether it passes:
- Pass rate went down slightly: 87.7% against 88.6% on Opus 5.
- 27.5% fewer lines of code for the same tasks, and 40% fewer output tokens.
- Blocker-severity bugs dropped 41% per million lines, and blocker vulnerabilities dropped 53%.
- Overall bug density rose 12%, and concurrency issues rose 44% per million lines.
- Comment density fell from 10.5% to 3.1%.
My reading: Opus 5.5 writes tighter code with fewer catastrophic mistakes, but it spends fewer lines on each problem, so the bugs that remain are packed closer together. The concurrency number is the one I would watch. If your agents write Go with goroutines or anything with shared state, review those diffs more carefully than usual. The drop in comments will please some people and annoy others; if you want comments, you now have to ask for them.
The price, and the catch in the price
Opus 5.5 costs $4 per million input tokens and $20 per million output, down from Opus 5’s $5/$25. Cache reads dropped from $0.50 to $0.20, batch is half price at $2/$10, and fast mode runs at $8/$40 on the Claude API only. The 1M context window and 128k max output are unchanged.
Fable 5.1 and GPT-6 Astra both list at $10/$50, which is where the “60% cheaper” headlines come from. Anthropic’s own claim is more careful: about 40% less than Opus 5 on typical workloads, because the model also uses fewer tokens per task. The Sonar data supports that.
The catch is effort. Opus 5.5 thinks more per turn than Opus 5 at the same effort level, most of all at xhigh and max. Kingy AI’s comparison estimates about 119,000 output tokens per task at max effort, against about 27,000 for Astra. I have not verified that figure, but if it is even roughly correct, a lower per-token price does not guarantee a lower bill at max effort. The default effort also dropped from high to medium, so a request that never set effort now does less thinking, not more. Set it explicitly and measure the cost per task, not the cost per token.
What breaks when you swap the model string
Opus 5 changed defaults. Opus 5.5 removes options. The migration guide lists four breaking changes, and each one returns a 400 instead of failing silently, which is at least easy to spot:
- Thinking can’t be disabled. Both
thinking: {type: "disabled"}and manualbudget_tokensare rejected. Effort is the only control now; uselowwhere you used to turn thinking off. - Forced tool use is gone.
tool_choiceofanyortoolis rejected. Useautowithstrict: trueor structured outputs, and name the tool in the prompt. - Thinking blocks are tied to the model and the conversation. On accounts created on or after August 31, 2026, replaying a thinking block after you edit the system prompt, the tools, or an earlier message returns a 400. Keep conversations append-only.
- The old
computer_20251124tool is rejected on the Claude API and Google Cloud. Move tocomputer_toolset_20260801. Bedrock still accepts the old tool.
A fifth change doesn’t error at all, which makes it the easiest to miss. The short text the model writes between tool calls now arrives in thinking blocks, and at the default display setting those blocks are empty. If your UI shows “now reading the config file…” style progress updates, it will go quiet with no error. Set thinking.display to "updates" (beta) to get them back.
Point 3 also matters for fallback, which The New Stack headlined as “your agent calls might secretly get routed to an older model”. The documented mechanism is this: Opus 5.5 now runs biology classifiers in addition to cyber ones, and with server-side fallback enabled, a declined request is retried on another model. Only Fable 5.1 and Mythos 5.1 can read Opus 5.5’s thinking blocks, so a fallback to anything else continues the conversation without the earlier reasoning. It is documented, and it is opt-in, but if you enabled fallbacks: "default" on Opus 5 and never thought about it again, check your logs for which model actually answered.
What I am doing
- Swap to
claude-opus-5-5in a dev branch and grep for"disabled",budget_tokens, andtool_choice. Those are the 400s. - Set
effortexplicitly everywhere. The default moved down, and the thinking per level moved up, so no carried-over setting means the same thing it did last week. - Compare cost per completed task against Opus 5 on a real workload before switching production traffic, not cost per token.
- Add a review step for concurrency code, based on the Sonar numbers.
- Wait for Sonnet 5.5 before touching anything that runs on Sonnet 5. If it follows the same pattern, it will come with the same breaking changes.
The model is a real step forward, and even the independent numbers put it at or near the top of every agentic benchmark at a lower price than its competitors. Anthropic’s table just makes the lead look larger than a neutral harness does. Plan for “tied with Astra, cheaper, and better at knowledge work,” and treat anything beyond that as a bonus.
Sources: Introducing Claude Opus 5.5 (Anthropic, September 22, 2026); What’s new in Claude Opus 5.5 and the migration guide (Claude Platform Docs); Artificial Analysis; Sonar; VentureBeat; Kingy AI. Figures were reported on launch day and have not been independently reproduced.