Opus 5.5 beats Fable 5.1 at 40% of the price. Read the benchmark table twice.

ai llm anthropic claude opus benchmarks

Anthropic released Opus 5.5 today, September 22, 2026. Two months ago I wrote that Opus 5’s real story was the API changes, not the benchmarks. This time the benchmarks deserve a closer look, partly because they are big and partly because the first independent run already disagrees with one of them.

One thing first, since I went looking for it: there is no Sonnet 5.5 yet. The announcement says Sonnet 5.5 and Haiku 5.5 will follow “in the following weeks” with similar improvements. So this post is about Opus 5.5 alone, and Sonnet 5 is still the current Sonnet.

Quick picks

If you are…PickWhy
Running long agentic coding sessionsOpus 5.5Leads Terminal-Bench 4.0 and FrontierCode in every table, including the independent one
Paying Fable 5.1 prices for coding workOpus 5.5Higher scores at $4/$20 against Fable’s $10/$50
Doing agentic science or SaaS workflow automationTest GPT-6 Astra tooAstra leads Terminal-Bench-Science and AutomationBench
Relying on thinking: disabled or forced tool useStay on Opus 5 until you migrateBoth now return a 400
Waiting for a cheaper mid-tier upgradeSonnet 5 for nowSonnet 5.5 is announced, not released

Anthropic’s table

Here is the full comparison from the launch page:

BenchmarkOpus 5.5Fable 5.1Opus 5GPT-6 AstraGPT-5.6 Sol
Terminal-Bench 4.066.4%55.8%52.3%57.9%37.3%
FrontierCode v1.1 (Main)54.4%50.3%48.0%53.3%47.5%
CursorBench 4.057.8%51.8%46.6%n/a41.7%
GDPval-AA v2.1 (Elo)18461735170815421588
AutomationBench40.0%31.4%26.9%41.4%28.8%
Humanity’s Last Exam (tools)67.7%65.6%63.6%57.2%n/a
Terminal-Bench-Science 0.158.7%52.6%29.0%64.6%22.4%
OSWorld 2.0 (partial)81.8%80.7%74.0%n/an/a
Chartography (tools)89.0%88.4%83.4%n/an/a

The headline is the first row. Terminal-Bench measures an agent working in a real shell on multi-step tasks, which is the closest public benchmark to what I actually do with these models all day. A jump from 52.3% to 66.4% in two months, and a 10-point lead over Fable 5.1, would be the largest single-release gain on that benchmark I have seen.

GDPval-AA is the other big one. It scores professional deliverables (reports, spreadsheets, analyses) head to head and reports an Elo rating. A 300-point gap over GPT-6 Astra means Opus 5.5 wins most direct comparisons, not just a few.

Two rows go the other way, and Anthropic printed them anyway, which I appreciate. Astra leads AutomationBench (business workflows across SaaS tools) by 1.4 points, a gap too small to matter. It leads Terminal-Bench-Science by about 6 points, which does matter if your agents do scientific computing.

The independent run is smaller

Artificial Analysis ran its own evaluations the same day. Their summary: Opus 5.5 at max effort scores 58 on their Intelligence Index, the highest they have measured “by several points,” and it leads six of their ten component evaluations.

But their Terminal-Bench 4.0 score is 59.6%, not 66.4%. That is still 11 points over Opus 5, and it ties GPT-6 Astra at xhigh effort. What it does not show is a 10-point lead. Their Humanity’s Last Exam number is 61.4% against Anthropic’s 67.7%, although that gap is at least partly because the two runs use different tool setups.

This is normal. Vendors tune the harness, the effort level, and the timeout budget, and independent evaluators use one standard harness for every model. Neither number is wrong. They measure different things: Anthropic’s shows what the model can do in a setup built for it, and Artificial Analysis shows how it compares with everything else under the same conditions. If you are choosing between Opus 5.5 and Astra for terminal agents, “tied, at a lower price” is the honest summary. That is still a good result.

Less code, not always better code

Sonar ran its code-quality evaluation across 4,444 Java tasks from HumanEval, MBPP, and ComplexCodeEval and compared Opus 5.5 against Opus 5. This is the most interesting data from launch day, because it measures what the code looks like, not just whether it passes:

  • Pass rate went down slightly: 87.7% against 88.6% on Opus 5.
  • 27.5% fewer lines of code for the same tasks, and 40% fewer output tokens.
  • Blocker-severity bugs dropped 41% per million lines, and blocker vulnerabilities dropped 53%.
  • Overall bug density rose 12%, and concurrency issues rose 44% per million lines.
  • Comment density fell from 10.5% to 3.1%.

My reading: Opus 5.5 writes tighter code with fewer catastrophic mistakes, but it spends fewer lines on each problem, so the bugs that remain are packed closer together. The concurrency number is the one I would watch. If your agents write Go with goroutines or anything with shared state, review those diffs more carefully than usual. The drop in comments will please some people and annoy others; if you want comments, you now have to ask for them.

The price, and the catch in the price

Opus 5.5 costs $4 per million input tokens and $20 per million output, down from Opus 5’s $5/$25. Cache reads dropped from $0.50 to $0.20, batch is half price at $2/$10, and fast mode runs at $8/$40 on the Claude API only. The 1M context window and 128k max output are unchanged.

Fable 5.1 and GPT-6 Astra both list at $10/$50, which is where the “60% cheaper” headlines come from. Anthropic’s own claim is more careful: about 40% less than Opus 5 on typical workloads, because the model also uses fewer tokens per task. The Sonar data supports that.

The catch is effort. Opus 5.5 thinks more per turn than Opus 5 at the same effort level, most of all at xhigh and max. Kingy AI’s comparison estimates about 119,000 output tokens per task at max effort, against about 27,000 for Astra. I have not verified that figure, but if it is even roughly correct, a lower per-token price does not guarantee a lower bill at max effort. The default effort also dropped from high to medium, so a request that never set effort now does less thinking, not more. Set it explicitly and measure the cost per task, not the cost per token.

What breaks when you swap the model string

Opus 5 changed defaults. Opus 5.5 removes options. The migration guide lists four breaking changes, and each one returns a 400 instead of failing silently, which is at least easy to spot:

  1. Thinking can’t be disabled. Both thinking: {type: "disabled"} and manual budget_tokens are rejected. Effort is the only control now; use low where you used to turn thinking off.
  2. Forced tool use is gone. tool_choice of any or tool is rejected. Use auto with strict: true or structured outputs, and name the tool in the prompt.
  3. Thinking blocks are tied to the model and the conversation. On accounts created on or after August 31, 2026, replaying a thinking block after you edit the system prompt, the tools, or an earlier message returns a 400. Keep conversations append-only.
  4. The old computer_20251124 tool is rejected on the Claude API and Google Cloud. Move to computer_toolset_20260801. Bedrock still accepts the old tool.

A fifth change doesn’t error at all, which makes it the easiest to miss. The short text the model writes between tool calls now arrives in thinking blocks, and at the default display setting those blocks are empty. If your UI shows “now reading the config file…” style progress updates, it will go quiet with no error. Set thinking.display to "updates" (beta) to get them back.

Point 3 also matters for fallback, which The New Stack headlined as “your agent calls might secretly get routed to an older model”. The documented mechanism is this: Opus 5.5 now runs biology classifiers in addition to cyber ones, and with server-side fallback enabled, a declined request is retried on another model. Only Fable 5.1 and Mythos 5.1 can read Opus 5.5’s thinking blocks, so a fallback to anything else continues the conversation without the earlier reasoning. It is documented, and it is opt-in, but if you enabled fallbacks: "default" on Opus 5 and never thought about it again, check your logs for which model actually answered.

What I am doing

  1. Swap to claude-opus-5-5 in a dev branch and grep for "disabled", budget_tokens, and tool_choice. Those are the 400s.
  2. Set effort explicitly everywhere. The default moved down, and the thinking per level moved up, so no carried-over setting means the same thing it did last week.
  3. Compare cost per completed task against Opus 5 on a real workload before switching production traffic, not cost per token.
  4. Add a review step for concurrency code, based on the Sonar numbers.
  5. Wait for Sonnet 5.5 before touching anything that runs on Sonnet 5. If it follows the same pattern, it will come with the same breaking changes.

The model is a real step forward, and even the independent numbers put it at or near the top of every agentic benchmark at a lower price than its competitors. Anthropic’s table just makes the lead look larger than a neutral harness does. Plan for “tied with Astra, cheaper, and better at knowledge work,” and treat anything beyond that as a bonus.


Sources: Introducing Claude Opus 5.5 (Anthropic, September 22, 2026); What’s new in Claude Opus 5.5 and the migration guide (Claude Platform Docs); Artificial Analysis; Sonar; VentureBeat; Kingy AI. Figures were reported on launch day and have not been independently reproduced.