Essay/Analysis·Sep 22, 2026

Claude Opus 5.5 Review: Lower Costs, Stronger Agents, Migration Gotchas

Opus 5.5 cuts token prices and improves reported agent performance. Our release review explains the 40% cost claim, benchmark caveats, and API migration changes.

Luke Thompson
Luke ThompsonSep 22, 2026 · 4 min read
In this article
Claude Opus 5.5 Review: Lower Costs, Stronger Agents, Migration Gotchas
Claude Opus 5.5 deserves a place in your next evaluation run. Anthropic released it on September 22, 2026, with lower prices and stronger reported results. The practical question is whether your team can finish the same work with fewer retries, lower costs, and less review time.
Field note

Review scope: This is AI-assisted analysis of Anthropic's release announcement and documentation. The Claude Insider has not independently tested Opus 5.5 for this article. Benchmark results below come from Anthropic's launch materials.

The price cut has two parts

Anthropic lists these standard rates per million tokens on its pricing page:

  • Input: $4 for Opus 5.5, down from $5 for Opus 5.
  • Output: $20, down from $25.
  • Cache reads: $0.20, down from $0.50.
  • Five-minute cache writes: $5, down from $6.25.

Input and output prices fell 20%. Cache-read prices fell 60%. Anthropic's estimate of 40% lower costs on typical workloads also reflects lower token consumption. Treat that estimate as a workload-dependent claim, not a discount on every request.

For a simple calculation, hold usage at one million uncached input tokens and one million output tokens. Opus 5 costs $30; Opus 5.5 costs $24. That saves $6, or 20%, before any change in token usage. Tool charges, cache writes, and infrastructure costs sit outside this example.

Fast mode costs $8 per million input tokens and $40 per million output tokens, with advertised speed up to 2.5 times faster. Budget for that premium separately. For a background job, paying more for speed only makes sense when finishing sooner has measurable value.

Stronger results, with exceptions

Anthropic's release comparison reports these results for Opus 5.5 versus Opus 5:

  • Terminal-Bench 4.0: 66.4% versus 52.3%.
  • CursorBench 4.0: 57.8% versus 46.6%.
  • AutomationBench: 40.0% versus 26.9%.

Opus 5.5 does not lead every listed comparison. GPT-6 Astra scores 41.4% on AutomationBench and 64.6% on Terminal-Bench-Science, versus 40.0% and 58.7% for Opus 5.5. These point estimates do not establish meaningful differences by themselves.

Read the evaluation notes alongside the scores. Most Opus 5.5 results use maximum effort; Terminal-Bench uses xhigh. The Terminal-Bench standard error for Opus 5.5 is ±2.6 percentage points. Some evaluations substitute older Claude models when safeguards intervene. AutomationBench uses no fallback and counts those interventions as failures.

That matters when buying an agent for a specific job. A published result can depend on the model, tools, fallback behavior, and testing budget together. Compare the whole setup you intend to run. Do not turn one benchmark score into a promise about your own completion rate.

What you can use today

The Opus product page lists access for Pro, Max, Team, and Enterprise users, plus the Claude Platform, AWS, Google Cloud, and Microsoft Foundry. Anthropic reports output generation more than 30% faster than Opus 5. That does not mean every end-to-end task finishes 30% sooner.

The model documentation specifies a one-million-token context window, up to 128,000 output tokens, and the API model ID claude-opus-5-5. Adaptive thinking stays on. Default effort is medium, compared with high on Opus 5. Set effort explicitly when comparing results.

For subscription users, Anthropic also announced higher five-hour limits on Pro, Max, Team, and seat-based Enterprise plans, plus a rate-limit reset users can save. Sonnet 5.5 and Haiku 5.5 are planned for the coming weeks.

Check your integration before changing the model ID

Anthropic's migration guide documents changes that affect existing Messages API integrations:

  • Thinking settings: disabled thinking and manual thinking budgets return errors. Use the supported effort control.
  • Forced tools: tool_choice values any and tool are unsupported. Review auto tool selection and strict schemas.
  • Computer use: the Claude API and Google Cloud require computer_toolset_20260801 in place of computer_20251124, plus agent-loop changes. Amazon Bedrock retains the older tool.
  • Progress updates: narration between tool calls moves into thinking blocks. Interfaces that read only text blocks can stop showing progress.

Preserved thinking adds another integration check. Changing earlier messages, system instructions, or tool definitions can invalidate returned thinking blocks. Prefix validation applies by default to API accounts created on or after August 31, 2026. Anthropic says Claude Code, claude.ai, Managed Agents, and the Claude Agent SDK already handle this behavior.

Our assessment

Start with a small set of completed tasks from your own work. Include a stubborn bug, a document with figures you can verify, and a workflow that uses several tools. Keep inputs and acceptance criteria fixed. Run repeated trials at explicit effort settings so you can see variation.

Track whether each run passes review. Record elapsed time, token costs, tool charges, retries, and human correction time. Calculate cost per successful task by dividing the cost of all attempts, including failures, by the number that pass. Report undefined when none pass.

For teams already using Opus 5, the lower rates justify testing Opus 5.5 now. Keep your current routing until the new model meets your quality bar on representative work. Expand its share of traffic when the savings survive the cost of failures and human review.

Sources

THE CLAUDE INSIDER

Get the Claude playbook in your inbox.

One weekly email for Claude and Claude Code users. Real workflows, no hype. Subscribe and we send you The Claude Power-User Cheatsheet.

GUIDES AND COMPARISONS

— ¶ —

Luke Thompson

Luke Thompson

Editor-in-Chief · The Claude Insider

Luke Thompson is the founder of The Operations Guide, LLC and editor of The Claude Insider. Based in Jonesborough, Tennessee, he has spent years building AI-augmented business systems and automation workflows for operators and teams. He began working with large language models in production well before the current wave of consumer AI tools, integrating them into client workflows, content pipelines, and operational infrastructure. At The Claude Insider, he writes about Claude with the specificity of someone who uses it daily as a professional tool — not as a reviewer or commentator, but as a builder. His coverage focuses on what actually works: prompt patterns, API integration strategies, agentic workflows, and the real-world tradeoffs that practitioners face. He is not affiliated with Anthropic, PBC.

Articles are researched and drafted with AI assistance, reviewed and edited by Luke Thompson.

From Reading to Action

Know where AI can pay off in your company.

Take the free two-minute AI Readiness Assessment. See your score, the two gaps holding you back, and the next move worth making.

Get your readiness score

Instant report · No account to start

Related reading

View archive →