Essay/Claude·Apr 12, 2026

Claude Sonnet 4.6: Opus-Level Power at Sonnet Pricing

Anthropic's Sonnet 4.6 posts 79.6% on SWE-bench, 72.5% on OSWorld, and 60.4% on ARC-AGI-2. How it stacks up against Gemini 3 Deep Think and GPT-5.2.

Luke Thompson
Luke ThompsonApr 12, 2026 · 10 min read
In this article
Claude Sonnet 4.6: Opus-Level Power at Sonnet Pricing
Anthropic just released Claude Sonnet 4.6, and the gap between Sonnet and Opus has never been smaller. Two weeks after launching Opus 4.6, Anthropic is bringing much of that flagship intelligence down to the Sonnet tier - with the same $3/$15 per million token pricing that made Sonnet the workhorse model for developers. Sonnet 4.6 is now the default model for Free and Pro users on claude.ai and Claude Cowork. If you've been using Claude, you're already on it. And the numbers back the hype: in Claude Code testing, developers preferred Sonnet 4.6 over Sonnet 4.5 70% of the time - and over the previous flagship Opus 4.5 59% of the time.

Why This Release Matters

Here's the thing about the Sonnet tier: it's where most of the actual work happens. Opus gets the headlines, but Sonnet is what developers hit millions of times a day in production. It's the model powering the majority of Claude API calls, the default in Claude Code, and the go-to for teams that need speed and capability without burning through budgets.

Sonnet 4.6 changes the calculus for when you need to reach for Opus. Developers with early access have been preferring Sonnet 4.6 over its predecessor by a wide margin. The surprise: many preferred it to Claude Opus 4.5, the smartest model from November 2025. Performance that previously required an Opus-class model is now available at Sonnet pricing.

What's New in Sonnet 4.6

1M Token Context Window (Beta)

Sonnet 4.6 doubles the previous Sonnet context limit to 1 million tokens in beta. That's enough to hold entire codebases, lengthy contracts, or dozens of research papers in a single request. When Opus 4.6 launched with 1M context two weeks ago, many developers wondered when Sonnet would follow. The answer: now.

Related essay
Claude Opus 4.6 Released: Agent Teams and 1M Context

For practical use, this means you can analyze a full Next.js application's source code, review an entire quarter of SEC filings, or process a complete technical documentation site - all in one conversation without chunking or RAG workarounds.

Coding That Closes the Gap

Sonnet 4.6 brings much-improved coding skills to the Sonnet tier. The model excels at complex code fixes, especially when searching across large codebases is essential. It handles multi-file refactors, understands project-level architecture, and catches its own mistakes more reliably than Sonnet 4.5.

The numbers: 79.6% on SWE-bench Verified, up from Sonnet 4.5's 77.2% and within 1.2 points of Opus 4.6's 80.8%. That's the narrowest Sonnet-to-Opus gap in any Claude generation. With a prompt modification, Anthropic saw it hit 80.2%. The SWE-bench score was averaged over 10 trials with thinking turned off - meaning this is the model's baseline coding ability, not its ceiling.

Computer Use Hits Human-Level Tasks

OSWorld is the standard benchmark for AI computer use, and Sonnet models have been making steady gains across sixteen months. Early Sonnet 4.6 users are seeing human-level capability in tasks like navigating complex spreadsheets, filling out multi-step web forms, and coordinating actions across multiple browser tabs.

Related essay
Claude vs ChatGPT (Feb 2026): Benchmarks, Pricing, Verdict

Sonnet 4.6 hit 72.5% on OSWorld-Verified - within 0.2 points of Opus 4.6's 72.7%, and nearly double GPT-5.2's 38.2%. In real-world testing, insurance company Pace reported 94% accuracy on their computer use tasks. For anyone building browser automation or desktop agent workflows, this is the most capable Sonnet model for the job.

Document Comprehension Matches Opus

On OfficeQA - which measures how well a model can read enterprise documents like charts, PDFs, and tables - Sonnet 4.6 matches Opus 4.6 performance. That's a meaningful upgrade for document comprehension workloads. If you're building document processing pipelines, you can now get Opus-quality extraction at one-third the cost.

Design and Frontend Output

Early access customers independently described visual outputs from Sonnet 4.6 as notably more polished, with better layouts, animations, and design sensibility. The model also needed fewer rounds of iteration to reach production-quality results. For frontend developers and designers using Claude to generate UI code, this means less back-and-forth to get something that actually looks good.

Read the original source on anthropic.com

Benchmark Performance

The numbers paint a clear picture. Sonnet 4.6 posts improvements across the board, with standout results on reasoning and coding benchmarks.

79.6%
SWE-bench Verified
72.5%
OSWorld-Verified
60.4%
ARC-AGI-2 (High Effort)
74.1%
GPQA Diamond
89%
Math Benchmark

The story across these benchmarks is consistent: Sonnet 4.6 is within striking distance of Opus on coding and computer use, while posting major gains in reasoning and math. The 1.2-point SWE-bench gap and 0.2-point OSWorld gap are essentially rounding errors. Where Opus still leads meaningfully is scientific reasoning (GPQA Diamond: 91.3% vs 74.1%) and the hardest novel reasoning tasks. For a model at $3 per million input tokens, these numbers rewrite what's possible at the Sonnet price point.

Arena Rankings: Context

As of today, Opus 4.6 sits at the top of the Arena leaderboard (formerly LMArena/LMSYS Chatbot Arena) with a 1,506 Elo score in the thinking variant. Sonnet 4.5 previously held positions #11 and #12 on the text leaderboard. Sonnet 4.6 is brand new - expect Arena rankings to update as users cast votes over the coming days.

Given the benchmark improvements over Sonnet 4.5 and the early-access preference data showing it competing with Opus 4.5, there's reason to expect Sonnet 4.6 will climb significantly on the Arena leaderboard.

Gemini 3 Deep Think: The Reasoning Rival

Five days before Sonnet 4.6 launched, Google released a major upgrade to Gemini 3 Deep Think on February 12, 2026. It's the clearest competitor to Claude's reasoning capabilities - but the two models are solving different problems.

On pure reasoning benchmarks, Deep Think is ahead. Its 84.6% on ARC-AGI-2 (verified by the ARC Prize Foundation) towers over Sonnet 4.6's 60.4% and even Opus 4.6's ~65%. On Humanity's Last Exam - the benchmark designed to be unsolvable by current AI - Deep Think scored 48.4%, blowing past GPT-5 Pro's ~31.6%. Its Codeforces Elo of 3,455 puts it at elite competitive programming levels. Google also reports gold medal-level performance on the International Physics and Chemistry Olympiads.

Field note

Gemini 3 Deep Think key scores: 84.6% ARC-AGI-2 | 48.4% Humanity's Last Exam | 3,455 Codeforces Elo | Gold medal-level on International Physics and Chemistry Olympiads

But here's the tradeoff: Deep Think is a specialized "slow reasoner." It's designed for asynchronous scientific discovery and mathematical proof verification - problems where you're willing to wait minutes for an answer. Sonnet 4.6 is a general-purpose workhorse that responds in seconds. They occupy different positions in the emerging split between fast conversational AI and deep reasoning engines.

Where Sonnet 4.6 dominates the comparison is practical software engineering and computer use. Its 72.5% on OSWorld-Verified is nearly double GPT-5.2's 38.2%, and Deep Think doesn't even compete on that benchmark. SWE-bench Verified at 79.6% reflects real-world coding ability that Deep Think's Codeforces ranking (competitive algorithm puzzles) doesn't capture. For the 95% of developers who need a model to refactor code, debug production issues, and navigate browser-based workflows, Sonnet 4.6 is the better tool.

Availability also differs. Deep Think is restricted to Google AI Ultra subscribers and a limited API early access program. Sonnet 4.6 is available to everyone on claude.ai Free and Pro tiers, the full API, and Claude Code. For teams evaluating which model to build on, accessibility and pricing matter as much as peak benchmark scores.

New Under the Hood: Adaptive Thinking and Context Compaction

Two beta features shipped alongside the model upgrade. Adaptive thinking lets Sonnet 4.6 dynamically adjust how much reasoning effort it applies to a problem - spending more compute on hard questions and less on simple ones. This is similar in concept to what powers Deep Think's extended reasoning, but calibrated for real-time use rather than marathon problem-solving sessions.

Context compaction automatically summarizes older conversation context, effectively enabling unlimited conversation length within the 1M token window. Instead of hitting a wall when context fills up, the model compresses earlier turns while preserving key information. For long coding sessions or multi-hour research workflows, this means fewer restarts and less context management overhead.

Safety and Prompt Injection Resistance

Anthropic ran extensive safety evaluations. The results: Sonnet 4.6 is as safe as, or safer than, other recent Claude models. Their safety researchers described it as having "a broadly warm, honest, prosocial, and at times funny character, very strong safety behaviors, and no signs of major concerns around high-stakes forms of misalignment."

The real improvement is in prompt injection resistance. Computer use is powerful but risky - malicious actors can hide instructions on websites to hijack the model. Sonnet 4.6 is a major improvement over Sonnet 4.5 in resisting these attacks, performing similarly to Opus 4.6. If you're building computer use agents, this directly reduces your security surface.

Field note

Prompt injection resistance in Sonnet 4.6 now matches Opus 4.6 levels. For computer use and agentic workflows where models interact with untrusted content, this is a meaningful safety upgrade over Sonnet 4.5.

Pricing and Availability

Sonnet 4.6 is available now across all channels:

  • claude.ai - Default model for Free and Pro users, effective today
  • Claude Cowork - Default model for desktop knowledge work
  • Claude API - Model ID: claude-sonnet-4-6
  • Pricing - $3 per million input tokens, $15 per million output tokens (unchanged from Sonnet 4.5)
  • Context - 1M token window in beta

Same pricing, substantially more capable. Anthropic is following the pattern they set with Opus 4.6 - delivering more intelligence at the same cost per token. For production workloads, this is a straightforward model swap with the ID claude-sonnet-4-6.

How It Fits the Claude Lineup

Anthropic has maintained a rapid release pace. Here's where Sonnet 4.6 sits in the current model family:

  • Haiku 4.5 (October 2025) - Budget tier. Sonnet 4-level performance at one-third the cost.
  • Sonnet 4.5 (September 2025) - Previous workhorse. Now superseded by 4.6.
  • Opus 4.5 (November 2025) - Previous flagship. Strong general intelligence, but Sonnet 4.6 now competes in many tasks.
  • Opus 4.6 (February 5, 2026) - Current flagship. Agent teams, 1M context, top Arena scores.
  • Sonnet 4.6 (February 17, 2026) - New sweet spot. Near-Opus intelligence at Sonnet pricing.

An updated Haiku model is expected to follow in the coming weeks, which would complete the 4.6 generation across all three tiers.

What This Means for Developers

If you're building on the Claude API, the practical impact is clear. The 1M context window means simpler architectures - fewer RAG pipelines needed when you can just feed the full document set. The coding improvements mean fewer iterations per task. The prompt injection resistance means safer computer use agents.

For teams that were considering upgrading to Opus for specific workloads, Sonnet 4.6 narrows the gap enough that many use cases no longer justify the 5x price premium. Document comprehension matching Opus on OfficeQA is a concrete example - there's no reason to pay Opus pricing for PDF extraction if Sonnet delivers the same accuracy.

The model ID for the API is claude-sonnet-4-6. If you're on claude.ai Free or Pro, you're already using it.

Quick Takeaway

Claude Sonnet 4.6 delivers near-Opus intelligence at Sonnet pricing: 79.6% SWE-bench, 72.5% OSWorld, 60.4% ARC-AGI-2, and 89% on math. The 1M token context window, adaptive thinking, context compaction, and Opus-matching document comprehension make this the most capable mid-tier model available.

Gemini 3 Deep Think wins on pure reasoning benchmarks (84.6% ARC-AGI-2, 48.4% Humanity's Last Exam) but operates as a specialized slow reasoner with limited availability. For the practical work most developers actually do - coding, computer use, document processing, and general-purpose AI tasks - Sonnet 4.6 at $3/$15 per million tokens is the new default. Opus 4.6 still leads on the hardest scientific reasoning, but the gap has never been thinner.

Related essay
GPT-5.4 vs Claude Opus 4.6: Full Benchmark & Feature Comparison (March 2026)
Related essay
Claude Opus 4.7 Is Here: Anthropic's Most Capable Coding Model Yet

THE CLAUDE INSIDER

Get the Claude playbook in your inbox.

One weekly email for Claude and Claude Code users. Real workflows, no hype. Subscribe and we send you The Claude Power-User Cheatsheet.

GUIDES AND COMPARISONS

— ¶ —

Luke Thompson

Luke Thompson

Editor-in-Chief · The Claude Insider

Luke Thompson is the founder of The Operations Guide, LLC and editor of The Claude Insider. Based in Jonesborough, Tennessee, he has spent years building AI-augmented business systems and automation workflows for operators and teams. He began working with large language models in production well before the current wave of consumer AI tools, integrating them into client workflows, content pipelines, and operational infrastructure. At The Claude Insider, he writes about Claude with the specificity of someone who uses it daily as a professional tool — not as a reviewer or commentator, but as a builder. His coverage focuses on what actually works: prompt patterns, API integration strategies, agentic workflows, and the real-world tradeoffs that practitioners face. He is not affiliated with Anthropic, PBC.

Articles are researched and drafted with AI assistance, reviewed and edited by Luke Thompson.

From Reading to Action

Know where AI can pay off in your company.

Take the free two-minute AI Readiness Assessment. See your score, the two gaps holding you back, and the next move worth making.

Get your readiness score

Instant report · No account to start

Related reading

View archive →