Essay/Claude·Apr 12, 2026

GPT-5.4 vs Claude Opus 4.6: Full Benchmark & Feature Comparison (March 2026)

OpenAI launched GPT-5.4 on March 5, 2026. Here's a complete head-to-head against Claude Opus 4.6 across benchmarks, computer use, coding, pricing, and real-world professional tasks.

Luke Thompson
Luke ThompsonApr 12, 2026 · 7 min read
In this article
GPT-5.4 vs Claude Opus 4.6: Full Benchmark & Feature Comparison (March 2026)
OpenAI released GPT-5.4 on March 5, 2026 — its "most capable and efficient frontier model for professional work." That puts a direct target on Claude Opus 4.6, which only a month ago claimed the #1 spot on the Chatbot Arena for the first time since Gemini 3 established its reign. So where does GPT-5.4 actually beat Opus 4.6, where does Claude still lead, and which model should you reach for in March 2026? Here's the full breakdown.
Watch: GPT-5.4 vs Claude Opus 4.6 — independent head-to-head comparison
Field note

Published: March 6, 2026. This article covers GPT-5.4 (released March 5, 2026) vs Claude Opus 4.6 (released February 5, 2026). GPT-5.4 Arena rankings are not yet available — it takes 2–4 weeks for a new model to accumulate enough Arena votes for a reliable ELO score.

What's New in GPT-5.4

GPT-5.4 is not an incremental update. OpenAI describes it as a unification of the best advances from recent releases — the coding muscle of GPT-5.3-Codex, the reasoning depth of GPT-5.2 Thinking, and new native computer-use that no prior mainline model had. There are three versions: Claude Sonnet 4.5 Is Gone: Your Migration Playb...

  • GPT-5.4 — Standard version, rolling out in ChatGPT and the API. Available to Plus, Team, and Pro subscribers.
  • GPT-5.4 Thinking — Reasoning-optimized variant with upfront thinking plans. Replaces GPT-5.2 Thinking (which goes away in 3 months). Available to Plus, Team, and Pro.
  • GPT-5.4 Pro — Maximum capability version. Available to Pro and Enterprise plans only.

OpenAI highlighted six key improvement areas in the GPT-5.4 release:

  • Coding, document understanding, tool use, and instruction following
  • Image perception and multimodal tasks
  • Long-running task execution and multi-step agent workflows
  • Token efficiency and end-to-end performance on tool-heavy workloads
  • Agentic web search and multi-source synthesis for hard-to-locate information
  • Upfront thinking plans so users can redirect mid-response
Field note

The biggest new capability: GPT-5.4 is the first mainline OpenAI model with built-in native computer use — it can interact directly with software by writing code or executing mouse and keyboard commands based on screenshots. Claude Opus 4.6 has had computer use since launch, and Anthropic led here, but GPT-5.4 closes the gap significantly.

Benchmark Comparison

No single model wins across every benchmark. GPT-5.4 wins five categories, Gemini 3.1 Pro wins four, and Claude Opus 4.6 wins three in current evaluations. The pattern: GPT-5.4 dominates professional knowledge work and computer-use workloads; Opus 4.6 leads on reasoning-heavy academic and scientific benchmarks.

Professional Knowledge Work (GDPval)

GDPval tests agents across 44 occupations — writing, research, financial analysis, legal work — comparing AI outputs against what industry professionals produce. GPT-5.4 scores 83.0%, meaning it meets or exceeds human professional output in more than 4 out of 5 cases. This compares to 70.9% for GPT-5.2. Claude Opus 4.6's GDPval score was not separately disclosed at launch, but Sonnet 4.6 holds the #1 position on GDPval-AA (expert office work) at 1,633 Elo.

Spreadsheet & Financial Modeling

On OpenAI's internal benchmark for investment banking analyst–grade spreadsheet modeling, GPT-5.4 scores 87.3% versus 68.4% for GPT-5.2. This is a major jump. Claude Opus 4.6 has historically been competitive on financial analysis tasks, but GPT-5.4's spreadsheet performance is the strongest public result any model has reported for this category.

Reasoning & Science (GPQA Diamond, MMLU Pro)

Claude Opus 4.6 leads the reasoning benchmarks that measure graduate-level scientific and academic knowledge:

  • GPQA Diamond: Opus 4.6 at 77.3%. Gemini 3.1 Pro matches GPT-5.4 Pro at 94.3% on this benchmark at a fraction of the cost — a significant data point.
  • MMLU Pro: Opus 4.6 at 85.1%, leading the field on this broad multi-discipline test.
  • TAU-bench: Opus 4.6 leads this tool-use and autonomous reasoning benchmark.

APEX-Agents (Professional Services)

GPT-5.4 is #1 on the APEX-Agents benchmark, which measures performance on professional services work. According to the benchmark creators: "It excels at creating long-horizon deliverables such as slide decks, financial models, and legal analysis, delivering top performance while running faster and at a lower cost than competitive frontier models."

Coding Performance

Coding is where both models have the most to prove, and the picture here is nuanced.

SWE-bench Verified

SWE-bench Verified tests AI on real GitHub issues from open-source Python projects. The current leaderboard:

  • Claude Opus 4.5: 80.9% (still #1)
  • Claude Opus 4.6: 80.8%
  • MiniMax M2.5: 80.2%
  • GPT-5.2: 80.0%
  • Claude Sonnet 4.6: 79.6%
  • GPT-5.3-Codex: leads SWE-Bench Pro at 56.8% (different, harder benchmark)
  • Gemini 3 Pro: 76.2%
Field note

GPT-5.4's SWE-bench Verified score has not yet been published. Given GPT-5.4 incorporates GPT-5.3-Codex's capabilities, expect a significant jump from GPT-5.2's 80.0% score. We'll update this table when official results are available.

There's an important context note: OpenAI has stopped reporting SWE-bench Verified scores and now recommends SWE-bench Pro as the more reliable benchmark, citing concerns that 59.4% of the hardest unsolved problems had flawed test cases and that models could reproduce verbatim test patches. On SWE-bench Pro, GPT-5.3-Codex leads at 56.8%.

Real-World Coding: What Users Report

In independent 48-hour tests across rapid-fire coding challenges, Claude Opus 4.6 scored 220/220 on one developer's benchmark set — a perfect score. That said, GPT-5.4's strength is in large-scale, throughput-sensitive coding where speed and token efficiency matter. OpenAI says GPT-5.4's Tool Search system reduced token usage by 47% while maintaining accuracy on tool-heavy workloads — a major efficiency win for production pipelines.

Computer Use & Agentic Tasks

Both models now support direct computer interaction — clicking, typing, navigating interfaces — but with different approaches and track records.

  • Claude Opus 4.6: 72.5% on OSWorld computer use. Has had full computer-use support since launch, including a build-run-verify-fix loop. Agent teams (released simultaneously) let multiple Claude instances work in parallel.
  • GPT-5.4: Debuts native computer use for the first time in a mainline OpenAI model. Claims record scores on OSWorld-Verified and WebArena Verified, though specific percentages were not published at launch.
  • GPT-5.4 Thinking (upfront plans): Users can see and adjust the model's thinking plan mid-response before it finalizes its output — a workflow improvement with no Claude equivalent yet.

Context Window

  • Claude Opus 4.6: 200K standard, 1M token context in beta. 128K max output.
  • GPT-5.4: 1M token context in the API (experimentally). Tokens over 272K are billed at double the rate.
  • ChatGPT (consumer): 400K context for GPT-5.2; GPT-5.4 context limits in the ChatGPT app not yet published.

Pricing

OpenAI described GPT-5.4 as coming at their "highest per-token price yet." Exact API pricing at press time:

  • Claude Opus 4.6: $5 per million input tokens / $25 per million output tokens
  • GPT-5.4: Full pricing not yet published. Expected to exceed GPT-5.2's rates, which ran ~$10/$30 per million tokens.
  • GPT-5.3-Codex (for comparison): $3/$15 per million tokens via Codex app — cheaper than Opus at near-equivalent SWE-bench performance.
  • Claude Sonnet 4.6: $3/$15 per million tokens — exceptional value at 79.6% SWE-bench Verified.
Field note

Value play: If you're choosing purely on cost/coding performance, Claude Sonnet 4.6 at $3/$15 per million tokens with 79.6% SWE-bench Verified is arguably the best deal in the market right now — beating GPT-5.2 at a fraction of projected GPT-5.4 pricing.

Safety Philosophy

Both companies have invested heavily in safety, but with distinctly different approaches. Anthropic built Claude's safety constraints through Constitutional AI — safety baked into model training. Claude consistently refuses harmful requests more reliably in red-team evaluations.

OpenAI classifies GPT-5.4 as "High capability" for cybersecurity tasks, acknowledging its dual-use potential. They've responded by launching a Trusted Access for Cyber framework with a $10M fund to promote AI-powered cyber defense. For enterprise and regulated industries, Claude's more conservative safety posture remains a meaningful differentiator.

Verdict: Which Should You Use?

No single model wins everything in March 2026. The smartest approach is task-routing:

  • Choose Claude Opus 4.6 for deep reasoning, academic and scientific tasks, quality-focused code generation, long-document analysis, safety-sensitive applications, and agentic workflows with agent teams.
  • Choose GPT-5.4 for professional knowledge work (especially spreadsheets and financial models), computer-use automation, speed-critical large-scale coding pipelines, and scenarios where the upfront thinking plan feature saves iterations.
  • Consider Claude Sonnet 4.6 if budget is a concern — it delivers near-Opus coding performance at Sonnet pricing, making it the best value-per-dollar in the frontier tier.
  • Consider GPT-5.4 Thinking for Pro users if you're on the $200/month ChatGPT Pro plan and do heavy document + agentic work — it now represents a genuine Opus competitor.

The competitive gap between OpenAI and Anthropic has narrowed meaningfully with this release. The Chatbot Arena ELO for GPT-5.4 will be the next definitive data point — expect those rankings to emerge over the next 2–4 weeks as user votes accumulate.

Related essay
Claude vs ChatGPT (March 2026): Benchmarks, Pricing, Verdict
Related essay
Chatbot Arena Leaderboard March 2026: Full Rankings & Analysis
Related essay
SWE-Bench 2026 Leaderboard: Claude, ChatGPT & Gemini Coding Performance Ranked
Related essay
Claude 4 Arrives: Everything You Need to Know About Anthropic's Most Capable Model

THE CLAUDE INSIDER

Get the Claude playbook in your inbox.

One weekly email for Claude and Claude Code users. Real workflows, no hype. Subscribe and we send you The Claude Power-User Cheatsheet.

GUIDES AND COMPARISONS

— ¶ —

Luke Thompson

Luke Thompson

Editor-in-Chief · The Claude Insider

Luke Thompson is the founder of The Operations Guide, LLC and editor of The Claude Insider. Based in Jonesborough, Tennessee, he has spent years building AI-augmented business systems and automation workflows for operators and teams. He began working with large language models in production well before the current wave of consumer AI tools, integrating them into client workflows, content pipelines, and operational infrastructure. At The Claude Insider, he writes about Claude with the specificity of someone who uses it daily as a professional tool — not as a reviewer or commentator, but as a builder. His coverage focuses on what actually works: prompt patterns, API integration strategies, agentic workflows, and the real-world tradeoffs that practitioners face. He is not affiliated with Anthropic, PBC.

Articles are researched and drafted with AI assistance, reviewed and edited by Luke Thompson.

From Reading to Action

Know where AI can pay off in your company.

Take the free two-minute AI Readiness Assessment. See your score, the two gaps holding you back, and the next move worth making.

Get your readiness score

Instant report · No account to start

Related reading

View archive →