Essay/Claude·Apr 12, 2026

SWE-Bench 2026 Leaderboard: Claude vs GPT vs Gemini

A source-labeled SWE-bench leaderboard showing current Pro scores, reported Verified results, and why harness differences matter more than a single rank.

Luke Thompson
Luke ThompsonApr 12, 2026 · 7 min read
In this article
SWE-Bench 2026 Leaderboard: Claude vs GPT vs Gemini
GPT-5.4 currently leads Claude Opus 4.6 and Gemini 3.1 Pro among those three model families on Scale AI’s public SWE-bench Pro leaderboard. SWE-bench Verified tells a different—and less useful—story: vendor-reported frontier scores are packed near 80%, evaluation setups differ, and OpenAI now says the benchmark no longer measures frontier coding capability reliably. This page separates the two benchmarks and labels every score by source and setup. That is the only responsible way to use a coding leaderboard when choosing between Claude, GPT, and Gemini.
Field note

Checked July 11, 2026. Leaderboards change. The Pro table below reflects Scale AI’s public leaderboard on that date; the Verified table reproduces results reported in Anthropic’s Claude Opus 4.6 system card. Follow the primary-source links before making a procurement decision.

The short answer

  • SWE-bench Pro leader among GPT, Claude, and Gemini: GPT-5.4 (xHigh), 59.1%. Scale marks the entry with an asterisk; review the leaderboard notes and exact setup.
  • Highest Claude entry in that comparison: Claude Opus 4.6 (thinking), 51.9%. Agent, reasoning setting, and harness affect the result.
  • Highest Gemini entry in that comparison: Gemini 3.1 Pro (thinking), 46.1%. Do not compare this directly with a Verified score.
  • SWE-bench Verified leader: There is no durable model-only answer. The official leaderboard includes systems and scaffolds; vendor launch results may use different setups.
  • Better frontier buying signal: Pro is the better starting point, but your repository acceptance test matters more than either leaderboard.

SWE-bench Pro leaderboard: GPT vs Claude vs Gemini

SWE-bench Pro is designed around longer, more difficult software-engineering tasks across public repositories. Scale says it addresses contamination, narrow task coverage, oversimplified problems, and unreliable testing found in earlier benchmarks. Its lower scores leave more room to distinguish frontier systems.

The following is a focused comparison of the leading GPT, Claude, and Gemini entries visible on Scale’s public leaderboard. It is not a claim that these are the top entries across every provider.

  • 1. GPT-5.4 (xHigh) — OpenAI: 59.1% resolved.*
  • 2. Claude Opus 4.6 (thinking) — Anthropic: 51.9% resolved.*
  • 3. Gemini 3.1 Pro (thinking) — Google: 46.1% resolved.*
  • 4. Claude Opus 4.5 — Anthropic: 45.89% resolved.
  • 5. Claude Sonnet 4.5 — Anthropic: 43.6% resolved.
  • 6. Gemini 3 Pro Preview — Google: 43.3% resolved.
  • 7. GPT-5.2 Codex — OpenAI: 41.04% resolved.

On this evaluation, GPT-5.4’s 59.1% is 7.2 percentage points above Claude Opus 4.6 and 13 points above Gemini 3.1 Pro. That is evidence about these submitted systems on this public test—not proof that GPT will produce the best accepted change in every repository.

The labels matter. “Thinking,” “xHigh,” and the asterisks identify configuration or leaderboard qualifications that belong with the score. Removing them turns a system result into a misleading model claim.

SWE-bench Verified results: useful history, weak frontier ranking

SWE-bench Verified contains 500 human-reviewed issues from open-source Python repositories. The official SWE-bench site now provides a bash-only view that runs language models through a shared mini-SWE-agent environment, plus a broader leaderboard covering different agent systems. Those views answer different questions.

Anthropic’s Claude Opus 4.6 system card reports the following Verified results in one comparison table. Treat this as vendor-reported evaluation evidence, not a live reproduction of the official leaderboard.

  • Claude Opus 4.5: 80.9% reported SWE-bench Verified.
  • Claude Opus 4.6: 80.8% reported SWE-bench Verified.
  • GPT-5.2: 80.0% reported SWE-bench Verified.
  • Claude Sonnet 4.5: 77.2% reported SWE-bench Verified.
  • Gemini 3 Pro: 76.2% reported SWE-bench Verified.

A spread of only 4.7 points separates the five entries. More importantly, the benchmark is near saturation and its remaining cases do not cleanly measure model limits. In February 2026, OpenAI said it had stopped using Verified for frontier launches and recommended SWE-bench Pro.

Why OpenAI stopped reporting Verified

OpenAI audited 138 Verified problems that one of its models did not consistently solve. It reported that at least 59.4% of that audited subset contained material issues in test design or problem description. Correct solutions could fail narrow tests, while broad tests could reward solutions that did not fully address the issue. OpenAI also identified contamination risk.

That finding does not make every Verified score worthless. Verified still records a meaningful period of coding-agent progress, and the official bash-only setup can support controlled comparisons. It does mean a one-decimal difference near 80% should not drive a purchase or support a headline declaring one model universally best.

The harness can change the winner

SWE-bench evaluates a working system. The model receives an issue and a repository, but an agent loop decides how it inspects files, runs commands, edits code, uses context, retries, and spends tokens. A stronger scaffold can outperform a stronger raw model; a larger reasoning budget can improve the score while increasing cost and latency.

For a fair comparison, hold the dataset, agent, prompt, tool permissions, runtime, retry policy, reasoning setting, and pass criteria constant. If any of those change, label the result as a system comparison. The official SWE-bench bash-only view is valuable precisely because it attempts to standardize the environment for language-model comparisons.

How to read a changing rank

A leaderboard position can change because a provider released a better model, submitted a stronger agent, increased its reasoning budget, corrected an evaluation, or because another entrant was added. The rank alone does not reveal which event occurred. Save the model identifier, score, configuration label, benchmark version, source URL, and access date together.

Also distinguish percentage points from percent improvement. Moving from 50% to 55% is a five-point increase and a 10% relative improvement. Neither number tells you whether the extra solved tasks resemble your work or justify additional latency and cost. Inspect per-task results when available, then reproduce the comparison on your own acceptance set.

What the leaderboard means for Claude Code buyers

The Pro ranking says GPT-5.4 deserves a place in any frontier coding evaluation. Claude Opus 4.6 and Gemini 3.1 Pro also solve substantial portions of a difficult public benchmark. It does not tell you which product will produce the most reviewable pull requests in your codebase.

Claude Code, Codex, Gemini’s coding tools, IDE agents, and custom harnesses expose different permissions, context management, review surfaces, integrations, and billing. A model that scores lower in one public scaffold can still win when its product workflow reduces reviewer effort or avoids regressions.

Run a repository acceptance test

Two-week evaluation

Bottom line

For the current public SWE-bench Pro results, GPT-5.4 leads the named Claude and Gemini entries, followed by Claude Opus 4.6 and Gemini 3.1 Pro. For SWE-bench Verified, resist the temptation to crown a winner: the scores are compressed, setups vary, and OpenAI’s audit found serious limitations at the frontier.

Use Pro to create a shortlist. Use a controlled, repository-specific acceptance test to choose the system. Preserve the model version, agent, reasoning budget, date, and source beside every score so the comparison remains auditable when the leaderboard changes.

Primary sources

SWE-bench, “SWE-bench Verified” https://www.swebench.com/verified.html

Scale AI, “SWE-bench Pro (Public Dataset)” https://labs.scale.com/leaderboard/swe_bench_pro_public

OpenAI, “Why SWE-bench Verified no longer measures frontier coding capabilities” https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Anthropic, “Claude Opus 4.6 System Card” https://www-cdn.anthropic.com/14e4fb01875d2a69f646fa5e574dea2b1c0ff7b5.pdf

Related essay
Claude Code vs Codex CLI: Which Coding Agent Fits?
Related essay
Claude vs Gemini in 2026: Workflows, Coding, Long Context
Related essay
Claude Code in VS Code: Setup, Workflows, and Safety

THE CLAUDE INSIDER

Get the Claude playbook in your inbox.

One weekly email for Claude and Claude Code users. Real workflows, no hype. Subscribe and we send you The Claude Power-User Cheatsheet.

GUIDES AND COMPARISONS

— ¶ —

Luke Thompson

Luke Thompson

Editor-in-Chief · The Claude Insider

Luke Thompson is the founder of The Operations Guide, LLC and editor of The Claude Insider. Based in Jonesborough, Tennessee, he has spent years building AI-augmented business systems and automation workflows for operators and teams. He began working with large language models in production well before the current wave of consumer AI tools, integrating them into client workflows, content pipelines, and operational infrastructure. At The Claude Insider, he writes about Claude with the specificity of someone who uses it daily as a professional tool — not as a reviewer or commentator, but as a builder. His coverage focuses on what actually works: prompt patterns, API integration strategies, agentic workflows, and the real-world tradeoffs that practitioners face. He is not affiliated with Anthropic, PBC.

Articles are researched and drafted with AI assistance, reviewed and edited by Luke Thompson.

From Reading to Action

Know where AI can pay off in your company.

Take the free two-minute AI Readiness Assessment. See your score, the two gaps holding you back, and the next move worth making.

Get your readiness score

Instant report · No account to start

Related reading

View archive →