As of: Apr 2026Quarterly-refreshed

Claude 3.5 Sonnet vs GPT-4o: Version-by-Version Benchmark Comparison

Claude 3.5 Sonnet beat GPT-4o on coding at launch in June 2024 — the latest versions in both families have since shipped. This page tracks every release.

Latest versions (Apr 2026): Claude Sonnet 4.5 (Sep 2025) vs GPT-5 (May 2025). For an always-current breakdown, check our full AI models directory at zerotwo.ai/models.
TL;DR

Claude 3.5 Sonnet vs GPT-4o: When Claude 3.5 Sonnet launched in June 2024, it scored 92.0% on HumanEval vs GPT-4o's 90.2%— the first time a Claude model clearly outpaced OpenAI's flagship on coding. Both families have shipped multiple versions since then; the newest Claude Sonnet 4.5 and GPT-5 are the current benchmarks. Claude wins on cost and coding; GPT-4o's successors lead on multimodal and math. Use this page to compare every version honestly, then test them side by side on ZeroTwo.

Release Timeline: Every Claude Sonnet and GPT-4o Version

Both model families have shipped multiple revisions since their flagship launches. Claude 3.5 Sonnet set the coding standard in June 2024; GPT-4o followed with audio and vision upgrades through late 2024. Track the full progression below.

Anthropic / Claude Sonnet
1
Claude 3.5 Sonnet
Jun 2024
Beats GPT-4o on HumanEval at launch
2
Claude 3.5 Sonnet (new)
Oct 2024
Raises SWE-bench to 49.0%
3
Claude 3.7 Sonnet
Feb 2025
Extended thinking; reasoning leap
4
Claude Sonnet 4
May 2025
Strong coding + multimodal upgrade
5
Claude Sonnet 4.5
Sep 2025
Latest — agentic task leader
OpenAI / GPT-4o family
1
GPT-4o (May)
May 2024
Multimodal flagship launch
2
GPT-4o (Aug)
Aug 2024
Structured output improvements
3
GPT-4o (Nov)
Nov 2024
Vision + audio refinements
4
GPT-4.1
Apr 2025
Instruction-following & tool use
5
GPT-5
May 2025
Latest — multimodal reasoning flagship

Sources: Anthropic model card · OpenAI GPT-4o launch

Need to switch between versions mid-conversation? Browse all available Claude Sonnet and GPT-4o versions on ZeroTwo — pick the exact release you want in one click.

Benchmark Scorecards: Coding, Reasoning, Multimodal, Cost

Four categories where Claude 3.5 Sonnet and GPT-4o differ most. Scores from official Anthropic and OpenAI model cards, June 2024 release versions.

HumanEval
Coding
Claude92%
GPT-4o90.2%
Claude leads on function-level code generation
Claude leads
GPQA Diamond
Reasoning
Claude59.4%
GPT-4o53.6%
Claude edges GPT-4o on graduate-level science
Claude leads
MMMU
Multimodal
Claude68.3%
GPT-4o69.1%
GPT-4o has a slight multimodal edge
GPT-4o leads
$/M tokens
Cost
Claude$3 input
GPT-4o$5 input
Claude 3.5 Sonnet is 40% cheaper on input tokens
Claude leads

Run the same prompt through Claude and GPT-4o — right now

ZeroTwo lets you chat with Claude Sonnet 4.5 and GPT-5 side by side in one window. No separate subscriptions. Free to start.

Start free comparison

"Claude 3.5 Sonnet is the best available LLM right now... it beats GPT-4o on almost every benchmark, and is available at a fraction of the price of Claude 3 Opus."

How Does Claude 3.5 Sonnet vs GPT-4o Compare on Coding?

Claude 3.5 Sonnet leads on code generation. It scored 92.0% on HumanEval — the most-cited coding benchmark — versus GPT-4o's 90.2% at launch, according to Anthropic's model card. The gap is most visible on SWE-bench Verified, a real-world GitHub issue-resolution benchmark: Claude 3.5 Sonnet (New) reached 49.0% vs GPT-4o's ~33%.

For developers using AI coding tools, the practical difference shows up in tasks like writing test suites, refactoring large modules, and multi-file context reasoning. Claude consistently produces tighter, fewer-hallucination code blocks in controlled evaluations on Aider's LLM coding leaderboard.

Practical tip: Both models perform well on standard LeetCode-style problems. The Claude advantage is most pronounced on longer, multi-function generation tasks and autonomous bug-fixing (SWE-bench). If you need the best AI for coding tasks, Claude Sonnet is the current leader.

Reasoning and Knowledge: MMLU, GPQA, MATH

Both models tied on MMLU at 88.7% each — the graduate-level multitask exam considered a proxy for general knowledge breadth. Claude 3.5 Sonnet pulled ahead on GPQA Diamond (59.4% vs GPT-4o's 53.6%), a harder science PhD–level reasoning benchmark. GPT-4o maintained an edge on MATH (76.6% vs 71.1%), particularly on formal proof-style problems.

For most professional use cases — legal research, medical literature review, financial analysis — the 88.7% MMLU tie suggests comparable general knowledge. GPQA Diamond scores matter for domain experts in hard sciences. MATH scores matter for quantitative researchers and engineers. LMSYS Chatbot Arena human preference voting has historically placed both models in a statistical tie for overall quality.

Pricing: Claude 3.5 Sonnet Is 40% Cheaper on Input

At launch, Claude 3.5 Sonnet cost $3.00 per million input tokens vs GPT-4o's $5.00 — a 40% difference. Output tokens were identical at $15.00/M. For teams running high-volume inference, this difference compounds quickly.

ModelInput $/MOutput $/MNote
Claude 3.5 Sonnet$3.00$15.00Anthropic API, Jun 2024
GPT-4o (May 2024)$5.00$15.00OpenAI API, May 2024
GPT-4.1 (Apr 2025)$2.00$8.00OpenAI reduced prices
GPT-5 (May 2025)$10.00$30.00Premium flagship pricing

Running cost comparisons across model versions? Use ZeroTwo's live AI pricing comparison tool to see per-token and per-subscription costs side by side.

Full Version-by-Version Benchmark Matrix

All versions across 7 benchmarks. Newer Claude/GPT values marked with ~ are estimated from official announcements and independent evaluations. Artificial Analysis and LMSYS Arena are the canonical live sources.

ModelHumanEvalSWE-benchMMLUGPQAMATHMMMUInput $/MOutput $/M
Claude 3.5 Sonnet (Jun 2024)92.0%49.0%*88.7%59.4%71.1%68.3%$3.00$15.00
Claude 3.5 Sonnet New (Oct 2024)92.0%49.0%88.7%59.4%71.1%68.3%$3.00$15.00
Claude 3.7 Sonnet (Feb 2025)~93%62.3%~89%~68%~78%~70%$3.00$15.00
Claude Sonnet 4 (May 2025)~94%~72%~90%~72%~82%~73%$3.00$15.00
Claude Sonnet 4.5 (Sep 2025)~95%~78%~91%~74%~84%~75%$3.00$15.00
GPT-4o (May 2024)90.2%~33%88.7%53.6%76.6%69.1%$5.00$15.00
GPT-4o (Aug 2024)90.2%~36%88.7%53.6%76.6%69.1%$5.00$15.00
GPT-4o (Nov 2024)90.2%~38%88.7%53.6%76.6%69.1%$5.00$15.00
GPT-4.1 (Apr 2025)~92%~55%~90%~60%~80%~71%$2.00$8.00
GPT-5 (May 2025)~96%~80%~92%~76%~88%~80%$10.00$30.00

* SWE-bench Verified. ~ indicates estimate from public announcements. Prices are API list prices at launch date.

Which Is Better Today? An Honest Verdict (Apr 2026)

Based on the full version history: Claude's Sonnet line leads on autonomous coding and is cheaper at comparable capability tiers. GPT's newer models (GPT-4.1, GPT-5) match or exceed Claude on multimodal tasks and mathematics.

Pick Claude if…
  • → Writing or reviewing code autonomously (SWE-bench lead)
  • → Cost is a primary constraint (40% cheaper input)
  • → Graduate-level reasoning matters (GPQA edge)
  • → You prefer longer, structured outputs with fewer hallucinations
  • → You want agentic task completion (Sonnet 4.5 is the current leader)
Pick GPT-4o/GPT-5 if…
  • → Multimodal workflows: images, audio, real-time voice
  • → Math-heavy tasks (MATH benchmark advantage)
  • → Deep OpenAI ecosystem integration (Assistants API, DALL·E, etc.)
  • → You need the absolute latest flagship: GPT-5
  • → You want structured output / JSON mode out of the box

7 Benchmark Statistics Worth Knowing

92.0%
Claude 3.5 Sonnet HumanEval
Source: Anthropic model card, Jun 2024
90.2%
GPT-4o HumanEval
Source: OpenAI model card, May 2024
88.7%
MMLU — both models tied
Source: Anthropic + OpenAI, 2024
59.4%
Claude 3.5 Sonnet GPQA Diamond
Source: Anthropic model card
53.6%
GPT-4o GPQA Diamond
Source: OpenAI model card
76.6%
GPT-4o MATH (vs Claude 71.1%)
Source: Model cards
$3 vs $5
Input price per million tokens (Claude vs GPT-4o)
Source: API pricing at launch

Frequently Asked Questions

Key Takeaways

  • 1Claude 3.5 Sonnet scored 92.0% on HumanEval vs GPT-4o's 90.2% at launch in June 2024 — the first clear coding win for Anthropic against OpenAI's flagship.
  • 2Both models tied on MMLU at 88.7%; Claude leads GPQA Diamond (59.4% vs 53.6%); GPT-4o leads MATH (76.6% vs 71.1%).
  • 3Claude 3.5 Sonnet costs 40% less on input tokens ($3/M vs $5/M) while matching GPT-4o's output pricing.
  • 4Since launch, both families have shipped major revisions — Claude Sonnet 4.5 and GPT-5 are the current flagships as of April 2026.
  • 5For most coding and cost-sensitive workloads, Claude Sonnet is the better default; for multimodal and voice-heavy tasks, GPT-4o's successors hold the edge.

Related Comparisons

Author: ZeroTwo AI Research Team — AI model comparisons and benchmark analysis specialists
Published: Nov 1, 2024Updated: Apr 22, 2026

Test Claude Sonnet and GPT-4o yourself

ZeroTwo gives you every Claude Sonnet version and every GPT-4o variant in one place. No credit card required to start.

Start free — no credit card