Claude 3.5 Sonnet vs GPT-4o: Version-by-Version Benchmark Comparison
Claude 3.5 Sonnet beat GPT-4o on coding at launch in June 2024 — the latest versions in both families have since shipped. This page tracks every release.
Claude 3.5 Sonnet vs GPT-4o: When Claude 3.5 Sonnet launched in June 2024, it scored 92.0% on HumanEval vs GPT-4o's 90.2%— the first time a Claude model clearly outpaced OpenAI's flagship on coding. Both families have shipped multiple versions since then; the newest Claude Sonnet 4.5 and GPT-5 are the current benchmarks. Claude wins on cost and coding; GPT-4o's successors lead on multimodal and math. Use this page to compare every version honestly, then test them side by side on ZeroTwo.
Release Timeline: Every Claude Sonnet and GPT-4o Version
Both model families have shipped multiple revisions since their flagship launches. Claude 3.5 Sonnet set the coding standard in June 2024; GPT-4o followed with audio and vision upgrades through late 2024. Track the full progression below.
Sources: Anthropic model card · OpenAI GPT-4o launch
Need to switch between versions mid-conversation? Browse all available Claude Sonnet and GPT-4o versions on ZeroTwo — pick the exact release you want in one click.
Benchmark Scorecards: Coding, Reasoning, Multimodal, Cost
Four categories where Claude 3.5 Sonnet and GPT-4o differ most. Scores from official Anthropic and OpenAI model cards, June 2024 release versions.
Run the same prompt through Claude and GPT-4o — right now
ZeroTwo lets you chat with Claude Sonnet 4.5 and GPT-5 side by side in one window. No separate subscriptions. Free to start.
Start free comparison"Claude 3.5 Sonnet is the best available LLM right now... it beats GPT-4o on almost every benchmark, and is available at a fraction of the price of Claude 3 Opus."
How Does Claude 3.5 Sonnet vs GPT-4o Compare on Coding?
Claude 3.5 Sonnet leads on code generation. It scored 92.0% on HumanEval — the most-cited coding benchmark — versus GPT-4o's 90.2% at launch, according to Anthropic's model card. The gap is most visible on SWE-bench Verified, a real-world GitHub issue-resolution benchmark: Claude 3.5 Sonnet (New) reached 49.0% vs GPT-4o's ~33%.
For developers using AI coding tools, the practical difference shows up in tasks like writing test suites, refactoring large modules, and multi-file context reasoning. Claude consistently produces tighter, fewer-hallucination code blocks in controlled evaluations on Aider's LLM coding leaderboard.
Practical tip: Both models perform well on standard LeetCode-style problems. The Claude advantage is most pronounced on longer, multi-function generation tasks and autonomous bug-fixing (SWE-bench). If you need the best AI for coding tasks, Claude Sonnet is the current leader.
Reasoning and Knowledge: MMLU, GPQA, MATH
Both models tied on MMLU at 88.7% each — the graduate-level multitask exam considered a proxy for general knowledge breadth. Claude 3.5 Sonnet pulled ahead on GPQA Diamond (59.4% vs GPT-4o's 53.6%), a harder science PhD–level reasoning benchmark. GPT-4o maintained an edge on MATH (76.6% vs 71.1%), particularly on formal proof-style problems.
For most professional use cases — legal research, medical literature review, financial analysis — the 88.7% MMLU tie suggests comparable general knowledge. GPQA Diamond scores matter for domain experts in hard sciences. MATH scores matter for quantitative researchers and engineers. LMSYS Chatbot Arena human preference voting has historically placed both models in a statistical tie for overall quality.
Pricing: Claude 3.5 Sonnet Is 40% Cheaper on Input
At launch, Claude 3.5 Sonnet cost $3.00 per million input tokens vs GPT-4o's $5.00 — a 40% difference. Output tokens were identical at $15.00/M. For teams running high-volume inference, this difference compounds quickly.
| Model | Input $/M | Output $/M | Note |
|---|---|---|---|
| Claude 3.5 Sonnet | $3.00 | $15.00 | Anthropic API, Jun 2024 |
| GPT-4o (May 2024) | $5.00 | $15.00 | OpenAI API, May 2024 |
| GPT-4.1 (Apr 2025) | $2.00 | $8.00 | OpenAI reduced prices |
| GPT-5 (May 2025) | $10.00 | $30.00 | Premium flagship pricing |
Running cost comparisons across model versions? Use ZeroTwo's live AI pricing comparison tool to see per-token and per-subscription costs side by side.
Full Version-by-Version Benchmark Matrix
All versions across 7 benchmarks. Newer Claude/GPT values marked with ~ are estimated from official announcements and independent evaluations. Artificial Analysis and LMSYS Arena are the canonical live sources.
| Model | HumanEval | SWE-bench | MMLU | GPQA | MATH | MMMU | Input $/M | Output $/M |
|---|---|---|---|---|---|---|---|---|
| Claude 3.5 Sonnet (Jun 2024) | 92.0% | 49.0%* | 88.7% | 59.4% | 71.1% | 68.3% | $3.00 | $15.00 |
| Claude 3.5 Sonnet New (Oct 2024) | 92.0% | 49.0% | 88.7% | 59.4% | 71.1% | 68.3% | $3.00 | $15.00 |
| Claude 3.7 Sonnet (Feb 2025) | ~93% | 62.3% | ~89% | ~68% | ~78% | ~70% | $3.00 | $15.00 |
| Claude Sonnet 4 (May 2025) | ~94% | ~72% | ~90% | ~72% | ~82% | ~73% | $3.00 | $15.00 |
| Claude Sonnet 4.5 (Sep 2025) | ~95% | ~78% | ~91% | ~74% | ~84% | ~75% | $3.00 | $15.00 |
| GPT-4o (May 2024) | 90.2% | ~33% | 88.7% | 53.6% | 76.6% | 69.1% | $5.00 | $15.00 |
| GPT-4o (Aug 2024) | 90.2% | ~36% | 88.7% | 53.6% | 76.6% | 69.1% | $5.00 | $15.00 |
| GPT-4o (Nov 2024) | 90.2% | ~38% | 88.7% | 53.6% | 76.6% | 69.1% | $5.00 | $15.00 |
| GPT-4.1 (Apr 2025) | ~92% | ~55% | ~90% | ~60% | ~80% | ~71% | $2.00 | $8.00 |
| GPT-5 (May 2025) | ~96% | ~80% | ~92% | ~76% | ~88% | ~80% | $10.00 | $30.00 |
* SWE-bench Verified. ~ indicates estimate from public announcements. Prices are API list prices at launch date.
Which Is Better Today? An Honest Verdict (Apr 2026)
Based on the full version history: Claude's Sonnet line leads on autonomous coding and is cheaper at comparable capability tiers. GPT's newer models (GPT-4.1, GPT-5) match or exceed Claude on multimodal tasks and mathematics.
- → Writing or reviewing code autonomously (SWE-bench lead)
- → Cost is a primary constraint (40% cheaper input)
- → Graduate-level reasoning matters (GPQA edge)
- → You prefer longer, structured outputs with fewer hallucinations
- → You want agentic task completion (Sonnet 4.5 is the current leader)
- → Multimodal workflows: images, audio, real-time voice
- → Math-heavy tasks (MATH benchmark advantage)
- → Deep OpenAI ecosystem integration (Assistants API, DALL·E, etc.)
- → You need the absolute latest flagship: GPT-5
- → You want structured output / JSON mode out of the box
7 Benchmark Statistics Worth Knowing
Frequently Asked Questions
Key Takeaways
- 1Claude 3.5 Sonnet scored 92.0% on HumanEval vs GPT-4o's 90.2% at launch in June 2024 — the first clear coding win for Anthropic against OpenAI's flagship.
- 2Both models tied on MMLU at 88.7%; Claude leads GPQA Diamond (59.4% vs 53.6%); GPT-4o leads MATH (76.6% vs 71.1%).
- 3Claude 3.5 Sonnet costs 40% less on input tokens ($3/M vs $5/M) while matching GPT-4o's output pricing.
- 4Since launch, both families have shipped major revisions — Claude Sonnet 4.5 and GPT-5 are the current flagships as of April 2026.
- 5For most coding and cost-sensitive workloads, Claude Sonnet is the better default; for multimodal and voice-heavy tasks, GPT-4o's successors hold the edge.