Part 1 — Knowledge benchmark (MMLU), 2022–2026
MMLU was the dominant benchmark from 2022 to 2025, tracking broad knowledge across 57 academic disciplines. Neither lab reports it in launch materials any more: both sit above 96% and the spread between them is inside the noise. The historical arc is still the clearest picture of how fast the gap closed — Claude entered 2023 thirteen points behind GPT-4 and had caught up by mid-2024.
+27.8pp
OpenAI MMLU gain
GPT-3.5 → GPT-6 Astra (3.8 yrs)
+24.6pp
Claude MMLU gain
Claude 1 → Claude Fable 5.1
8 releases
Anthropic 2026 cadence
Opus 4.6 → Fable 5.1, Feb–Sep 2026
ARC-AGI-3
Widest live gap
99.9% vs 30.2% — reversed in September
OpenAI / ChatGPTAnthropic / ClaudeShaded region = estimated, no longer reported
MMLU BENCHMARK (%) — HIGHER IS BETTER
OpenAI / ChatGPT
Sep '26GPT-6 AstraLATEST
99.9% ARC-AGI-3 on a stateful harness; 72.6% OSWorld 2.0; 97.6% FrontierMath T4
97.8% est Jul '26GPT-5.6 Sol
Sol / Terra / Luna tiers; 96.2% SWE-bench Verified; $5/$30 per MTok
97.4% est Apr '26GPT-5.5
Codename 'Spud'; first fully-retrained base since GPT-4.5
96.5% est Mar '26GPT-5.4
Native computer use; 1M context; merges the Codex line
97.2% est Feb '26GPT-5.2
Improved polish; spreadsheet & finance tasks
96.8% est Nov '25GPT-5.1
Adaptive reasoning; 8 personality presets
96.4% May '25GPT-5
Unified reasoning + conversational AI
96% Apr '25o3 / o4-mini
Top AIME 2024/25 benchmark scores
96.7% Jan '25o3-mini
Compact, fast reasoning
93% Dec '24o1
Full o1 + $200/mo Pro tier
92.3% Sep '24o1-preview
Chain-of-thought reasoning model
90.8% May '24GPT-4o
Omni-modal; free-tier access
88.7% Nov '23GPT-4 Turbo
128K context; updated knowledge cutoff
87.4% Mar '23GPT-4
Major leap; launched in Bing and ChatGPT
86.4% Nov '22GPT-3.5
ChatGPT launch — 1M users in 5 days
70% Anthropic / Claude
Sep '26Claude Fable 5.1LATEST
81.2% SWE-bench Pro; cache reads cut 4x to $0.25/MTok; Mythos 5.1 is the same model, restricted
97.6% est Jul '26Claude Opus 5
96.0% SWE-bench Verified; 1M context; Opus pricing held at $5/$25
97.5% est Jun '26Claude Sonnet 5
GA 30 Jun; 85.2% SWE-bench Verified at $2/$10 per MTok
96.4% est Jun '26Claude Fable 5
95.0% SWE-bench Verified, 80.0% Pro; export-control pause 12 Jun – 1 Jul
97.2% est May '26Claude Opus 4.8
88.6% SWE-bench Verified; 41-day release cadence
96.9% est Apr '26Claude Opus 4.7
87.6% SWE-bench Verified; xhigh effort tier
96.7% est Mar '26Claude Sonnet 4.6
Cost-efficient tier; 12 days after Opus 4.6
96.2% est Feb '26Claude Opus 4.6
1M context beta; adaptive thinking; agent teams
96.5% est Nov '25Claude 4.5 Opus
First model >80% SWE-bench (80.9%)
95.8% Sep '25Claude 4.5 Sonnet
77.2% SWE-bench; 30+ hour task focus
95% May '25Claude 4 Opus
ASL-3 safety classification; Claude Code
94.5% Feb '25Claude 3.7 Sonnet
First hybrid reasoning model
91% Oct '24Claude 3.5 Sonnet v2
Computer use capability; upgraded Haiku
89.5% Jun '24Claude 3.5 Sonnet
Beats Opus at Sonnet price; Artifacts
88.7% Mar '24Claude 3 Opus
Multimodal; noted self-awareness in tests
86.8% Nov '23Claude 2.1
200K context; reduced hallucination rate
80% Jul '23Claude 2
General public access; 100K context
78.5% Mar '23Claude 1
First public Claude via limited API
73% Part 2 — Where the frontier is contested (September 2026)
MMLU measures what a model knows. These measure what it can do: ship a repository-level change, drive a terminal, operate a desktop, and reason about problems it has never seen. Both current flagships — OpenAI's GPT-5.6 Sol (9 July) and Anthropic's Claude Opus 5 (24 July) — cleared 96% on SWE-bench Verified, so the live comparison has moved to SWE-bench Pro, OSWorld 2.0, Frontier-Bench and ARC-AGI-3. Hover any bar for the source note.
Benchmarks moved under the models this year. OSWorld went to the harder 2.0 set, Terminal-Bench to 2.1, SWE-bench to Pro, and GDPval-AA to v2. Scores below are not comparable to the 2025 numbers that carried the same benchmark names.
OpenAIAnthropicSCORE (%) — HIGHER IS BETTER
GDPval-AA v2
Economically valuable knowledge work across 44 professions, scored as an Elo rating
Opus 5 leads by 125 Elo — roughly a 67% expected win rate head-to-head. Elo is an interval scale and is not plotted alongside the percentage benchmarks. Bars start at 1,650 Elo, not zero.
SWE-bench Verified — saturated
Reported for reference, not plotted: the spread no longer separates the frontier.
GPT-5.6 Sol96.2%
Claude Opus 596.0%
Claude Fable 595.0%
Claude Sonnet 585.2%
Four points separate the two flagships. Verified is finished as a differentiator; the contested version is Pro. Neither September release adds a row — Anthropic published no Verified score for Fable 5.1, and the widely-quoted 95.0% belongs to Fable 5.
SWE-bench Pro
Repository-level software engineering — harder and less contaminated than Verified
GPT-5.6 Sol64.6%
Claude Fable 5.181.2% ▲
Fable 5.1 (Sep 2026) takes the Anthropic figure to 81.2, up from 79.2 for Opus 5. OpenAI published no SWE-bench Pro score for GPT-6 Astra, so the OpenAI bar is still GPT-5.6 Sol and the gap shown is not a like-for-like September comparison.
Terminal-Bench
Agentic coding in real terminal environments — planning, iteration, tool coordination
GPT-5.6 Sol89.5% ▲
Claude Opus 589.1%
Effectively a tie (0.4 points, Sol at xhigh vs Opus 5 at max effort). GPT-5.6 Sol Ultra reaches 91.9%. Not updated for September: Fable 5.1 reports against Terminal-Bench 4.0 (55.8%) and Terminal-Bench-Science (52.6%), which are different, harder sets and cannot be put on this bar.
OSWorld 2.0
Autonomous desktop navigation via screenshots and keyboard/mouse
GPT-6 Astra72.6%
Claude Fable 5.177.9% ▲
Both September flagships, but read the settings before reading the gap: Fable 5.1 scores 77.9 on the partial setting and only 41.7 strict, and Anthropic ran with production safeguards on, taking zeros where a classifier intervened. OpenAI reports 72.6 for Astra against 65.7 for Sol. OSWorld moved to the harder 2.0 set in 2026, so none of this is comparable to the ~80% figures on OSWorld-Verified.
Frontier-Bench
Long-horizon agentic coding on unsolved problems
GPT-5.6 Sol34.4%
Claude Opus 543.3% ▲
Both labs are under 50%. The newest benchmark on this page and the least saturated.
ARC-AGI-3
Novel reasoning on problems absent from any training set
GPT-6 Astra99.9% ▲
Claude Opus 530.2%
Astra saturates ARC-AGI-3 at 99.9% and turns the widest gap on this page completely around — but the headline number depends on a stateful, expensive harness, and stateless API calls score far lower. Anthropic has published no ARC-AGI-3 figure for Fable 5.1, so the Anthropic bar is still Opus 5. Treat this pairing as the least settled on the page.
Benchmark notes: MMLU is official where published. Entries marked est (February 2026 onward, shaded on the chart) are estimates positioned from the ArtificialAnalysis Intelligence Index — neither lab reports MMLU in launch materials any more. Capability scores are from each model's launch post and the comparison round-ups that followed, and reflect each lab's highest published effort setting. They are not all from the same release: where a lab published no figure for a benchmark with its September model, the bar names the last model that did, and that benchmark's note says so. Anthropic's September figures were measured with production safeguards on, taking zeros where a classifier intervened, which makes them conservative. GDPval-AA v2 is an Elo rating and is shown on its own scale rather than on the percentage axis. Benchmark versions changed during 2026 — OSWorld 2.0, Terminal-Bench 2.1, SWE-bench Pro, GDPval-AA v2 — so these are not comparable with earlier figures published under the same names. Last updated 11 September 2026.
Live tracking: llm-stats.com · Artificial Analysis · ARC Prize leaderboard