AI benchmarks

Every frontier AI model on every major benchmark

Frontier AI benchmark scores as of September 8, 2026: on ARC-AGI-2, GPT-6 Astra leads at 95%; on GPQA, GPT-6 Astra leads at 96%; on SWE Pro, Claude Fable 5.1 leads at 81.2%. Every score below cites the lab’s announcement post or an independent re-runner.

Last verified: September 8, 2026.

ARC-AGI-2 leader
95%
GPQA leader
96%
SWE Pro leader
81.2%
Coverage
8 × 7
frontier models × benchmarks

How to read this page

● 87.2lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.

□ 86.4independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.

⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).

— the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.

Headline matrix

Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.

Anthropic · September 1, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
DeepSeek · August 13, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Google · September 2, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
OpenAI · September 3, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Grok 4.6Closed
xAI · August 12, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Mistral AI · April 28, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Meta · September 2, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Alibaba · August 3, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified

By benchmark

Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.

ARC-AGI-2

ReasoningDiscriminating

Successor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. Two labs now name an ARC-AGI-2 figure in their own launch materials, and in both cases the number matches the ARC Prize Foundation's verified board rather than describing a separate run: Anthropic's Claude Fable 5.1 system card gives ARC-AGI-1 97.5 and ARC-AGI-2 90 and says whose numbers they are, and OpenAI's GPT-6 Astra post carries a full ARC-AGI-1/2/3 table whose 95.0 for Astra is the same figure the Foundation publishes. So every value in this column traces back to one measurement per model, taken from that board at each model's highest verified reasoning-effort variant. Coverage is four of the eight current flagships. That is thin against the two knowledge columns, which both carry all eight, but it now beats SWE-Bench Pro's three — a reversal from earlier in 2026, when the coding column was the most widely reported here and this one was the sparsest. GPT-6 Astra leads at 95.0, then Claude Fable 5.1 90.0, Grok 4.6 67.1 and DeepSeek-V4-Pro 61.3 — a 33.7-point band, the second-widest on this page. Astra's arrival widened it from 31.2 by lifting only the top, which is the pattern this column keeps producing. That spread is the thing to watch: the labs that publish nothing on this benchmark are the ones that score lowest when somebody else measures them. The four flagships without a submission are Gemini 3.8 Flash, Muse Spark 1.3, Mistral Medium 3.5 and Qwen3.8-Max. Google's gap is the fresh one — the retired Gemini 3.7 Flash held 84.6 and its successor has no submission — while the other three labs have never placed a model anywhere but the floor of this board, appearing only through older siblings: two Llama 4 builds and three Magistral builds at zero, and one Qwen3-235B build at 1.3. The board rebuilds irregularly rather than on a schedule — it sat frozen for eleven days across August, then regenerated three times in four days at the start of September — so a missing row means a model has not been run yet, never that it ran and scored nothing.

Author
François Chollet · ARC Prize Foundation
Human baseline
Every eval task was solved pass@2 by at least two people in a 400+ participant calibration study; ARC Prize sets 85% as the target a system must reach

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
OpenAI reports ARC-AGI-2 in its own launch materials for the first time on a flagship: the Astra post's Abstract reasoning table reads 95.0 for GPT-6 Astra, 92.5 for GPT-5.6 Sol, 90.4 for Claude Opus 5, 90.0 for Claude Fable 5.1, 89.2 for Claude Fable 5 and an em-dash for Gemini 3.8 Flash, alongside ARC-AGI-3 and ARC-AGI-1 rows. The 5.6 generation's abstract-reasoning slot had been ARC-AGI-3 alone. The 95.0 is identical to the ARC Prize Foundation's verified max-effort score in this cell's independent half, so the two halves describe one measurement rather than two — the same reading the page gives Claude Fable 5.1, whose system card restates the Foundation's number explicitly. OpenAI's tables are footed 'Evaluation scores are the maximum at any effort', which is the configuration the Foundation's Max row also describes.
Model
Score
Notes
Anthropic reports ARC-AGI for the first time on a Fable-class model: the system card's capability evaluation summary gives ARC-AGI-2 90% and ARC-AGI-1 97.5%, both at max effort on the semi-private validation sets, where the Fable 5 card carried no abstract-reasoning row of any kind. The card attributes the figures to the ARC Prize Foundation rather than to an Anthropic run — 'The ARC Prize Foundation reports that Fable 5.1 achieved a verified score of 97.5% on ARC-AGI-1 and 90% on ARC-AGI-2 at max effort' — so the lab and independent values in this cell describe one measurement, not two.
Model
Score
Notes
Grok 4.6's launch post reports no ARC-AGI benchmark. Its Evals table runs AA Intelligence Index, GDPVal-AA v2, CursorBench v3.2, DeepSWE v1.1, FrontierCode v1.1, APEX-Agents, Terminal-Bench v3.0, APEX-SWE, AA-Briefcase and Harvey LAB — nothing testing abstract reasoning.
Model
Score
Notes
DeepSeek reports no ARC-AGI benchmark of any kind. Neither of the model card's two benchmark tables carries one, and neither does the change-log entry announcing the model's general availability.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no ARC-AGI benchmark, as the 3.7 and 3.6 cards did not. Google's published eval set leads on agentic coding, computer use and enterprise knowledge workflows instead.
Model
Score
Not reported
Notes
Mistral reports no ARC-AGI benchmark for Medium 3.5.
Model
Score
Not reported
Notes
Meta's Muse Spark 1.3 launch post reports no ARC-AGI benchmark — and reports the one evaluation summary it does publish as a chart image rather than as text, so none of its figures can be quoted as a lab claim. The image's alt text stops at 'Benchmark scorecard comparing Muse Spark 1.3, Muse Spark 1.2, GPT 5.6 Sol (max), and Opus 5 (max) across agent, coding, instruction-following, and long-context evaluations', with no numbers in it. This page does not read numbers off pictures.
Model
Score
Not reported
Notes
Neither Alibaba's Qwen3.8-Max launch posts nor the 31-row benchmark table in the open-weights sibling's model card — the source of this row's other lab figures — reports any ARC-AGI benchmark.

Humanity's Last Exam

KnowledgeDiscriminating

Crowdsourced expert-level exam of 2,500 questions across more than a hundred subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools; this column records the no-tools figure. One of two benchmarks on this page with a published value for all eight current flagships — GPQA Diamond is the other — and much the wider band of the two: Claude Fable 5.1 tops the field at 60.9 and Mistral Medium 3.5 anchors it at 13.8, a 47.1-point spread, with the middle six clustered between 42.7 and 54.7. Some of that width is one straggler, but less than it used to be. Drop Mistral and the remaining seven still span 18.2 points, against 12.8 before September 2026, because two labs added points at the top in the same fortnight while the floor did not move: Anthropic's new flagship took the lead to 60.9, and GPT-6 Astra then entered 5.2 points above the retired GPT-5.6 Sol at 54.7. The shape that produced is a ladder rather than a cluster — leader, then Astra 6.2 points back, then a five-model band inside 6 points — so the approaching-saturation reading this column used to invite no longer holds at the top even though it still describes the middle. Three of the eight values are lab claims; the rest come from Artificial Analysis, which runs the full exam on every current flagship it charts. The column's first head-to-head between a lab and an independent evaluator on the same model landed in September 2026, and they agree: Anthropic's 60.9 against Artificial Analysis's 59.1, a 1.8-point gap well inside the divergence threshold. A second evaluator has since arrived on this column and it does not agree at all. Epoch AI's file had for months been a mirror of other people's published results with no current flagship on it; it now carries Claude Fable 5.1 at its extended-reasoning setting, scoring 46.5 and topping Epoch's own board — 12.2 points below what Artificial Analysis measures for the same model at the same setting, and 14.4 below what Anthropic reports for it. That is the widest evaluator-versus-evaluator gap anywhere on this page, twice the 6.7 points that separate the two boards on GPQA Diamond, and it is worth more attention than any of the lab-versus-independent gaps this page flags: the cells here are only comparable to each other because they nearly all come from one evaluator, and the moment a second one measures the same model the column's apparent precision turns out to be a house style. Two caveats on the other end of the table. Google publishes only HLE-Verified (3.8 Flash 54.9), a filtered 1,811-item subset rather than the full exam, so it stays in the cell note rather than in the cell — and the two instruments disagree about direction, since Google's subset shows a 1.3-point generational gain where the full-exam re-run has the new model a twentieth of a point behind its predecessor. Meta's figure is recorded at a max-effort setting Artificial Analysis has charted but Meta says is still pending safety testing.

Author
Center for AI Safety · Scale AI
Human baseline
No aggregate human score is published; the questions were written and cross-checked by close to a thousand expert contributors and are built to sit past the reach of any one specialist

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
Anthropic ran Humanity's Last Exam on Fable 5.1 itself, where the Fable 5 card had an em-dash in the Fable column and reported Mythos 5 alone. The capability summary gives 60.9% without tools and 65.0% with, against Fable 5's 57.8 and 63.8 and Claude Opus 5's 56.6 and 63.6; this column records the no-tools figure. Both configurations use adaptive thinking at max effort, averaged over five trials, with a 1M total-token cap and no context compaction. The with-tools run adds web search, a fetch tool restricted to URLs already in the conversation, programmatic tool calling and code execution, blocklists sources known to discuss the exam, and re-grades as incorrect any transcript found to have retrieved an answer.
Model
Score
Notes
OpenAI publishes an HLE figure for Astra but not the one this column tracks. The Academic table's row is headed 'Humanity's Last Exam (w/ tools)' and gives Astra 57.2, where this column records the no-tools figure, so it stays out of the cell — the same near-miss discipline that keeps Google's HLE-Verified 54.9 out of the Gemini cell. That the row is the with-tools instrument is checkable inside OpenAI's own table, which lists Claude Fable 5.1 at 65.0: exactly Anthropic's published with-tools number, against the 60.9 no-tools figure in this page's Anthropic cell. OpenAI has never published a no-tools HLE score for any model, and GPT-5.6 Sol's cell in this same row is an em-dash.
Model
Score
Notes
Meta published no Humanity's Last Exam figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text, so nothing here is quotable as a lab claim.
Model
Score
Notes
Google publishes no plain Humanity's Last Exam number for Gemini 3.8 Flash. The launch post states in text that the model 'achieves a 54.9% on HLE-Verified', but HLE-Verified is a filtered subset rather than the full exam this column tracks — Google's own methodology defines it as 1,811 items, being 668 verified questions from the original set plus 1,143 revised ones, with 689 uncertain items excluded — so it stays a note rather than becoming the cell. That is the same variant discipline this page applies to SWE-Bench Pro against Verified and MMMU against MMMU-Pro. The two instruments also disagree about whether there was a generational gain at all: Google's subset has 3.8 Flash improving on its predecessor by 1.3 points, where the full-exam re-run has the two models level to within a twentieth of a point, with the successor fractionally behind.
Model
Score
Notes
HLE 43.6, from the same open-weights model card's Benchmark Results table, whose Qwen3.8-Max column reads against Opus 4.8 45.7, Fable 5 53.3, GPT 5.6 Sol (max) 47.2 and Qwen3.7-Max 41.4. The table also gives HLE with tools at 56.2; this column records the no-tools figure. Alibaba's launch posts publish no HLE number.
Model
Score
Notes
Grok 4.6's launch post reports no Humanity's Last Exam figure. Its Evals table is agentic-coding and knowledge-work first, with no academic-exam benchmark on it.
Model
Score
Notes
The only cell on this row that describes the current GA build. DeepSeek's change-log entry announcing general availability states in body text 'HLE (wo / w tools): 42.7/60.0', and this column records the no-tools figure. That supersedes the model card's 37.7, which is the April preview build's Max-mode number (the card's ladder reads Non-Think 7.7, High 34.5, Max 37.7, and 48.2 with tools) and which the card still shows because it has not been re-issued since June. The newer primary source wins.
Model
Score
Notes
Mistral published no Humanity's Last Exam figure for Medium 3.5; its release body text names only SWE-Bench Verified and τ³-Telecom.

SWE-Bench Pro

CodingDiscriminating

Contamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases — 731 tasks in the public set, drawn from copyleft repositories to make training-data contamination legally awkward — that the model must patch end-to-end. Down to three of the eight current flagships, all three lab claims run on that lab's own harness rather than on Scale's standardized leaderboard, and no longer the most widely reported column on this page — ARC-AGI-2 passed it in September 2026 by holding four. Claude Fable 5.1's 81.2 sits over a 55–68 pair (Qwen3.8-Max 67.7, DeepSeek-V4-Pro 55.4) — a 25.8-point band, so what remains still discriminates at the production frontier, but on a thinner sample than it used to. Anthropic is the only lab moving on this benchmark rather than away from it: Fable 5.1 raised its predecessor's 80 to 81.2 and added SWE-bench Multilingual and Multimodal alongside it, while OpenAI's GPT-6 Astra dropped the SWE-Bench Pro row GPT-5.6 Sol had claimed at 64.6 and reports Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1 instead, Google's Gemini 3.8 Flash reports no SWE-Bench variant of any kind for a second consecutive release, xAI's Grok 4.6 reports none where the retired Grok 4.5 claimed 64.7, and Meta and Mistral publish nothing here. Google's software-engineering slot is DeepSWE v1.1 and Terminal-Bench instead. The independent side is empty for a structural reason rather than a timing one: Scale AI stewards the benchmark, neither Artificial Analysis nor Epoch AI evaluates it at all, and Scale's own public leaderboard has no row for any of the eight — it is topped by Muse Spark 1.1, two generations behind Meta's current flagship, and its own summary prose still describes top models as scoring around 23% on the public set, which is where the board's methodology write-up was frozen. Cross-lab corroboration is this column's only check, and it thinned twice over in September 2026: OpenAI's and Alibaba's comparison tables had both carried the retired Claude Fable 5's 80, but the one competitor document that names Anthropic's newest flagship — OpenAI's GPT-6 Astra post — carries no SWE-Bench row of any variant, so 81.2 rests on one lab's scaffold until somebody publishes one. The same absence cost the column its cleanest worked example of corroboration, since Alibaba's open-weights card had independently reproduced OpenAI's own 64.6 on its own re-run and there is now no OpenAI figure left to check.

Human baseline
No human pass-rate is published; the authors describe each task as hours-to-days of work for a professional software engineer, and every one was human-verified as solvable before it entered the set

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
The system card's capability summary reads 81.2 for Fable 5.1, against Fable 5's 80, Claude Opus 5's 79.2 and GPT-5.6 Sol's 64.6, and the body text restates it: 'Fable 5.1 achieved 81.2%.' Run on Anthropic's own scaffold and averaged over five trials, so not directly comparable to Scale's standardized leaderboard. Anthropic reports two sibling variants alongside it, SWE-bench Multilingual at 89.1 and SWE-bench Multimodal at 54.7, and no SWE-bench Verified.
Model
Score
Notes
SWE-bench Pro 67.7, from the same open-weights model card's Benchmark Results table, whose Qwen3.8-Max column reads against Opus 4.8 69.2, Fable 5 80.0, GPT 5.6 Sol (max) 64.6 and Qwen3.7-Max 60.6. The card's footnote gives the configuration: 'Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks corrected and all baselines evaluated on the refined benchmark.' That is Alibaba's own harness on a corrected task set, so it is not directly comparable to Scale's standardized leaderboard — though the Fable 5 and GPT-5.6 Sol figures in the same row match those labs' own published numbers exactly, which is some evidence Alibaba is not rescaling the benchmark. Alibaba's own launch write-up published multi-day autonomous-coding case studies and no numeric table at all.
Model
Score
Notes
'SWE Pro (Resolved)' 55.4 for the Max thinking mode, from the model card's frontier-comparison table, where it sits against Opus-4.6 Max 57.3, GPT-5.4 xHigh 57.7, Gemini-3.1-Pro High 54.2, K2.6 Thinking 58.6 and GLM-5.1 58.4; the mode ladder gives Non-Think 52.1, High 54.4, Max 55.4. Run on DeepSeek's internal agent scaffold, so not directly comparable to Scale's standardized leaderboard. Like the GPQA cell this is a pre-GA figure: DeepSeek's GA announcement leads on DeepSWE 62.7 and Terminal Bench 2.1 87.9 for its coding slots and does not restate SWE-Bench Pro.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no SWE-Bench benchmark of any kind, a second consecutive Flash release with none. Google's software-engineering slot is DeepSWE v1.1 at 73.7, Terminal-bench 2.1 at 89.4 and Terminal-bench 4.0 at 19.1 — near neighbours on the same skill, none of them this benchmark. The 58.7 the 3.6 Flash card carried on this benchmark belongs to a model two generations back.
Model
Score
Not reported
Notes
OpenAI has left this benchmark. The Astra post's Coding table is Terminal-Bench 4.0, DeepSWE v1.1, FrontierCode 1.1 Extended, FrontierCode 1.1 Main, Internal Database Migration Tasks and the Artificial Analysis Coding Agent Index v1.4, with no SWE-Bench row of any variant, where the GPT-5.6 post it replaces claimed SWE-Bench Pro 64.6. None of the substitutes is this benchmark: DeepSWE v1.1 and Terminal-Bench are different instruments, and the AA Coding Agent Index is a composite. The retired figure is not carried forward.
Model
Score
Not reported
Notes
Grok 4.6's launch post publishes no SWE-Bench number of any variant. Its coding slots are DeepSWE v1.1 65.9, FrontierCode v1.1 Extended 61.3, APEX-SWE 56.4, CursorBench v3.2 69.9 and Terminal-Bench v3.0 26.0 — near neighbours, none of them this benchmark. The retired Grok 4.5 claimed 64.7 here, but a predecessor's score is not evidence about this model. Worth holding against xAI's own description of Grok 4.6 as its flagship model for code.
Model
Score
Not reported
Notes
Mistral published SWE-Bench Verified for Medium 3.5 but not the harder Pro variant — the opposite of the direction most of the frontier has moved.
Model
Score
Not reported
Notes
Meta published no SWE-Bench Pro number for Muse Spark 1.3. Its software-engineering evaluations are DeepSWE v1.1, Terminal-Bench 2.1 and SWE-Atlas Codebase QnA — the last of which is a Scale AI benchmark of code comprehension, not the patch-writing SWE-Bench Pro this column tracks — and all of them are reported as a chart image with no figures in the text.

AIME 2025

MathApproaching saturation

American Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for several refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes HMMT 2026 Feb, Anthropic publishes ArXivMath and CritPT-Corrected, and Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. The labs' own framing has hardened over two releases. Anthropic's Fable 5 system card named this benchmark to explain the substitution, introducing USAMO as ‘the next step of the math olympiad track in the US after the AIME, which was a popular AI benchmark last year but is now saturated’; the Fable 5.1 card drops the olympiad track entirely for research-level mathematics, preferring a benchmark drawn monthly from recent arXiv abstracts because it is ‘more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks’. The independent side is empty too, and that is checked rather than assumed: Artificial Analysis still maintains an AIME 2025 leaderboard, but it is both a full generation old — topped by GPT-5.2 at 99.0 — and shrinking, down from seven charted rows to six in September 2026 when its last Anthropic entry dropped off, and it contains none of the eight. Epoch AI publishes no AIME data of any year. That distinction matters for whether the column should survive. A leaderboard that exists but has not been refreshed could pick up a current flagship next month; a benchmark nobody maintains an evaluation page for will not. AIME is in the first category and MMMU is in the second, which is why the two zero-coverage columns on this page are not the same case. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.

Author
Mathematical Association of America
Human baseline
Strong high-school competitors solve ~50%; AIME qualifiers (top USAMO contenders) ~80%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no competition-math benchmark for Fable 5.1, and there is no AIME row anywhere in the system card. Its math slot is research-level instead — ArXivMath, a MathArena benchmark whose problems are extracted monthly from recent arXiv abstracts, and CritPT-Corrected, a 71-problem physics-research test — and the card explains the preference directly: ArXivMath is 'more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks.'
Model
Score
Not reported
Notes
DeepSeek reports no AIME of any year. Its competition-math slots are HMMT 2026 Feb, where V4-Pro Max scores 95.2, and IMOAnswerBench at 89.8, and no change-log entry has added a math benchmark since.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no competition-math benchmark, matching the 3.7 and 3.6 cards. Its nearest scientific-reasoning slots are HLE-Verified, BioMysteryBench and LABBench2.
Model
Score
Not reported
Notes
OpenAI's competition-math slot is FrontierMath, where Astra scores 97.6% on Tier 4 (v2) against GPT-5.6 Sol's 83.0%, and the post describes it as saturating that tier at 98%. It publishes no AIME figure of any year.
Model
Score
Not reported
Notes
Grok 4.6's launch post reports no competition-math benchmark; the launch leads on agentic coding and knowledge work instead.
Model
Score
Not reported
Notes
Mistral published no AIME 2025 figure for Medium 3.5, and no competition-math benchmark of any kind.
Model
Score
Not reported
Notes
Meta published no competition-math figure for Muse Spark 1.3; the launch is agent- and coding-first, and its benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Alibaba publishes no AIME 2025 number for Qwen3.8-Max, and the open-weights model card carries no competition-math row of any kind across all 31 of its benchmarks.

GPQA Diamond

KnowledgeApproaching saturation

Graduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier, and all eight current flagships carry a value. The frontier has effectively cleared it: seven of the eight sit inside a 5.9-point band from DeepSeek-V4-Pro's 90.1 to GPT-6 Astra's 96.0, and Anthropic has now acted on the position its Fable 5 system card set out — that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it — by publishing no GPQA figure at all for Claude Fable 5.1, where the benchmark survives only as one of three question sets used to measure chain-of-thought controllability. The full eight-model band reads 21.2 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a crossing of the 20-point threshold that is entirely an artifact of the floor. That arithmetic is recomputed rather than assumed on every refresh, and it has come out the same way every time it has been run. Only three of the eight values are lab claims; five come from Artificial Analysis. This is also the one column where two independent evaluators can be compared directly, and the comparison is a caution against reading the top of it too closely: Epoch AI runs its own GPQA Diamond evaluations with published per-run logs, and where it and Artificial Analysis have measured the same model it has agreed to within a point on most of them while putting the retired Claude Fable 5 at 85.9 against AA's 92.6 — a 6.7-point disagreement between two careful evaluators on the same model and the same benchmark, larger than the entire band separating the other seven. That check is partly restored at the top of the board: Epoch now runs GPT-6 Astra at 95.8 and has it leading its own 311-row board, against Artificial Analysis's 96.3 for the same model — a 0.5-point agreement, and the closest the two evaluators have come on a current leader. It still cannot be run on the rest of the top, because Epoch's newest Google entry is Gemini 3.7 Flash at 94.8 and its newest Anthropic entry is Claude Opus 5 at 93.9, both a generation behind what Artificial Analysis is charting. So the two boards now agree on who leads this column and disagree about who is second.

Author
Rein et al. · NYU · Cohere
Human baseline
PhD-level experts in matched domains score ~65%; non-experts with web access ~34%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
From the Academic table of OpenAI's GPT-6 Astra post, whose six columns read Astra 96.0, GPT-5.6 Sol 94.6, Claude Fable 5.1 93.7, Claude Fable 5 92.6, Claude Opus 5 93.7 and Gemini 3.8 Flash 95.3 — the same column set every table in that post uses, so there is no Ultra-style extra column to mis-align against. The post's body text adds that at a lower-cost setting Astra scores 94.9, still above Sol's best 94.6, at roughly 37% lower estimated API cost. Scores are the maximum at any effort per the tables' own footer. This is the highest lab-claimed figure on the column and a 1.4-point gain over the retired Sol row.
Model
Score
Notes
The Gemini 3.8 Flash model card reports no GPQA Diamond, the same omission the 3.7 and 3.6 cards had. Its published results table runs to thirteen benchmarks and contains no multiple-choice science evaluation of any kind — the nearest entries are HLE-Verified and two biology-research sets, BioMysteryBench and LABBench2. OpenAI's GPT-6 Astra comparison table does list the model at 95.3, but that is a competitor restating Artificial Analysis's number rather than a figure Google published.
Model
Score
Notes
Grok 4.6's launch post reports no GPQA Diamond. Its Evals table has no multiple-choice science row of any kind.
Model
Score
Notes
Meta published no GPQA Diamond figure for Muse Spark 1.3. Its single launch benchmark figure is a chart image with no numbers in the surrounding text, and the companion methodology PDF carries no figures at all, so nothing here is quotable as a lab claim.
Model
Score
Notes
Anthropic publishes no GPQA Diamond figure for Fable 5.1. The system card's capability summary has no GPQA row, and the benchmark survives in the document only as one of three question sets used to measure chain-of-thought controllability. That follows through on the Fable 5 card's stated intent: Anthropic 'consider[s] GPQA Diamond to be a saturated evaluation and plan[s] to stop reporting the performance of future models on it.' OpenAI's GPT-6 Astra comparison table does list Fable 5.1 at 93.7, but that is a competitor restating Artificial Analysis's number rather than a figure Anthropic published.
Model
Score
Notes
GPQA Diamond 92.6, from the Benchmark Results table in the model card for Qwen3.8-2.4T-A95B, the open-weights build Alibaba released nine days after Qwen3.8-Max. That card carries a column headed Qwen3.8-Max because, in its own words, 'Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features'; the row reads Opus 4.8 92.0, Fable 5 92.6, GPT 5.6 Sol (max) 94.1, Qwen3.7-Max 92.4, Qwen3.8-Max 92.6. Alibaba's own launch posts publish no benchmark table at all — their headline claims are LMArena ranks — so the open-weights card is the only text source for this figure.
Model
Score
Notes
GPQA Diamond (Pass@1) 90.1 for the Max thinking mode, from the DeepSeek-V4-Pro model card's frontier-comparison table; the card's mode ladder gives Non-Think 72.9, High 89.1 and Max 90.1. Read this as a pre-GA figure: the card has not been re-issued since June and describes the April preview weights, while the `deepseek-v4-pro` id now resolves to the August GA build. DeepSeek's GA announcement restated HLE but not GPQA, so there is no newer lab number to move to.
Model
Score
Notes
Mistral published no GPQA Diamond for Medium 3.5. Its release names two numbers and no others — SWE-Bench Verified 77.6 and τ³-Telecom 91.4 — skipping the whole MMLU / GPQA / AIME / HumanEval / MATH set that most launches lead with.

MMMU

MultimodalApproaching saturation

Massive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: Google reports CharXiv Reasoning and LVBench, Anthropic reports Chartography and GDP.pdf, Meta publishes no multimodal figure of any kind, and no current flagship publishes a plain MMMU number. OpenAI's abandonment went a step further in September 2026 — where the GPT-5.6 post at least had a Multimodal table reporting MMMU-Pro, the GPT-6 Astra post has no multimodal table at all, and files its image work under computer use as ScreenSpot-Pro and OSWorld 2.0. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row. The abandonment extends to the independent side, and that is the strongest evidence that this column has run its course: Artificial Analysis has no plain-MMMU evaluation page at all — the URL returns a 404 and only MMMU-Pro exists — and Epoch AI's benchmark corpus carries no MMMU file either. The benchmark's own authors are the third piece of the same picture, though not in the way the front page suggests. The leaderboard shows a ‘Last updated: 09/05/2025’ stamp and renders its table client-side, but the data file that table is drawn from is live: 211 rows, with entries as recent as July 2026. Every row added since November 2025 carries an MMMU-Pro score and leaves the plain-MMMU cell empty, and the last model to receive a plain-MMMU number there was Claude Opus 4.5. So the authors did not stop maintaining the board — they moved it to the harder variant, which is the same move the labs made. Three of the eight flagships are currently charted on Artificial Analysis's MMMU-Pro board — GPT-6 Astra 86.9 at the top of it, Gemini 3.8 Flash 85.6 and Mistral Medium 3.5 64.9 — but that is a different and harder instrument, so importing one would break the column's internal comparability. Muse Spark 1.3, charted at 82.0 a cycle ago, has fallen outside that board's eighteen-row window while staying live in Artificial Analysis's registry, which is a windowing artifact rather than a retraction. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and the strongest removal candidate on the page.

Author
MMMU Benchmark · University of Waterloo + collaborators
Human baseline
No single number: the authors publish three expert tiers, scoring 88.6% at best, 82.6% in the middle and 76.2% at worst on the validation set

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no plain-MMMU number for Fable 5.1. Its multimodal section runs on Chartography, BenchCAD, OSWorld 2.0 and GDP.pdf — no MMMU of any variant, and no MMMU-Pro either.
Model
Score
Not reported
Notes
Neither of the model card's benchmark tables contains an MMMU row, and no change-log entry names one. The V-series is text-and-code first; DeepSeek's separate vision model is a different tier and is not this row.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card uses CharXiv Reasoning at 86.2, LVBench at 87.8 agentic and 87.1 static, and GDP.PDF at 35.0 as its multimodal signals, and reports neither MMMU nor MMMU-Pro. Chart reasoning, long-video understanding and PDF comprehension each test a slice of what MMMU covers; none of them is the college-exam instrument this column tracks.
Model
Score
Not reported
Notes
OpenAI publishes no multimodal benchmark table for Astra at all — a change from the GPT-5.6 post, which had one carrying MMMU-Pro (Sol 83% without tools, 84.6% with) and gdp.pdf. Astra's image-understanding results appear only as computer-use figures, ScreenSpot-Pro 92.7 and OSWorld 2.0 72.6, neither of which is MMMU or a variant of it. So there is no lab figure to consider for this cell, near-miss or otherwise.
Model
Score
Not reported
Notes
Grok 4.6's launch post reports no MMMU. The model takes text and image input and returns text only, but xAI published no multimodal benchmark number with it.
Model
Score
Not reported
Notes
Mistral published no MMMU figure for Medium 3.5. Its release describes the model as the first Mistral flagship to merge instruction-following, reasoning, coding and a from-scratch vision encoder into a single set of weights, but attaches no multimodal benchmark to the claim.
Model
Score
Not reported
Notes
Meta published no MMMU figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Qwen3.8-Max is natively vision-language, and Alibaba still publishes no plain-MMMU number for it — not in either launch post and not in the open-weights model card, whose multimodal coverage stops at describing the vision input the hosted model adds over the open build.

SWE-Bench Verified

CodingApproaching saturation

The 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025, and now down to two of the eight current flagships — DeepSeek-V4-Pro at 80.6 and Mistral Medium 3.5 at 77.6 — after Anthropic dropped it. That departure is the column's clearest signal yet, because Anthropic held the highest figure on the page at 95 and its replacement flagship's system card reports SWE-bench Pro, Multilingual and Multimodal with no Verified row of any kind. OpenAI, Google, xAI, Meta and Alibaba had already left. Alibaba's open-weights model card is the neatest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. With two values the band reads 3 points, which is far too thin a sample to move a saturation label on, so the prior label stands. The column's one independent value comes from Epoch AI rather than from Aider: Epoch runs the benchmark itself and publishes the results in its bulk data download, and its board carries DeepSeek-V4-Pro at 77.6 against that lab's own 80.6 — a 3-point gap, under this page's divergence threshold, and a like-for-like comparison because both figures describe the same April build. Epoch's board is the reason to be careful about reading the labs' abandonment as neglect rather than substitution: it runs to 35 rows and its newest entries from OpenAI, Anthropic, Google and Alibaba are all at least one generation behind those labs' current flagships, which is what you would expect if the benchmark is still being run but is no longer what the launch posts lead with. Aider's polyglot leaderboard, the other board that historically covered this benchmark, has added nothing since two DeepSeek V3.2-Exp runs in October 2025 and carries none of the eight.

Human baseline
Subset construction targets human-solvable issues; success rate not directly comparable to model pass@1

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
'SWE Verified (Resolved)' 80.6 for the Max thinking mode, from the model card, level with Opus-4.6 Max at 80.8 and Gemini-3.1-Pro High at 80.6; the mode ladder gives Non-Think 73.6, High 79.4, Max 80.6. Same pre-GA caveat as the GPQA and SWE Pro cells — DeepSeek's GA announcement does not restate it.
Model
Score
Notes
From Mistral's own body text: 'Mistral Medium 3.5 scores 77.6% on SWE-Bench Verified, ahead of Devstral 2 and models like Qwen3.5 397B A17B.' The only one of this page's seven benchmarks Mistral publishes a number for at all.
Model
Score
Not reported
Notes
Anthropic no longer reports SWE-bench Verified. The Fable 5.1 system card's software-engineering section covers three variants — SWE-bench Pro, Multilingual and Multimodal — and carries no Verified row anywhere, where the Fable 5 card had published 95. Anthropic's Verified figure was already the secondary, near-saturated number behind its Pro result; dropping it leaves the harder variant as its only SWE-bench claim.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no SWE-Bench variant at all — not Verified, and not Pro.
Model
Score
Not reported
Notes
OpenAI no longer reports SWE-Bench Verified, the benchmark it originally published: Verified was dropped from its launch materials in February 2026, and the Astra post's Coding table now carries no SWE-Bench row of any variant at all — leading instead on Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1.
Model
Score
Not reported
Notes
Grok 4.6's launch post reports no SWE-Bench Verified. xAI has not published a Verified number since before Grok 4.5.
Model
Score
Not reported
Notes
Meta published no SWE-Bench Verified figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Alibaba publishes no SWE-Bench Verified number for Qwen3.8-Max. The open-weights model card that supplies this row's other lab figures runs to 31 benchmark rows, including SWE-bench Pro, with no Verified row among them — the cleanest single illustration on this page of the frontier moving wholly to the Pro variant. The retired Qwen3.7-Max claimed 80.4 here, but a predecessor's score is not evidence about this model.

About this page

Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.

Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.

Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.

Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.

Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.

Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.

Last verified: September 8, 2026. 8 frontier models · 7 benchmarks · 8 labs.