AI benchmarks

Every frontier AI model on every major benchmark

Frontier AI benchmark scores for each lab’s current flagship, as of September 29, 2026: on ARC-AGI-2, GPT-6 Astra leads at 95%; on GPQA, GPT-6 Astra leads at 96%; on SWE Pro, Claude Fable 5.1 leads at 81.2%. Every score below cites the lab’s announcement post or an independent re-runner.

Last verified: September 29, 2026.

ARC-AGI-2 leader
95%
GPQA leader
96%
SWE Pro leader
81.2%
Coverage
8 × 7
frontier models × benchmarks

How to read this page

● 87.2 — lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.

□ 86.4 — independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.

⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).

— — the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.

Headline matrix

Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.

Anthropic · September 1, 2026
ARC-AGI-2
SWE Pro
AIME
—
MMMU
—
SWE Verified
—
DeepSeek · September 10, 2026
ARC-AGI-2
—
SWE Pro
—
AIME
—
MMMU
—
SWE Verified
—
Google · September 2, 2026
ARC-AGI-2
SWE Pro
—
AIME
—
MMMU
—
SWE Verified
—
OpenAI · September 3, 2026
ARC-AGI-2
SWE Pro
—
AIME
—
MMMU
—
SWE Verified
—
Grok 4.7Closed
xAI · September 21, 2026
ARC-AGI-2
—
SWE Pro
—
AIME
—
GPQA
—
MMMU
—
SWE Verified
—
Mistral AI · April 28, 2026
ARC-AGI-2
—
SWE Pro
—
AIME
—
MMMU
—
SWE Verified
Meta · September 2, 2026
ARC-AGI-2
—
SWE Pro
—
AIME
—
MMMU
—
SWE Verified
—
Alibaba · August 3, 2026
ARC-AGI-2
—
SWE Pro
AIME
—
MMMU
—
SWE Verified
—

By benchmark

Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.

ARC-AGI-2

ReasoningDiscriminating

Successor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. Two labs now name an ARC-AGI-2 figure in their own launch materials, and in both cases the number matches the ARC Prize Foundation's verified board rather than describing a separate run: Anthropic's Claude Fable 5.1 system card gives ARC-AGI-1 97.5 and ARC-AGI-2 90 and says whose numbers they are, and OpenAI's GPT-6 Astra post carries a full ARC-AGI-1/2/3 table whose 95.0 for Astra is the same figure the Foundation publishes. So every value in this column traces back to one measurement per model, taken from that board at each model's highest verified reasoning-effort variant. Coverage is three of the eight current flagships: GPT-6 Astra leads at 95.0, Claude Fable 5.1 follows at 90.0, and Gemini 3.8 Flash — which entered the board on September 24, 2026 without Google publishing anything — sits just under at 89.2. That is a 5.8-point band across three cells, far too thin a sample to move a saturation label on, and more a fact about which labs happen to have a current board row than about how far apart the field is. The band moved four times in a month for reasons that have nothing to do with capability: GPT-6 Astra's arrival widened it from 31.2 by lifting only the top, DeepSeek's flagship swap narrowed it to 27.9 by removing the bottom, xAI's swap collapsed it to 5.0 by removing the bottom again, and the September 24 rebuild widened it slightly to 5.8 by adding a row near the floor. The five flagships still without a submission are Grok 4.7, DeepSeek-V4.1-Flash, Muse Spark 1.3, Mistral Medium 3.5 and Qwen3.8-Max. The board has rebuilt twice since Grok 4.7 shipped — on September 24 and again on September 28, 2026, when it added OpenAI's GPT-6 Sol — and neither rebuild ran Grok 4.7 or DeepSeek-V4.1-Flash, which had been out for eighteen days by the second. Whether that is a submission not made or a queue not reached is not something the board publishes, and the page does not guess. The other three labs have never placed a model anywhere but the floor of this board, appearing only through older siblings: two Llama 4 builds and three Magistral builds at zero, and one Qwen3-235B build at 1.25. The board rebuilds irregularly rather than on a schedule, so a missing row means a model has not been run, never that it ran and scored nothing.

Author
François Chollet · ARC Prize Foundation
Human baseline
Every eval task was solved pass@2 by at least two people in a 400+ participant calibration study; ARC Prize sets 85% as the target a system must reach

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
OpenAI reports ARC-AGI-2 in its own launch materials for the first time on a flagship: the Astra post's Abstract reasoning table reads 95.0 for GPT-6 Astra, 92.5 for GPT-5.6 Sol, 90.4 for Claude Opus 5, 90.0 for Claude Fable 5.1, 89.2 for Claude Fable 5 and an em-dash for Gemini 3.8 Flash, alongside ARC-AGI-3 and ARC-AGI-1 rows. The 5.6 generation's abstract-reasoning slot had been ARC-AGI-3 alone. The 95.0 is identical to the ARC Prize Foundation's verified max-effort score in this cell's independent half, so the two halves describe one measurement rather than two — the same reading the page gives Claude Fable 5.1, whose system card restates the Foundation's number explicitly. OpenAI's tables are footed 'Evaluation scores are the maximum at any effort', which is the configuration the Foundation's Max row also describes.
Model
Score
Notes
Anthropic reports ARC-AGI for the first time on a Fable-class model: the system card's capability evaluation summary gives ARC-AGI-2 90% and ARC-AGI-1 97.5%, both at max effort on the semi-private validation sets, where the Fable 5 card carried no abstract-reasoning row of any kind. The card attributes the figures to the ARC Prize Foundation rather than to an Anthropic run — 'The ARC Prize Foundation reports that Fable 5.1 achieved a verified score of 97.5% on ARC-AGI-1 and 90% on ARC-AGI-2 at max effort' — so the lab and independent values in this cell describe one measurement, not two.
Model
Score
Notes
The Gemini 3.8 Flash model card reports no ARC-AGI benchmark, as the 3.7 and 3.6 cards did not. Google's published eval set leads on agentic coding, computer use and enterprise knowledge workflows instead.
Model
Score
Not reported
Notes
DeepSeek publishes no ARC-AGI figure of any generation for this model. Neither of the model card's benchmark tables carries an ARC row, and the release announcement publishes no numbers at all.
Model
Score
Not reported
Notes
Grok 4.7's launch post reports no ARC-AGI benchmark. Its comparison table runs CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, the Harvey Legal Agent Benchmark and HealthBench Professional, alongside a GDPval Elo chart — agentic coding and professional knowledge work throughout, with nothing testing abstract reasoning.
Model
Score
Not reported
Notes
Mistral reports no ARC-AGI benchmark for Medium 3.5.
Model
Score
Not reported
Notes
Meta's Muse Spark 1.3 launch post reports no ARC-AGI benchmark — and reports the one evaluation summary it does publish as a chart image rather than as text, so none of its figures can be quoted as a lab claim. The image's alt text stops at 'Benchmark scorecard comparing Muse Spark 1.3, Muse Spark 1.2, GPT 5.6 Sol (max), and Opus 5 (max) across agent, coding, instruction-following, and long-context evaluations', with no numbers in it. This page does not read numbers off pictures.
Model
Score
Not reported
Notes
Neither Alibaba's Qwen3.8-Max launch posts nor the 31-row benchmark table in the open-weights sibling's model card — the source of this row's other lab figures — reports any ARC-AGI benchmark.

Humanity's Last Exam

KnowledgeDiscriminating

Crowdsourced expert-level exam of 2,500 questions across more than a hundred subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools; this column records the no-tools figure. The only benchmark on this page with a published value for all eight current flagships, and by some way the widest band of any column here: Claude Fable 5.1 tops this page's flagships at 60.9 and Mistral Medium 3.5 anchors them at 13.8, a 47.1-point spread, with the middle six between 36.8 and 54.7. The top of this column is not the top of the field: Anthropic's Claude Opus 5.5, a higher tier this page's roster does not track, publishes 64.4 in its own system card. HLE took the all-eight distinction outright in September 2026, when GPQA Diamond — which had also carried all eight — lost a cell to xAI's flagship swap and has not got it back. Some of the width is one straggler, but less than it used to be. Drop Mistral and the remaining seven span 24.1 points, against 12.8 before September 2026, and both ends moved to get there: Anthropic's new flagship took the lead to 60.9, GPT-6 Astra entered 5.2 points above the retired GPT-5.6 Sol at 54.7, and DeepSeek then swapped a 42.7 for its successor's 36.8 — the only flagship swap on this page to lower a lab's own HLE figure. The shape that produced is a ladder rather than a cluster: leader, Astra 6.2 points back, a four-model band inside 5.6 points, DeepSeek a further 6.3 below that, then Mistral far under everything. Three of the eight values are lab claims; the rest come from Artificial Analysis, which runs the full exam on every current flagship it charts. Three rows carry both halves, and all three agree well inside the divergence threshold — Anthropic 60.9 against 59.1, Alibaba 43.6 against 43.1, DeepSeek 36.8 against 39.2, the last being the rare case where the independent evaluator is the more generous of the two. A second independent measurement exists for three of the eight, and it comes from the benchmark's own authors: Scale AI's official Humanity's Last Exam leaderboard, run on every public question with a fixed judge model, lists Claude Fable 5.1 at 46.5, GPT-6 Astra at 54.8 and Gemini 3.8 Flash at 44.5, and Epoch AI republishes the same figures. Set against Artificial Analysis's numbers for the same models, those are gaps of 12.2, 0.1 and 3.3 points, so the disagreement is specific to one model rather than a systematic difference between the two harnesses. The Anthropic result — 12.2 points below Artificial Analysis and 14.4 below Anthropic's own figure — is still the widest evaluator-versus-evaluator gap anywhere on this page, nearly twice the 6.7 points separating the two GPQA Diamond boards on the retired Claude Fable 5, and neither board explains it. The authors released HLE-Diamond on September 22, 2026, a refined 1,000-question subset whose first results run from GPT-6 Astra's 60.6 down to Grok 4.7's 23.4; it is a different instrument and stays out of this column. So does Google's HLE-Verified (3.8 Flash 54.9), a filtered 1,811-item subset rather than the full exam — and the two instruments disagree about direction, since Google's subset shows a 1.3-point generational gain where the full-exam re-run has the new model a twentieth of a point behind its predecessor.

Author
Center for AI Safety · Scale AI
Human baseline
No aggregate human score is published; the questions were written and cross-checked by close to a thousand expert contributors and are built to sit past the reach of any one specialist

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
Anthropic ran Humanity's Last Exam on Fable 5.1 itself, where the Fable 5 card had an em-dash in the Fable column and reported Mythos 5 alone. The capability summary gives 60.9% without tools and 65.0% with, against Fable 5's 57.8 and 63.8 and Claude Opus 5's 56.6 and 63.6; this column records the no-tools figure. Both configurations use adaptive thinking at max effort, averaged over five trials, with a 1M total-token cap and no context compaction. The with-tools run adds web search, a fetch tool restricted to URLs already in the conversation, programmatic tool calling and code execution, blocklists sources known to discuss the exam, and re-grades as incorrect any transcript found to have retrieved an answer.
Model
Score
Notes
OpenAI publishes an HLE figure for Astra but not the one this column tracks. The Academic table's row is headed 'Humanity's Last Exam (w/ tools)' and gives Astra 57.2, where this column records the no-tools figure, so it stays out of the cell — the same near-miss discipline that keeps Google's HLE-Verified 54.9 out of the Gemini cell. That the row is the with-tools instrument is checkable inside OpenAI's own table, which lists Claude Fable 5.1 at 65.0: exactly Anthropic's published with-tools number, against the 60.9 no-tools figure in this page's Anthropic cell. OpenAI has never published a no-tools HLE score for any model, and GPT-5.6 Sol's cell in this same row is an em-dash.
Model
Score
Notes
Meta published no Humanity's Last Exam figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text, so nothing here is quotable as a lab claim.
Model
Score
Notes
Google publishes no plain Humanity's Last Exam number for Gemini 3.8 Flash. The launch post states in text that the model 'achieves a 54.9% on HLE-Verified', but HLE-Verified is a filtered subset rather than the full exam this column tracks — Google's own methodology defines it as 1,811 items, being 668 verified questions from the original set plus 1,143 revised ones, with 689 uncertain items excluded — so it stays a note rather than becoming the cell. That is the same variant discipline this page applies to SWE-Bench Pro against Verified and MMMU against MMMU-Pro. The two instruments also disagree about whether there was a generational gain at all: Google's subset has 3.8 Flash improving on its predecessor by 1.3 points, where the full-exam re-run has the two models level to within a twentieth of a point, with the successor fractionally behind.
Model
Score
Notes
HLE 43.6, from the same open-weights model card's Benchmark Results table, whose Qwen3.8-Max column reads against Opus 4.8 45.7, Fable 5 53.3, GPT 5.6 Sol (max) 47.2 and Qwen3.7-Max 41.4. The table also gives HLE with tools at 56.2; this column records the no-tools figure. Alibaba's launch posts publish no HLE number.
Model
Score
Notes
Grok 4.7's launch post reports no Humanity's Last Exam figure. Its comparison table is agentic-coding and professional-knowledge-work first, with no academic-exam benchmark on it.
Model
Score
Notes
HLE (Pass@1) 36.8 at maximum reasoning effort, no tools, from the model card's frontier-comparison table. The card prints the cell as '36.8 (39.1†)' under a footnote reading 'Text-only subset of HLE', so the unmarked figure is the full exam and the parenthetical is a filtered subset this column does not track — the same distinction that keeps Google's HLE-Verified out of its cell. The card's separate 'HLE w/ tools' row gives 63.9, a different configuration. This is the lowest HLE figure DeepSeek has published for a flagship in 2026: its retired V4-Pro claimed 42.7, and the card's own text-only column puts V4-Pro at 42.7† and V4-Flash at 37.8†, so the generational comparison DeepSeek draws is between dagger figures rather than between full-exam ones.
Model
Score
Notes
Mistral published no Humanity's Last Exam figure for Medium 3.5; its release body text names only SWE-Bench Verified and τ³-Telecom.

SWE-Bench Pro

CodingDiscriminating

Contamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases — 731 tasks in the public set, drawn from copyleft repositories to make training-data contamination legally awkward — that the model must patch end-to-end. Down to two of the eight current flagships, both lab claims run on that lab's own harness rather than on Scale's standardized leaderboard, and no longer the most widely reported column on this page — ARC-AGI-2 passed it in September 2026, briefly fell level at two apiece when xAI's flagship swap cost that column a cell, and pulled ahead again on September 24 when the ARC Prize board rebuilt and picked up Gemini 3.8 Flash. Claude Fable 5.1's 81.2 sits over Qwen3.8-Max's 67.7 and nothing else: a 13.5-point band on two cells, which is far too thin a sample to move a saturation label on, so the label below is the one this column carried when it had enough data to compute one. Anthropic is the only lab moving on this benchmark rather than away from it: Fable 5.1 raised its predecessor's 80 to 81.2 and added SWE-bench Multilingual and Multimodal alongside it, and two Anthropic tiers this page's roster does not track have since gone past it in their own system cards — Claude Opus 5.5 at 89.9 and Claude Sonnet 5.5 at 81.3 — so the top of this column is the top of the tracked flagships, not of the field. Everyone else moved away: OpenAI's GPT-6 Astra dropped the SWE-Bench Pro row GPT-5.6 Sol had claimed at 64.6 and reports Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1 instead, DeepSeek-V4.1-Flash dropped the 55.4 the retired V4-Pro claimed and reports DeepSWE v1.1, NL2Repo-Bench and ProgramBench, Google's Gemini 3.8 Flash reports no SWE-Bench variant of any kind for a second consecutive release, xAI's Grok 4.7 reports none and nor did Grok 4.6 before it — the last xAI figure in this column was the retired Grok 4.5's 64.7, two flagships back — and Meta and Mistral publish nothing here. Three of the six departures happened inside a single fortnight, and each of the three replaced SWE-Bench Pro with an agentic-terminal or repository-scale benchmark rather than with nothing — the substitution is toward longer-horizon coding work, not away from measuring it. The independent side is empty for a structural reason rather than a timing one: Scale AI stewards the benchmark, neither Artificial Analysis nor Epoch AI evaluates it at all, and Scale's own leaderboard for the original task set — one board for the public dataset and one for the private commercial subset — has no row for any of the eight. Both boards are topped by Muse Spark 1.1, two generations behind Meta's current flagship, and the page's summary prose still describes top models as scoring around 23% on the public set. In September 2026 Scale added SWE-Bench Pro V2, a refreshed public split with a modified task set and a locked evaluation protocol, which does carry three of this page's flagships — Claude Fable 5.1 at 99.1, GPT-6 Astra at 96.9 and Gemini 3.8 Flash at 94.9 on its full split, and 92.2, 90.2 and 58.8 on its hard split, each on its own lab's agent harness. Those figures come from a different instrument from the one every lab number in this column was measured on, so they stay out of the cells rather than being read as independent checks on the lab claims. Cross-lab corroboration is therefore this column's only check, and it thinned twice over in September 2026: OpenAI's and Alibaba's comparison tables had both carried the retired Claude Fable 5's 80, but the one competitor document that names Anthropic's newest flagship — OpenAI's GPT-6 Astra post — carries no SWE-Bench row of any variant, so 81.2 rests on one lab's scaffold until somebody publishes one. The same absence cost the column its cleanest worked example of corroboration, since Alibaba's open-weights card had independently reproduced OpenAI's own 64.6 on its own re-run and there is now no OpenAI figure left to check.

Human baseline
No human pass-rate is published; the authors describe each task as hours-to-days of work for a professional software engineer, and every one was human-verified as solvable before it entered the set

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
The system card's capability summary reads 81.2 for Fable 5.1, against Fable 5's 80, Claude Opus 5's 79.2 and GPT-5.6 Sol's 64.6, and the body text restates it: 'Fable 5.1 achieved 81.2%.' Run on Anthropic's own scaffold and averaged over five trials, so not directly comparable to Scale's standardized leaderboard. Anthropic reports two sibling variants alongside it, SWE-bench Multilingual at 89.1 and SWE-bench Multimodal at 54.7, and no SWE-bench Verified.
Model
Score
Notes
SWE-bench Pro 67.7, from the same open-weights model card's Benchmark Results table, whose Qwen3.8-Max column reads against Opus 4.8 69.2, Fable 5 80.0, GPT 5.6 Sol (max) 64.6 and Qwen3.7-Max 60.6. The card's footnote gives the configuration: 'Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks corrected and all baselines evaluated on the refined benchmark.' That is Alibaba's own harness on a corrected task set, so it is not directly comparable to Scale's standardized leaderboard — though the Fable 5 and GPT-5.6 Sol figures in the same row match those labs' own published numbers exactly, which is some evidence Alibaba is not rescaling the benchmark. Alibaba's own launch write-up published multi-day autonomous-coding case studies and no numeric table at all.
Model
Score
Not reported
Notes
DeepSeek has dropped SWE-Bench from its reporting: the DeepSeek-V4.1-Flash card carries no SWE-Bench row of any variant, where the retired V4-Pro card published SWE Pro 55.4. Its software-engineering slots are DeepSWE v1.1 74.2 resolved, NL2Repo-Bench 64.0 and ProgramBench 20.3 instead — the NL2Repo figure being the model card's; DeepSeek's own change-log entry for the same release gives 65.4 for that benchmark, and it is the only row of the nineteen the two surfaces share where they disagree at all. Two rows on the card read like this benchmark and are not it: 'SEC-Bench Pro (Pass@1) 62.8' is a security benchmark run on the Claude Code harness, and DeepSWE v1.1 is a separate instrument with its own harness requirements.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no SWE-Bench benchmark of any kind, a second consecutive Flash release with none. Google's software-engineering slot is DeepSWE v1.1 at 73.7, Terminal-bench 2.1 at 89.4 and Terminal-bench 4.0 at 19.1 — near neighbours on the same skill, none of them this benchmark. The 58.7 the 3.6 Flash card carried on this benchmark belongs to a model two generations back.
Model
Score
Not reported
Notes
OpenAI has left this benchmark. The Astra post's Coding table is Terminal-Bench 4.0, DeepSWE v1.1, FrontierCode 1.1 Extended, FrontierCode 1.1 Main, Internal Database Migration Tasks and the Artificial Analysis Coding Agent Index v1.4, with no SWE-Bench row of any variant, where the GPT-5.6 post it replaces claimed SWE-Bench Pro 64.6. None of the substitutes is this benchmark: DeepSWE v1.1 and Terminal-Bench are different instruments, and the AA Coding Agent Index is a composite. The retired figure is not carried forward.
Model
Score
Not reported
Notes
Grok 4.7's launch post publishes no SWE-Bench number of any variant. Its coding slots are DeepSWE v1.1 71.0 (at high effort, per the table's own asterisk), CursorBench 4.0 46.3 and Terminal-Bench 4.0 37.6 — near neighbours, none of them this benchmark. Worth holding against xAI's description of Grok 4.7 as its most capable model for coding and knowledge work: the lab leads on coding and reports no SWE-Bench figure of any variant, a pattern now unbroken across three consecutive Grok flagships.
Model
Score
Not reported
Notes
Mistral published SWE-Bench Verified for Medium 3.5 but not the harder Pro variant — the opposite of the direction most of the frontier has moved.
Model
Score
Not reported
Notes
Meta published no SWE-Bench Pro number for Muse Spark 1.3. Its software-engineering evaluations are DeepSWE v1.1, Terminal-Bench 2.1 and SWE-Atlas Codebase QnA — the last of which is a Scale AI benchmark of code comprehension, not the patch-writing SWE-Bench Pro this column tracks — and all of them are reported as a chart image with no figures in the text.

AIME 2025

MathApproaching saturation

American Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for several refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes MathArena Apex and a Codeforces rating, Anthropic publishes ArXivMath and CritPT-Corrected, and Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. The labs' own framing has hardened over two releases. Anthropic's Fable 5 system card named this benchmark to explain the substitution, introducing USAMO as ‘the next step of the math olympiad track in the US after the AIME, which was a popular AI benchmark last year but is now saturated’; the Fable 5.1 card drops the olympiad track entirely for research-level mathematics, preferring a benchmark drawn monthly from recent arXiv abstracts because it is ‘more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks’. The independent side is empty too, and that is checked rather than assumed: Artificial Analysis still maintains an AIME 2025 leaderboard, but it is both a full generation old — topped by GPT-5.2 at 99.0 — and steadily shrinking, down from seven charted rows to six and then to five across September 2026 as its last Anthropic entry and then its floor row dropped off, and it contains none of the eight — every one of the five survivors is a model Artificial Analysis itself now marks deprecated. It loses rows from the bottom while the top stays put, which is what a board nobody is submitting to looks like. Epoch AI publishes no AIME data of any year. That distinction matters for whether the column should survive. A leaderboard that exists but has not been refreshed could pick up a current flagship next month; a benchmark nobody maintains an evaluation page for will not. AIME is in the first category and MMMU is in the second, which is why the two zero-coverage columns on this page are not the same case. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.

Author
Mathematical Association of America
Human baseline
Strong high-school competitors solve ~50%; AIME qualifiers (top USAMO contenders) ~80%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no competition-math benchmark for Fable 5.1, and there is no AIME row anywhere in the system card. Its math slot is research-level instead — ArXivMath, a MathArena benchmark whose problems are extracted monthly from recent arXiv abstracts, and CritPT-Corrected, a 71-problem physics-research test — and the card explains the preference directly: ArXivMath is 'more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks.'
Model
Score
Not reported
Notes
DeepSeek reports no AIME of any year, and the substitution moved again with this release: where the retired V4-Pro filled its math slot with HMMT 2026 Feb and IMOAnswerBench, the DeepSeek-V4.1-Flash card reports MathArena Apex 65.6 and a Codeforces rating of 3471 instead.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no competition-math benchmark, matching the 3.7 and 3.6 cards. Its nearest scientific-reasoning slots are HLE-Verified, BioMysteryBench and LABBench2.
Model
Score
Not reported
Notes
OpenAI's competition-math slot is FrontierMath, where Astra scores 97.6% on Tier 4 (v2) against GPT-5.6 Sol's 83.0%, and the post describes it as saturating that tier at 98%. It publishes no AIME figure of any year.
Model
Score
Not reported
Notes
Grok 4.7's launch post reports no competition-math benchmark; the launch leads on agentic coding and professional knowledge work instead.
Model
Score
Not reported
Notes
Mistral published no AIME 2025 figure for Medium 3.5, and no competition-math benchmark of any kind.
Model
Score
Not reported
Notes
Meta published no competition-math figure for Muse Spark 1.3; the launch is agent- and coding-first, and its benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Alibaba publishes no AIME 2025 number for Qwen3.8-Max, and the open-weights model card carries no competition-math row of any kind across all 31 of its benchmarks.

GPQA Diamond

KnowledgeApproaching saturation

Graduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier. Seven of the eight current flagships carry a value; Grok 4.7 does not, and that gap is new — xAI published no GPQA figure with it, and Artificial Analysis, which had scored every previous Grok, has ingested the model without running this benchmark on it. So this column lost the all-eight coverage it shared with Humanity's Last Exam, and HLE now holds that alone. The frontier has effectively cleared what remains: six of the seven sit inside a 5.1-point band from DeepSeek-V4.1-Flash's 90.9 to GPT-6 Astra's 96.0, and Anthropic has acted on the position its Fable 5 system card set out — that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it — by publishing no GPQA figure at all for Claude Fable 5.1, where the benchmark survives only as one of three question sets used to measure chain-of-thought controllability. The full band reads 21.2 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a crossing of the 20-point threshold that is entirely an artifact of the floor. That arithmetic is recomputed rather than assumed on every refresh, and it has come out the same way every time. Three of the seven values are lab claims; the other four come from Artificial Analysis. This is also the one column where two independent evaluators can be compared directly, and the comparison is a caution against reading the top of it too closely: Epoch AI runs its own GPQA Diamond evaluations with published per-run logs, and where it and Artificial Analysis have measured the same model it has agreed to within a point on most of them while putting the retired Claude Fable 5 at 85.9 against AA's 92.6 — a 6.7-point disagreement between two careful evaluators on the same model and the same benchmark, larger than the entire band separating the other seven. That check was being restored from the top down, and it has stalled. Epoch runs GPT-6 Astra at 95.8 and has it leading its own 314-row board, against Artificial Analysis's 96.3 for the same model — a 0.5-point agreement, the closest the two evaluators have come on a current leader — and Gemini 3.8 Flash at 95.4, within a tenth of a point of Artificial Analysis's 95.3, so the two boards agree on both the first and the second model here rather than only on the leader. But the only row Epoch has added since early September 2026 is Claude Opus 5.5 at 90.6, a higher Anthropic tier this page does not track and one Artificial Analysis has not run on this benchmark at all. Below the top two the comparison runs out and is not being restored: Epoch has no row for Anthropic's current flagship, for Meta's, for xAI's or for DeepSeek's. So the two independent boards agree closely wherever both have measured the same model, the 6.7-point Fable 5 disagreement remains the outlier rather than the pattern, and the set of models where a comparison is possible at all is not growing.

Author
Rein et al. · NYU · Cohere
Human baseline
PhD-level experts in matched domains score ~65%; non-experts with web access ~34%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
From the Academic table of OpenAI's GPT-6 Astra post, whose six columns read Astra 96.0, GPT-5.6 Sol 94.6, Claude Fable 5.1 93.7, Claude Fable 5 92.6, Claude Opus 5 93.7 and Gemini 3.8 Flash 95.3 — the same column set every table in that post uses, so there is no Ultra-style extra column to mis-align against. The post's body text adds that at a lower-cost setting Astra scores 94.9, still above Sol's best 94.6, at roughly 37% lower estimated API cost. Scores are the maximum at any effort per the tables' own footer. This is the highest lab-claimed figure on the column and a 1.4-point gain over the retired Sol row.
Model
Score
Notes
The Gemini 3.8 Flash model card reports no GPQA Diamond, the same omission the 3.7 and 3.6 cards had. Its published results table runs to thirteen benchmarks and contains no multiple-choice science evaluation of any kind — the nearest entries are HLE-Verified and two biology-research sets, BioMysteryBench and LABBench2. OpenAI's GPT-6 Astra comparison table does list the model at 95.3, but that is a competitor restating Artificial Analysis's number rather than a figure Google published.
Model
Score
Notes
Meta published no GPQA Diamond figure for Muse Spark 1.3. Its single launch benchmark figure is a chart image with no numbers in the surrounding text, and the companion methodology PDF carries no figures at all, so nothing here is quotable as a lab claim.
Model
Score
Notes
Anthropic publishes no GPQA Diamond figure for Fable 5.1. The system card's capability summary has no GPQA row, and the benchmark survives in the document only as one of three question sets used to measure chain-of-thought controllability. That follows through on the Fable 5 card's stated intent: Anthropic 'consider[s] GPQA Diamond to be a saturated evaluation and plan[s] to stop reporting the performance of future models on it.' OpenAI's GPT-6 Astra comparison table does list Fable 5.1 at 93.7, but that is a competitor restating Artificial Analysis's number rather than a figure Anthropic published.
Model
Score
Notes
GPQA Diamond 92.6, from the Benchmark Results table in the model card for Qwen3.8-2.4T-A95B, the open-weights build Alibaba released nine days after Qwen3.8-Max. That card carries a column headed Qwen3.8-Max because, in its own words, 'Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features'; the row reads Opus 4.8 92.0, Fable 5 92.6, GPT 5.6 Sol (max) 94.1, Qwen3.7-Max 92.4, Qwen3.8-Max 92.6. Alibaba's own launch posts publish no benchmark table at all — their headline claims are LMArena ranks — so the open-weights card is the only text source for this figure.
Model
Score
Notes
GPQA Diamond (Pass@1) 90.9 at maximum reasoning effort, read off the DS-V4.1-Flash column of the model card's frontier-comparison table, where it sits against GPT-5.6 Sol 94.1, Opus-5.0 93.4, Kimi K3 92.9 and GLM-5.3 88.1. The same table restates two retired DeepSeek models at 92.4 and 89.9 — figures produced on this card's harness at maximum effort, which is why the 92.4 it gives DeepSeek-V4-Pro does not match the 90.1 V4-Pro's own card published. Reasoning effort on this generation is a 1-100 integer rather than the three named levels the V4 line used, and every instruct figure on the card is taken at 100.
Model
Score
Notes
Mistral published no GPQA Diamond for Medium 3.5. Its release names two numbers and no others — SWE-Bench Verified 77.6 and τ³-Telecom 91.4 — skipping the whole MMLU / GPQA / AIME / HumanEval / MATH set that most launches lead with.
Model
Score
Not reported
Notes
Grok 4.7's launch post reports no GPQA Diamond. Its comparison table has no multiple-choice science row of any kind; the nearest thing on it is HealthBench Professional at 56.7%, a clinical-reasoning instrument rather than a graduate science exam.

MMMU

MultimodalApproaching saturation

Massive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: Google reports CharXiv Reasoning and LVBench, Anthropic reports Chartography and GDP.pdf, Meta publishes no multimodal figure of any kind, and no current flagship publishes a plain MMMU number. OpenAI's abandonment went a step further in September 2026 — where the GPT-5.6 post at least had a Multimodal table reporting MMMU-Pro, the GPT-6 Astra post has no multimodal table at all, and files its image work under computer use as ScreenSpot-Pro and OSWorld 2.0. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row, and DeepSeek's first natively multimodal flagship arrived in September 2026 with Chartography, BabyVision and ZeroBench in the visual slots and no MMMU row either — the only MMMU string anywhere on its card is an MMMU-Pro figure for the base checkpoint. The abandonment extends to the independent side, and that is the strongest evidence that this column has run its course: Artificial Analysis publishes an MMMU-Pro evaluation and no plain-MMMU one at all, and Epoch AI's benchmark corpus carries no MMMU file either. The benchmark's own authors are the third piece of the same picture, though not in the way the front page suggests. The leaderboard shows a ‘Last updated: 09/05/2025’ stamp and renders its table client-side, but the data file that table is drawn from is live: 211 rows, with entries as recent as July 2026. Every row added since November 2025 carries an MMMU-Pro score and leaves the plain-MMMU cell empty, and the last model to receive a plain-MMMU number there was Claude Opus 4.5. So the authors did not stop maintaining the board — they moved it to the harder variant, which is the same move the labs made. Five of the eight flagships are now charted on Artificial Analysis's MMMU-Pro board — GPT-6 Astra 86.9, Gemini 3.8 Flash 85.6, Qwen3.8-Max 82.8, DeepSeek-V4.1-Flash 77.0 and Mistral Medium 3.5 64.9 — but that is a different and harder instrument, so importing one would break the column's internal comparability. The board is not topped by any of them: Anthropic's Claude Opus 5.5, a tier head this page's roster does not track, leads it at 87.7. Muse Spark 1.3 is the one flagship Artificial Analysis has scored on MMMU-Pro without charting it there: the run it charts by default for the model is the max-effort one, which carries no MMMU-Pro score, and the 82.0 belongs to a non-default xhigh run. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and the strongest removal candidate on the page.

Author
MMMU Benchmark · University of Waterloo + collaborators
Human baseline
No single number: the authors publish three expert tiers, scoring 88.6% at best, 82.6% in the middle and 76.2% at worst on the validation set

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no plain-MMMU number for Fable 5.1. Its multimodal section runs on Chartography, BenchCAD, OSWorld 2.0 and GDP.pdf — no MMMU of any variant, and no MMMU-Pro either.
Model
Score
Not reported
Notes
DeepSeek's first natively multimodal flagship still publishes no plain-MMMU number. The model card's instruct table, which is the source for every DeepSeek figure on this page, has no MMMU row at all; the card's separate base-model table gives MMMU-Pro (EM, 4-shot) 56.5, which is two steps away from this column — a harder variant, and measured on the DeepSeek-V4.1-Flash-Base checkpoint rather than on the instruct model. The card's visual slots are Chartography 78.9, BabyVision 89.6 and ZeroBench-main 49.0, none of which this page tracks.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card uses CharXiv Reasoning at 86.2, LVBench at 87.8 agentic and 87.1 static, and GDP.PDF at 35.0 as its multimodal signals, and reports neither MMMU nor MMMU-Pro. Chart reasoning, long-video understanding and PDF comprehension each test a slice of what MMMU covers; none of them is the college-exam instrument this column tracks.
Model
Score
Not reported
Notes
OpenAI publishes no multimodal benchmark table for Astra at all — a change from the GPT-5.6 post, which had one carrying MMMU-Pro (Sol 83% without tools, 84.6% with) and gdp.pdf. Astra's image-understanding results appear only as computer-use figures, ScreenSpot-Pro 92.7 and OSWorld 2.0 72.6, neither of which is MMMU or a variant of it. So there is no lab figure to consider for this cell, near-miss or otherwise.
Model
Score
Not reported
Notes
Grok 4.7's launch post reports no MMMU. The model accepts text and images and returns text, but xAI published no multimodal benchmark number with it.
Model
Score
Not reported
Notes
Mistral published no MMMU figure for Medium 3.5. Its release describes the model as Mistral's first flagship merged model, handling instruction-following, reasoning and coding in a single set of weights with a vision encoder trained from scratch, but attaches no multimodal benchmark to it.
Model
Score
Not reported
Notes
Meta published no MMMU figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Qwen3.8-Max is natively vision-language, and Alibaba still publishes no plain-MMMU number for it — not in either launch post and not in the open-weights model card, whose multimodal coverage stops at describing the vision input the hosted model adds over the open build.

SWE-Bench Verified

CodingApproaching saturation

The 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025, and now down to a single one of the eight current flagships: Mistral Medium 3.5 at 77.6, and nothing else. Two labs left inside a fortnight. Anthropic's departure was the louder one, because it held the highest figure on the page at 95 and its replacement flagship's system card reports SWE-bench Pro, Multilingual and Multimodal with no Verified row of any kind; DeepSeek's was quieter and cost the column more, because the retired DeepSeek-V4-Pro carried both the higher of the two surviving values at 80.6 and the column's only independent measurement. OpenAI, Google, xAI, Meta and Alibaba had already gone. Alibaba's open-weights model card is the neatest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. One value defines no band at all, so the label below is the one this column carried when it last had enough data to compute one, and it should be read as a marker rather than as a measurement. The independent half is now empty too. Epoch AI runs the benchmark itself and publishes per-run results in its bulk data download, but its 35-row board's only DeepSeek entry is the retired April V4-Pro build at 77.6, and its newest entries from OpenAI, Anthropic, Google and Alibaba are all at least one generation behind those labs' current flagships. That pattern used to be the reason to read the labs' abandonment as substitution rather than neglect — the benchmark was still being run, just not led with. That reading is now harder to sustain: Epoch's board has added nothing since late June 2026, so the independent side is not quietly keeping the benchmark alive either. Aider's polyglot leaderboard, the other board that historically covered this benchmark, has added nothing since two DeepSeek V3.2-Exp runs in October 2025 and carries none of the eight. A column with one lab value, no independent value and no computable band is the strongest removal candidate on this page after MMMU.

Human baseline
No human pass-rate is published; each of the 500 tasks passed a three-annotator screen by professional developers for a well-specified issue and fair tests, and OpenAI expects a typical engineer's solve rate to be below 100%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
From Mistral's own body text: 'Mistral Medium 3.5 scores 77.6% on SWE-Bench Verified, ahead of Devstral 2 and models like Qwen3.5 397B A17B.' The only one of this page's seven benchmarks Mistral publishes a number for at all.
Model
Score
Not reported
Notes
Anthropic no longer reports SWE-bench Verified. The Fable 5.1 system card's software-engineering section covers three variants — SWE-bench Pro, Multilingual and Multimodal — and carries no Verified row anywhere, where the Fable 5 card had published 95. Anthropic's Verified figure was already the secondary, near-saturated number behind its Pro result; dropping it leaves the harder variant as its only SWE-bench claim.
Model
Score
Not reported
Notes
The DeepSeek-V4.1-Flash card has no SWE-bench Verified row, where the retired V4-Pro card published 80.6. The only 'Verified' string anywhere on the card is SimpleQA-Verified, a knowledge benchmark in the base-model table. DeepSeek's coding slots are DeepSWE v1.1, NL2Repo-Bench, ProgramBench and the Terminal-Bench 2.1 / 3.0 / 4.0 ladder.
Model
Score
Not reported
Notes
The Gemini 3.8 Flash model card reports no SWE-Bench variant at all — not Verified, and not Pro.
Model
Score
Not reported
Notes
OpenAI no longer reports SWE-Bench Verified, the benchmark it originally published: Verified was dropped from its launch materials in February 2026, and the Astra post's Coding table now carries no SWE-Bench row of any variant at all — leading instead on Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1.
Model
Score
Not reported
Notes
Grok 4.7's launch post reports no SWE-Bench Verified. xAI has not published a Verified number since before Grok 4.5.
Model
Score
Not reported
Notes
Meta published no SWE-Bench Verified figure for Muse Spark 1.3, and its launch benchmarks are a chart image rather than text.
Model
Score
Not reported
Notes
Alibaba publishes no SWE-Bench Verified number for Qwen3.8-Max. The open-weights model card that supplies this row's other lab figures runs to 31 benchmark rows, including SWE-bench Pro, with no Verified row among them — the cleanest single illustration on this page of the frontier moving wholly to the Pro variant. The retired Qwen3.7-Max claimed 80.4 here, but a predecessor's score is not evidence about this model.

About this page

Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.

Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.

Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.

Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.

Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.

Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.

Last verified: September 29, 2026. 8 frontier models · 7 benchmarks · 8 labs.