AI benchmarks

Every frontier AI model on every major benchmark

Frontier AI benchmark scores as of August 21, 2026: on ARC-AGI-2, GPT-5.6 Sol leads at 92.5%; on GPQA, Grok 4.6 leads at 94.9%; on SWE Pro, Claude Fable 5 leads at 80%. Every score below cites the lab’s announcement post or an independent re-runner.

Last verified: August 21, 2026.

ARC-AGI-2 leader
92.5%
GPQA leader
94.9%
SWE Pro leader
80%
Coverage
8 × 7
frontier models × benchmarks

How to read this page

● 87.2lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.

□ 86.4independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.

⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).

— the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.

Headline matrix

Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.

Anthropic · June 9, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
DeepSeek · April 24, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Google · August 13, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
OpenAI · July 9, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Grok 4.6Closed
xAI · August 12, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Mistral AI · April 28, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Meta · August 5, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified
Alibaba · August 3, 2026
ARC-AGI-2
SWE Pro
AIME
MMMU
SWE Verified

By benchmark

Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.

ARC-AGI-2

ReasoningDiscriminating

Successor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. No frontier lab publishes an ARC-AGI-2 number in its own launch materials any more — OpenAI and Anthropic have both moved their abstract-reasoning slot to ARC-AGI-3 — so every value here comes from the ARC Prize Foundation's own verified leaderboard, recorded at each model's highest verified reasoning-effort variant. Coverage rose from three flagships to four this cycle: the payload was regenerated on August 20 and picked up both of August's new flagships at once. GPT-5.6 Sol still leads at 92.5, Claude Fable 5 lands at 89.2, Gemini 3.7 Flash enters at 84.6 and Grok 4.6 at 67.1 — a 25-point band, tighter than last cycle's 32 but still discriminating. Both newcomers are large generational gains rather than new entries on a blank column: Gemini 3.7 Flash's 84.6 is 24.2 points above the 3.6 Flash it replaced, on a post-training step with no new pre-train and at 41% of the cost per task, and Grok 4.6's 67.1 is 14.5 above the retired Grok 4.5. The remaining four flagships — Muse Spark 1.2, DeepSeek-V4-Pro, Mistral Medium 3.5 and Qwen3.8-Max — have no verified submission at all; the only DeepSeek model on the board is V4-Flash.

Author
François Chollet · ARC Prize Foundation
Human baseline
Every eval task was solved pass@2 by at least two people in a 400+ participant calibration study; ARC Prize sets 85% as the target a system must reach

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
OpenAI's GPT-5.6 GA post reports ARC-AGI-3 (Sol 7.78%) in its 'Abstract reasoning' table instead of ARC-AGI-2; the ARC-AGI-2 row GPT-5.5 carried is not reported for the 5.6 generation.
Model
Score
Notes
Neither the Fable 5 launch post nor the Fable 5 / Mythos 5 system card reports ARC-AGI-2; the system card's Table 8.1.A has no ARC row of any kind. Re-read in full 2026-08-13 — still true.
Model
Score
Notes
The Gemini 3.7 Flash model card does not report ARC-AGI-2, and neither did the 3.6 card. Google's published eval set leads on agentic coding, computer use, and enterprise workflow instead. The retired Gemini 3.6 Flash row's 60.4 was an ARC Prize measurement of the predecessor and is not carried forward.
Model
Score
Notes
Grok 4.6's launch post does not report ARC-AGI-2. Its Evals table is AA Intelligence Index, GDPVal-AA v2, CursorBench v3.2, DeepSWE v1.1, FrontierCode v1.1, APEX-Agents, Terminal-Bench v3.0, APEX-SWE, AA-Briefcase and Harvey LAB — no abstract-reasoning benchmark of any kind.
Model
Score
Not reported
Notes
DeepSeek reports no ARC-AGI benchmark of any kind — neither the model card's two benchmark tables nor the 2026-08-13 GA change-log entry includes one.
Model
Score
Not reported
Notes
Mistral reports no ARC-AGI-2 for Medium 3.5.
Model
Score
Not reported
Notes
Meta's Muse Spark 1.2 launch post reports no ARC-AGI benchmark, and every evaluation it does report is a chart image — so no lab cell on this row is text-verifiable.
Model
Score
Not reported
Notes
Neither Alibaba's Qwen3.8-Max launch posts nor the 31-row benchmark table in the Qwen3.8-2.4T-A95B open-weights model card reports any ARC-AGI benchmark.

Humanity's Last Exam

KnowledgeDiscriminating

Crowdsourced expert-level exam across 100+ subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools. One of two benchmarks on this page with a published value for all eight current flagships — GPQA Diamond is the other — and by far the wider band of the two: Claude Fable 5 tops the field at 55.5 and Mistral Medium 3.5 anchors it at 13.8, a 42-point spread with the middle six clustered between 41.0 and 49.5. It is also the most volatile column here, and three refresh cycles running have proved it. Artificial Analysis re-ran most of the board in the second week of August 2026 and revised five independent values upward at once (Fable 5 53.3 to 55.5, GPT-5.6 Sol 47.2 to 49.5, Gemini 38.3 to 40.8, DeepSeek 35.9 to 39.3, Mistral 12.8 to 13.8); DeepSeek separately restated its own claim from 37.7 to 42.7 when V4-Pro left preview; and the third week moved DeepSeek's independent value again, 39.3 to 41.0. Google's entry changed for a different reason — Gemini 3.7 Flash replaced 3.6 Flash on August 13 and enters at 47.9, seven points above its predecessor. A value carried forward from a prior cycle on this column is almost certainly wrong. One caveat on Google specifically: the only HLE-shaped number Google publishes is HLE-Verified (3.7 Flash 53.6), a filtered subset rather than the full exam, so it is recorded in the cell note rather than in the cell.

Author
Center for AI Safety · Scale AI
Human baseline
Expert humans in their own domains score ~88%; broad humans far below

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
NULLED 2026-07-31 (was 59.0). The system card's Table 8.1.A gives Humanity's Last Exam as Mythos 5 59.0 no-tools / 64.5 with-tools, Mythos Preview 56.8 / 64.7, Opus 4.8 49.8 / 57.9 — and an em-dash in the Fable 5 column for both configurations. The 59.0 this page carried was Mythos 5's number, not Fable 5's. Section 8.14.1 confirms the HLE runs were done on Mythos 5.
Model
Score
Notes
OpenAI's GPT-5.6 GA post reports no Humanity's Last Exam figure for any GPT-5.6 tier — the academic slot is GPQA Diamond plus FrontierMath Tier 1-3 / Tier 4.
Model
Score
Notes
Google publishes no plain Humanity's Last Exam number for Gemini 3.7 Flash. The model card does report 'HLE-Verified' at 53.6% (against 3.6 Flash 51.2, Claude Sonnet 5 31.0, GPT-5.6 Terra 51.1) — but HLE-Verified is a filtered subset, not the full exam this column tracks, so it is recorded here as context rather than promoted into the cell. That is the same variant discipline this page applies to SWE-Bench Pro vs. Verified and MMMU vs. MMMU-Pro. Note the direction of the gap: on Google's own subset 3.7 Flash gains 2.4 points over 3.6 Flash, while Artificial Analysis's full-exam re-run shows a 7.1-point gain — a reminder that the two are not interchangeable.
Model
Score
Notes
Not text-verifiable in the Muse Spark 1.2 launch post (charts are images).
Model
Score
Notes
NEWLY POPULATED 2026-08-13 (Mimas). HLE 43.6 for Qwen3.8-Max, from the Qwen3.8-2.4T-A95B open-weights model card's 'Benchmark Results' table: Opus 4.8 45.7, Fable 5 53.3, GPT 5.6 Sol (max) 47.2, Qwen3.7-Max 41.4, Qwen3.8-Max 43.6. The same table gives 'HLE w/ tools' 56.2; this column records the no-tools figure. Neither August 3 launch post publishes an HLE number.
Model
Score
Notes
Grok 4.6's launch post does not report Humanity's Last Exam.
Model
Score
Notes
REVISED 2026-08-13 (Mimas) from 37.7 to 42.7, and the citation moved from the HuggingFace model card to DeepSeek's own Change Log. The change log's dated 2026-08-13 entry announcing the GA rollout of DeepSeek-V4-Pro states in body text 'HLE (wo / w tools): 42.7/60.0' — the no-tools figure is what this column records. That supersedes the model card's 37.7, which is the April preview build's Max-effort number (mode ladder Non-Think 7.7, High 34.5, Max 37.7, with-tools 48.2) and which the card still shows unchanged, because the card has not been re-issued since 2026-06-22. Newer primary source wins. For the record, prior cycles before 2026-07-31 carried 23.4 here, a figure that has never appeared in any DeepSeek artifact.
Model
Score
Notes
Not reported in the Medium 3.5 release post.

SWE-Bench Pro

CodingDiscriminating

Contamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases that the model must patch end-to-end. Still the most widely reported column on this page, but its coverage is now visibly eroding — four of the eight current flagships carry a value, down from five, and all four are lab claims run on the lab's own harness rather than Scale's standardized leaderboard. Claude Fable 5's 80 sits over a 55–68 cluster (Qwen3.8-Max 67.7, GPT-5.6 Sol 64.6, DeepSeek-V4-Pro 55.4) — a ~25pp band, so what remains still discriminates at the production frontier. Three consecutive August flagship swaps each took a cell off this column or replaced one: Alibaba added Qwen3.8-Max 67.7 in the open-weights card it posted on August 12; xAI's Grok 4.6 (August 12) reports no SWE-Bench number where the retired Grok 4.5 claimed 64.7; and Google's Gemini 3.7 Flash (August 13) does the same, dropping the 58.7 the 3.6 Flash card had published only three weeks earlier. Google's software-engineering slot moved to DeepSWE v1.1 and Terminal-bench instead. Meta and Mistral publish nothing here, and Scale AI's public leaderboard still has no row for any of the eight.

Human baseline
Human SWE pass-rate on a comparable subset estimated ~75% by the benchmark authors

Saturation: Frontier scores still span a wide band — this benchmark separates the labs.

Model
Score
Notes
CORRECTED 2026-07-31 from 80.3 to 80. The system card's Table 8.1.A reads Mythos 5 80.3, Fable 5 80, Mythos Preview 77.8, Opus 4.8 69.2, GPT-5.5 58.6, Gemini 3.1 Pro 54.2 — the 80.3 prior cycles carried was Mythos 5's column. OpenAI's independent GPT-5.6 comparison table also lists Claude Fable 5 at 80%. Anthropic scaffold figure; not directly comparable to Scale's standardized SEAL leaderboard.
Model
Score
Notes
NEWLY POPULATED 2026-08-13 (Mimas). SWE-bench Pro 67.7 for Qwen3.8-Max, from the Qwen3.8-2.4T-A95B open-weights model card's 'Benchmark Results' table: Opus 4.8 69.2, Fable 5 80.0, GPT 5.6 Sol (max) 64.6, Qwen3.7-Max 60.6, Qwen3.8-Max 67.7. The card's footnote 3 gives the configuration: 'Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks corrected and all baselines evaluated on the refined benchmark.' Alibaba's own harness, so not directly comparable to Scale's standardized SEAL leaderboard. The August 3 technical write-up published only multi-day autonomous-coding case studies (the oh-my-cli project) with no numeric table, which is why this cell was null last cycle.
Model
Score
Notes
SWE-Bench Pro 64.6 per the GPT-5.6 GA post's Coding table, re-read live 2026-08-13: Sol 64.6, Terra 63.4, Luna 62.7, GPT-5.5 59.4, Mythos 5 80.3, Mythos Preview 77.8, Fable 5 80, Opus 4.8 69.2, Gemini 3.1 Pro 54.2. OpenAI's own harness figure; not directly comparable to Scale's standardized SEAL leaderboard. Alibaba's August 2026 model card independently cites GPT-5.6 Sol (max) at 64.6 on its own re-run of the refined benchmark, which corroborates the number from a second lab.
Model
Score
Notes
'SWE Pro (Resolved)' 55.4 for DS-V4-Pro Max per the official model card's frontier-comparison table (Opus-4.6 Max 57.3, GPT-5.4 xHigh 57.7, Gemini-3.1-Pro High 54.2, K2.6 Thinking 58.6, GLM-5.1 58.4, DS-V4-Pro Max 55.4; mode table Non-Think 52.1, High 54.4, Max 55.4). Internal DeepSeek agent scaffold with problematic tasks corrected on the refined benchmark. Re-read live 2026-08-13 and unchanged. CAVEAT: like the GPQA cell, this is the April preview build's figure — the 2026-08-13 GA change-log entry for DeepSeek-V4-Pro-0813 publishes DeepSWE 62.7 and Terminal Bench 2.1 87.9 for its coding slots but does not restate SWE-Bench Pro.
Model
Score
Not reported
Notes
NULLED 2026-08-21 by the flagship swap, not by a revision. The Gemini 3.6 Flash row carried a lab claim of 58.7 (SWE-Bench Pro Public, from the 3.6 model card). The Gemini 3.7 Flash model card, published 2026-08-13 and read live this run, reports NO SWE-Bench benchmark of any kind — its software-engineering slot is DeepSWE v1.1 (65.3 vs. 3.6 Flash's 48.6), FrontierCode 1.1 Main (43.6 vs. 34.4), Terminal-bench 2.1 (85.8 vs. 78.0) and Terminal-bench 3.0 (14.9 vs. 5.4). 58.7 belongs to the predecessor and is not carried forward; the correct output is an empty cell.
Model
Score
Not reported
Notes
NULLED 2026-08-13 with the row swap. The retired Grok 4.5 row carried a lab-claimed 64.7 here, sourced from that launch post's text chart description. Grok 4.6's launch post publishes no SWE-Bench number at all — its coding slots are DeepSWE v1.1 65.9, FrontierCode v1.1 (Extended) 61.3, APEX-SWE 56.4, CursorBench v3.2 69.9 and Terminal-Bench v3.0 26.0 — and a predecessor's score is not evidence about this model.
Model
Score
Not reported
Notes
Mistral published SWE-Bench Verified but not SWE-Bench Pro for Medium 3.5.
Model
Score
Not reported
Notes
Meta published no SWE-Bench Pro number for Muse Spark 1.2; the coding evaluations it did publish are Terminal-Bench 2.1, DeepSWE v1.1 and an internal 440-task bench, all as chart images.

AIME 2025

MathApproaching saturation

American Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for two refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes HMMT 2026 Feb, and Anthropic, Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. Anthropic's Fable 5 system card says why in passing: AIME ‘was a popular AI benchmark last year but is now saturated’. With zero published values the band is undefined, so the label below is the one this column carried when it last had data. The column is kept for continuity and is a candidate for removal.

Author
Mathematical Association of America
Human baseline
Strong high-school competitors solve ~50%; AIME qualifiers (top USAMO contenders) ~80%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no competition-math benchmark for Fable 5. The system card's academic slot is HLE plus FrontierCode; there is no AIME row in Table 8.1.A.
Model
Score
Not reported
Notes
NULLED 2026-07-31 (was 91.8), still null 2026-08-13. DeepSeek's official model card reports no AIME of any year — the competition-math slot is HMMT 2026 Feb (V4-Pro Max 95.2) and IMOAnswerBench (89.8) — and the 2026-08-13 GA change-log entry adds no math benchmark either. The 91.8 was read off the launch chart image by an earlier cycle and is not text-verifiable in any DeepSeek artifact.
Model
Score
Not reported
Notes
The Gemini 3.7 Flash model card reports no competition-math benchmark, matching the 3.6 card.
Model
Score
Not reported
Notes
OpenAI's GPT-5.6 GA post reports FrontierMath (Sol: Tier 1-3 89% / Tier 4 83%) instead of AIME 2025.
Model
Score
Not reported
Notes
Grok 4.6's launch post reports no competition-math benchmark; the launch is agentic-coding and knowledge-work first.
Model
Score
Not reported
Notes
Mistral did not publish AIME 2025 for Medium 3.5.
Model
Score
Not reported
Notes
Not reported for Muse Spark 1.2 (launch charts are images).
Model
Score
Not reported
Notes
Alibaba publishes no AIME 2025 number for Qwen3.8-Max, and the 2026-08-12 open-weights model card carries no competition-math row of any kind.

GPQA Diamond

KnowledgeApproaching saturation

Graduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier, and all eight current flagships carry a value. The frontier has effectively cleared it: seven of the eight sit inside a 4.8-point band from DeepSeek-V4-Pro's 90.1 to Grok 4.6's 94.9, and Anthropic's Claude Fable 5 / Mythos 5 system card states outright that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it on future models. The full eight-model band technically reads 20.1 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a hairline crossing of the 20-point threshold.

Author
Rein et al. · NYU · Cohere
Human baseline
PhD-level experts in matched domains score ~65%; non-experts with web access ~34%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
Grok 4.6's launch post does not report GPQA Diamond.
Model
Score
Notes
GPQA Diamond 94.6 per the GPT-5.6 GA post's Academic table, re-read live 2026-08-13: Sol 94.6, Terra 92.9, Luna 92.3, GPT-5.5 93.6, Mythos 5 94.1, Mythos Preview 94.6, Fable 5 92.6, Opus 4.8 92%, Gemini 3.1 Pro Preview 94.3. Unchanged for a third cycle running.
Model
Score
Notes
The Gemini 3.7 Flash model card does not report GPQA Diamond, the same omission the 3.6 card had. Its 20-row eval table has no multiple-choice science benchmark at all.
Model
Score
Notes
NULLED 2026-07-31 (was 92.6). Anthropic's Fable 5 / Mythos 5 system card section 8.8 reports GPQA Diamond for MYTHOS 5 only (94.1%, averaged over 5 trials) and gives no Fable 5 figure; the Table 8.1.A summary has no GPQA row at all. The 92.6 that circulates for Fable 5 is not text-verifiable in any Anthropic artifact — it appears in the launch post's benchmark PNG and in OpenAI's GPT-5.6 comparison table, and it is what Artificial Analysis measures independently, so it is recorded in this cell as the independent value instead. The same system-card section states Anthropic 'consider[s] GPQA Diamond to be a saturated evaluation and plan[s] to stop reporting the performance of future models on it.'
Model
Score
Notes
NEWLY POPULATED 2026-08-13 (Mimas). GPQA Diamond 92.6 for Qwen3.8-Max, from the 'Benchmark Results' table in the Qwen3.8-2.4T-A95B open-weights model card Alibaba posted on 2026-08-12: Opus 4.8 92.0, Fable 5 92.6, GPT 5.6 Sol (max) 94.1, Qwen3.7-Max 92.4, Qwen3.8-Max 92.6. Neither August 3 launch post publishes a GPQA number — Alibaba's headline claims there are LMArena ranks — so the open-weights card is the first text-verifiable Alibaba source for this cell.
Model
Score
Notes
Not text-verifiable in the Muse Spark 1.2 launch post (benchmark charts are images).
Model
Score
Notes
GPQA Diamond (Pass@1) 90.1 for DS-V4-Pro Max, per the official DeepSeek-V4-Pro model card's 'DeepSeek-V4-Pro-Max vs Frontier Models' table (mode ladder: Non-Think 72.9, High 89.1, Max 90.1). CAVEAT ADDED 2026-08-13: this is the April preview build's figure. The model card was re-fetched live and is unchanged (HF API lastModified 2026-06-22), and DeepSeek's 2026-08-13 GA change-log entry for DeepSeek-V4-Pro-0813 does not restate GPQA, so there is no newer lab number to move to — but the value describes the pre-GA weights, not the build behind the `deepseek-v4-pro` id today.
Model
Score
Notes
Mistral did not publish GPQA Diamond for Medium 3.5. The launch post's body text names only SWE-Bench Verified (77.6) and tau-cubed-Telecom (91.4); multiple third-party reviewers noted at the time that Mistral skipped MMLU / GPQA / AIME / HumanEval / MATH for this release.

MMMU

MultimodalApproaching saturation

Massive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: OpenAI reports MMMU-Pro, Google reports CharXiv Reasoning, Anthropic reports GDP.pdf vision, and no current flagship publishes a plain MMMU number. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row. With zero published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.

Author
MMMU Benchmark · University of Waterloo + collaborators
Human baseline
College students with web access ~83%; domain experts ~90%

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Not reported
Notes
Anthropic reports no plain-MMMU number for Fable 5. Multimodal capability is described via GDP.pdf vision (Fable 5 leads at 29.8) and CharXiv Reasoning rather than MMMU.
Model
Score
Not reported
Notes
NULLED 2026-07-31 (was 76.4), still null 2026-08-13. Neither of the official model card's benchmark tables contains an MMMU row, and neither the announcement post nor the GA change-log entry names one. The V-series is text-and-code-first; the 76.4 was a chart read, not a DeepSeek claim.
Model
Score
Not reported
Notes
The Gemini 3.7 Flash model card reports CharXiv Reasoning (84.5% no tools / 88.7% with tools) and LVBench 85.4% as its multimodal signals — no MMMU and no MMMU-Pro. CharXiv is also the card's one acknowledged regression against 3.6 Flash (84.5 vs. 85.2 no-tools, 88.7 vs. 89.4 with tools).
Model
Score
Not reported
Notes
OpenAI's GPT-5.6 GA post reports MMMU-Pro (Sol 83% no tools / 84.6% with tools) and gdp.pdf, not the plain MMMU benchmark tracked here.
Model
Score
Not reported
Notes
Grok 4.6's launch post does not report MMMU. The model takes text and image input (text-only output) but no MMMU number was published.
Model
Score
Not reported
Notes
Mistral did not publish MMMU for Medium 3.5. The launch post describes it as the first Mistral flagship to merge multimodal vision with chat in a single weights set, but names no MMMU number.
Model
Score
Not reported
Notes
Not reported for Muse Spark 1.2 (launch charts are images).
Model
Score
Not reported
Notes
Qwen3.8-Max is natively vision-language but Alibaba publishes no plain-MMMU number for it — not in either August 3 launch post and not in the 2026-08-12 open-weights model card, whose multimodal coverage is limited to the vision capabilities the hosted Max adds over the text-only open build.

SWE-Bench Verified

CodingApproaching saturation

The 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025. Three of the eight current flagships still report it — Claude Fable 5 at 95, then DeepSeek-V4-Pro 80.6 and Mistral Medium 3.5 77.6 — while OpenAI, Google, xAI, Meta and Alibaba have all dropped it from their launch materials in favor of the harder Pro variant. Alibaba's August 2026 open-weights model card is the clearest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. With three values the band reads 17.4 points, but three cells out of eight is too thin a sample to move a saturation label on, so the prior cycle's label stands.

Human baseline
Subset construction targets human-solvable issues; success rate not directly comparable to model pass@1

Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.

Model
Score
Notes
CORRECTED 2026-07-31 from 95.5 to 95. Table 8.1.A: Mythos 5 95.5, Fable 5 95, Mythos Preview 93.9, Opus 4.8 88.6, Gemini 3.1 Pro 80.6 — 95.5 was Mythos 5's column. The harder Pro variant is the launch headline; Verified is the secondary, near-saturated number.
Model
Score
Notes
'SWE Verified (Resolved)' 80.6 for DS-V4-Pro Max per the official model card, on par with Opus-4.6 Max (80.8) and Gemini-3.1-Pro High (80.6). Mode table: Non-Think 73.6, High 79.4, Max 80.6. Re-read live 2026-08-13 and unchanged. Same preview-build caveat as the GPQA and SWE Pro cells: DeepSeek's 2026-08-13 GA note does not restate it.
Model
Score
Notes
Per the launch post body text, re-read live 2026-08-13: "Mistral Medium 3.5 scores 77.6% on SWE-Bench Verified, ahead of Devstral 2 and models like Qwen3.5 397B A17B." Unchanged, and still the only benchmark of this page's seven that Mistral publishes a number for.
Model
Score
Not reported
Notes
The Gemini 3.7 Flash model card reports no SWE-Bench variant at all — not Verified, and unlike the 3.6 card not Pro either.
Model
Score
Not reported
Notes
OpenAI's GPT-5.6 GA post does not report SWE-Bench Verified; its Coding table leads on SWE-Bench Pro, DeepSWE v1.1, and Terminal-Bench 2.1. OpenAI deprecated Verified in February 2026.
Model
Score
Not reported
Notes
Grok 4.6's launch post does not report SWE-Bench Verified. xAI has not published a Verified number since before Grok 4.5.
Model
Score
Not reported
Notes
Not reported for Muse Spark 1.2 (launch charts are images).
Model
Score
Not reported
Notes
Not published for Qwen3.8-Max. The 2026-08-12 open-weights model card that recovered this row's GPQA, HLE and SWE-Bench Pro cells has 31 benchmark rows and no SWE-bench Verified among them — Alibaba has moved wholly to the Pro variant. The retired Qwen3.7-Max row carried a lab-claimed 80.4 here; it is not carried forward, because a predecessor's score is not evidence about this model.

About this page

Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.

Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.

Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.

Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.

Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.

Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.

Last verified: August 21, 2026. 8 frontier models · 7 benchmarks · 8 labs.