AI benchmarks
Every frontier AI model on every major benchmark
Frontier AI benchmark scores as of September 8, 2026: on ARC-AGI-2, GPT-6 Astra leads at 95%; on GPQA, GPT-6 Astra leads at 96%; on SWE Pro, Claude Fable 5.1 leads at 81.2%. Every score below cites the lab’s announcement post or an independent re-runner.
Last verified: September 8, 2026.
How to read this page
● 87.2 — lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.
□ 86.4 — independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.
⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).
— — the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.
Headline matrix
Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.
By benchmark
Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.
ARC-AGI-2
ReasoningDiscriminatingSuccessor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. Two labs now name an ARC-AGI-2 figure in their own launch materials, and in both cases the number matches the ARC Prize Foundation's verified board rather than describing a separate run: Anthropic's Claude Fable 5.1 system card gives ARC-AGI-1 97.5 and ARC-AGI-2 90 and says whose numbers they are, and OpenAI's GPT-6 Astra post carries a full ARC-AGI-1/2/3 table whose 95.0 for Astra is the same figure the Foundation publishes. So every value in this column traces back to one measurement per model, taken from that board at each model's highest verified reasoning-effort variant. Coverage is four of the eight current flagships. That is thin against the two knowledge columns, which both carry all eight, but it now beats SWE-Bench Pro's three — a reversal from earlier in 2026, when the coding column was the most widely reported here and this one was the sparsest. GPT-6 Astra leads at 95.0, then Claude Fable 5.1 90.0, Grok 4.6 67.1 and DeepSeek-V4-Pro 61.3 — a 33.7-point band, the second-widest on this page. Astra's arrival widened it from 31.2 by lifting only the top, which is the pattern this column keeps producing. That spread is the thing to watch: the labs that publish nothing on this benchmark are the ones that score lowest when somebody else measures them. The four flagships without a submission are Gemini 3.8 Flash, Muse Spark 1.3, Mistral Medium 3.5 and Qwen3.8-Max. Google's gap is the fresh one — the retired Gemini 3.7 Flash held 84.6 and its successor has no submission — while the other three labs have never placed a model anywhere but the floor of this board, appearing only through older siblings: two Llama 4 builds and three Magistral builds at zero, and one Qwen3-235B build at 1.3. The board rebuilds irregularly rather than on a schedule — it sat frozen for eleven days across August, then regenerated three times in four days at the start of September — so a missing row means a model has not been run yet, never that it ran and scored nothing.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
Humanity's Last Exam
KnowledgeDiscriminatingCrowdsourced expert-level exam of 2,500 questions across more than a hundred subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools; this column records the no-tools figure. One of two benchmarks on this page with a published value for all eight current flagships — GPQA Diamond is the other — and much the wider band of the two: Claude Fable 5.1 tops the field at 60.9 and Mistral Medium 3.5 anchors it at 13.8, a 47.1-point spread, with the middle six clustered between 42.7 and 54.7. Some of that width is one straggler, but less than it used to be. Drop Mistral and the remaining seven still span 18.2 points, against 12.8 before September 2026, because two labs added points at the top in the same fortnight while the floor did not move: Anthropic's new flagship took the lead to 60.9, and GPT-6 Astra then entered 5.2 points above the retired GPT-5.6 Sol at 54.7. The shape that produced is a ladder rather than a cluster — leader, then Astra 6.2 points back, then a five-model band inside 6 points — so the approaching-saturation reading this column used to invite no longer holds at the top even though it still describes the middle. Three of the eight values are lab claims; the rest come from Artificial Analysis, which runs the full exam on every current flagship it charts. The column's first head-to-head between a lab and an independent evaluator on the same model landed in September 2026, and they agree: Anthropic's 60.9 against Artificial Analysis's 59.1, a 1.8-point gap well inside the divergence threshold. A second evaluator has since arrived on this column and it does not agree at all. Epoch AI's file had for months been a mirror of other people's published results with no current flagship on it; it now carries Claude Fable 5.1 at its extended-reasoning setting, scoring 46.5 and topping Epoch's own board — 12.2 points below what Artificial Analysis measures for the same model at the same setting, and 14.4 below what Anthropic reports for it. That is the widest evaluator-versus-evaluator gap anywhere on this page, twice the 6.7 points that separate the two boards on GPQA Diamond, and it is worth more attention than any of the lab-versus-independent gaps this page flags: the cells here are only comparable to each other because they nearly all come from one evaluator, and the moment a second one measures the same model the column's apparent precision turns out to be a house style. Two caveats on the other end of the table. Google publishes only HLE-Verified (3.8 Flash 54.9), a filtered 1,811-item subset rather than the full exam, so it stays in the cell note rather than in the cell — and the two instruments disagree about direction, since Google's subset shows a 1.3-point generational gain where the full-exam re-run has the new model a twentieth of a point behind its predecessor. Meta's figure is recorded at a max-effort setting Artificial Analysis has charted but Meta says is still pending safety testing.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
SWE-Bench Pro
CodingDiscriminatingContamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases — 731 tasks in the public set, drawn from copyleft repositories to make training-data contamination legally awkward — that the model must patch end-to-end. Down to three of the eight current flagships, all three lab claims run on that lab's own harness rather than on Scale's standardized leaderboard, and no longer the most widely reported column on this page — ARC-AGI-2 passed it in September 2026 by holding four. Claude Fable 5.1's 81.2 sits over a 55–68 pair (Qwen3.8-Max 67.7, DeepSeek-V4-Pro 55.4) — a 25.8-point band, so what remains still discriminates at the production frontier, but on a thinner sample than it used to. Anthropic is the only lab moving on this benchmark rather than away from it: Fable 5.1 raised its predecessor's 80 to 81.2 and added SWE-bench Multilingual and Multimodal alongside it, while OpenAI's GPT-6 Astra dropped the SWE-Bench Pro row GPT-5.6 Sol had claimed at 64.6 and reports Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1 instead, Google's Gemini 3.8 Flash reports no SWE-Bench variant of any kind for a second consecutive release, xAI's Grok 4.6 reports none where the retired Grok 4.5 claimed 64.7, and Meta and Mistral publish nothing here. Google's software-engineering slot is DeepSWE v1.1 and Terminal-Bench instead. The independent side is empty for a structural reason rather than a timing one: Scale AI stewards the benchmark, neither Artificial Analysis nor Epoch AI evaluates it at all, and Scale's own public leaderboard has no row for any of the eight — it is topped by Muse Spark 1.1, two generations behind Meta's current flagship, and its own summary prose still describes top models as scoring around 23% on the public set, which is where the board's methodology write-up was frozen. Cross-lab corroboration is this column's only check, and it thinned twice over in September 2026: OpenAI's and Alibaba's comparison tables had both carried the retired Claude Fable 5's 80, but the one competitor document that names Anthropic's newest flagship — OpenAI's GPT-6 Astra post — carries no SWE-Bench row of any variant, so 81.2 rests on one lab's scaffold until somebody publishes one. The same absence cost the column its cleanest worked example of corroboration, since Alibaba's open-weights card had independently reproduced OpenAI's own 64.6 on its own re-run and there is now no OpenAI figure left to check.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
AIME 2025
MathApproaching saturationAmerican Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for several refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes HMMT 2026 Feb, Anthropic publishes ArXivMath and CritPT-Corrected, and Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. The labs' own framing has hardened over two releases. Anthropic's Fable 5 system card named this benchmark to explain the substitution, introducing USAMO as ‘the next step of the math olympiad track in the US after the AIME, which was a popular AI benchmark last year but is now saturated’; the Fable 5.1 card drops the olympiad track entirely for research-level mathematics, preferring a benchmark drawn monthly from recent arXiv abstracts because it is ‘more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks’. The independent side is empty too, and that is checked rather than assumed: Artificial Analysis still maintains an AIME 2025 leaderboard, but it is both a full generation old — topped by GPT-5.2 at 99.0 — and shrinking, down from seven charted rows to six in September 2026 when its last Anthropic entry dropped off, and it contains none of the eight. Epoch AI publishes no AIME data of any year. That distinction matters for whether the column should survive. A leaderboard that exists but has not been refreshed could pick up a current flagship next month; a benchmark nobody maintains an evaluation page for will not. AIME is in the first category and MMMU is in the second, which is why the two zero-coverage columns on this page are not the same case. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
GPQA Diamond
KnowledgeApproaching saturationGraduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier, and all eight current flagships carry a value. The frontier has effectively cleared it: seven of the eight sit inside a 5.9-point band from DeepSeek-V4-Pro's 90.1 to GPT-6 Astra's 96.0, and Anthropic has now acted on the position its Fable 5 system card set out — that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it — by publishing no GPQA figure at all for Claude Fable 5.1, where the benchmark survives only as one of three question sets used to measure chain-of-thought controllability. The full eight-model band reads 21.2 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a crossing of the 20-point threshold that is entirely an artifact of the floor. That arithmetic is recomputed rather than assumed on every refresh, and it has come out the same way every time it has been run. Only three of the eight values are lab claims; five come from Artificial Analysis. This is also the one column where two independent evaluators can be compared directly, and the comparison is a caution against reading the top of it too closely: Epoch AI runs its own GPQA Diamond evaluations with published per-run logs, and where it and Artificial Analysis have measured the same model it has agreed to within a point on most of them while putting the retired Claude Fable 5 at 85.9 against AA's 92.6 — a 6.7-point disagreement between two careful evaluators on the same model and the same benchmark, larger than the entire band separating the other seven. That check is partly restored at the top of the board: Epoch now runs GPT-6 Astra at 95.8 and has it leading its own 311-row board, against Artificial Analysis's 96.3 for the same model — a 0.5-point agreement, and the closest the two evaluators have come on a current leader. It still cannot be run on the rest of the top, because Epoch's newest Google entry is Gemini 3.7 Flash at 94.8 and its newest Anthropic entry is Claude Opus 5 at 93.9, both a generation behind what Artificial Analysis is charting. So the two boards now agree on who leads this column and disagree about who is second.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
MMMU
MultimodalApproaching saturationMassive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: Google reports CharXiv Reasoning and LVBench, Anthropic reports Chartography and GDP.pdf, Meta publishes no multimodal figure of any kind, and no current flagship publishes a plain MMMU number. OpenAI's abandonment went a step further in September 2026 — where the GPT-5.6 post at least had a Multimodal table reporting MMMU-Pro, the GPT-6 Astra post has no multimodal table at all, and files its image work under computer use as ScreenSpot-Pro and OSWorld 2.0. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row. The abandonment extends to the independent side, and that is the strongest evidence that this column has run its course: Artificial Analysis has no plain-MMMU evaluation page at all — the URL returns a 404 and only MMMU-Pro exists — and Epoch AI's benchmark corpus carries no MMMU file either. The benchmark's own authors are the third piece of the same picture, though not in the way the front page suggests. The leaderboard shows a ‘Last updated: 09/05/2025’ stamp and renders its table client-side, but the data file that table is drawn from is live: 211 rows, with entries as recent as July 2026. Every row added since November 2025 carries an MMMU-Pro score and leaves the plain-MMMU cell empty, and the last model to receive a plain-MMMU number there was Claude Opus 4.5. So the authors did not stop maintaining the board — they moved it to the harder variant, which is the same move the labs made. Three of the eight flagships are currently charted on Artificial Analysis's MMMU-Pro board — GPT-6 Astra 86.9 at the top of it, Gemini 3.8 Flash 85.6 and Mistral Medium 3.5 64.9 — but that is a different and harder instrument, so importing one would break the column's internal comparability. Muse Spark 1.3, charted at 82.0 a cycle ago, has fallen outside that board's eighteen-row window while staying live in Artificial Analysis's registry, which is a windowing artifact rather than a retraction. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and the strongest removal candidate on the page.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
SWE-Bench Verified
CodingApproaching saturationThe 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025, and now down to two of the eight current flagships — DeepSeek-V4-Pro at 80.6 and Mistral Medium 3.5 at 77.6 — after Anthropic dropped it. That departure is the column's clearest signal yet, because Anthropic held the highest figure on the page at 95 and its replacement flagship's system card reports SWE-bench Pro, Multilingual and Multimodal with no Verified row of any kind. OpenAI, Google, xAI, Meta and Alibaba had already left. Alibaba's open-weights model card is the neatest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. With two values the band reads 3 points, which is far too thin a sample to move a saturation label on, so the prior label stands. The column's one independent value comes from Epoch AI rather than from Aider: Epoch runs the benchmark itself and publishes the results in its bulk data download, and its board carries DeepSeek-V4-Pro at 77.6 against that lab's own 80.6 — a 3-point gap, under this page's divergence threshold, and a like-for-like comparison because both figures describe the same April build. Epoch's board is the reason to be careful about reading the labs' abandonment as neglect rather than substitution: it runs to 35 rows and its newest entries from OpenAI, Anthropic, Google and Alibaba are all at least one generation behind those labs' current flagships, which is what you would expect if the benchmark is still being run but is no longer what the launch posts lead with. Aider's polyglot leaderboard, the other board that historically covered this benchmark, has added nothing since two DeepSeek V3.2-Exp runs in October 2025 and carries none of the eight.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
About this page
Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.
Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.
Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.
Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.
Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.
Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.
Last verified: September 8, 2026. 8 frontier models · 7 benchmarks · 8 labs.