AI benchmarks
Every frontier AI model on every major benchmark
Frontier AI benchmark scores for each lab’s current flagship, as of September 29, 2026: on ARC-AGI-2, GPT-6 Astra leads at 95%; on GPQA, GPT-6 Astra leads at 96%; on SWE Pro, Claude Fable 5.1 leads at 81.2%. Every score below cites the lab’s announcement post or an independent re-runner.
Last verified: September 29, 2026.
How to read this page
● 87.2 — lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.
□ 86.4 — independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.
⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).
— — the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.
Headline matrix
Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.
By benchmark
Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.
ARC-AGI-2
ReasoningDiscriminatingSuccessor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. Two labs now name an ARC-AGI-2 figure in their own launch materials, and in both cases the number matches the ARC Prize Foundation's verified board rather than describing a separate run: Anthropic's Claude Fable 5.1 system card gives ARC-AGI-1 97.5 and ARC-AGI-2 90 and says whose numbers they are, and OpenAI's GPT-6 Astra post carries a full ARC-AGI-1/2/3 table whose 95.0 for Astra is the same figure the Foundation publishes. So every value in this column traces back to one measurement per model, taken from that board at each model's highest verified reasoning-effort variant. Coverage is three of the eight current flagships: GPT-6 Astra leads at 95.0, Claude Fable 5.1 follows at 90.0, and Gemini 3.8 Flash — which entered the board on September 24, 2026 without Google publishing anything — sits just under at 89.2. That is a 5.8-point band across three cells, far too thin a sample to move a saturation label on, and more a fact about which labs happen to have a current board row than about how far apart the field is. The band moved four times in a month for reasons that have nothing to do with capability: GPT-6 Astra's arrival widened it from 31.2 by lifting only the top, DeepSeek's flagship swap narrowed it to 27.9 by removing the bottom, xAI's swap collapsed it to 5.0 by removing the bottom again, and the September 24 rebuild widened it slightly to 5.8 by adding a row near the floor. The five flagships still without a submission are Grok 4.7, DeepSeek-V4.1-Flash, Muse Spark 1.3, Mistral Medium 3.5 and Qwen3.8-Max. The board has rebuilt twice since Grok 4.7 shipped — on September 24 and again on September 28, 2026, when it added OpenAI's GPT-6 Sol — and neither rebuild ran Grok 4.7 or DeepSeek-V4.1-Flash, which had been out for eighteen days by the second. Whether that is a submission not made or a queue not reached is not something the board publishes, and the page does not guess. The other three labs have never placed a model anywhere but the floor of this board, appearing only through older siblings: two Llama 4 builds and three Magistral builds at zero, and one Qwen3-235B build at 1.25. The board rebuilds irregularly rather than on a schedule, so a missing row means a model has not been run, never that it ran and scored nothing.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
Humanity's Last Exam
KnowledgeDiscriminatingCrowdsourced expert-level exam of 2,500 questions across more than a hundred subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools; this column records the no-tools figure. The only benchmark on this page with a published value for all eight current flagships, and by some way the widest band of any column here: Claude Fable 5.1 tops this page's flagships at 60.9 and Mistral Medium 3.5 anchors them at 13.8, a 47.1-point spread, with the middle six between 36.8 and 54.7. The top of this column is not the top of the field: Anthropic's Claude Opus 5.5, a higher tier this page's roster does not track, publishes 64.4 in its own system card. HLE took the all-eight distinction outright in September 2026, when GPQA Diamond — which had also carried all eight — lost a cell to xAI's flagship swap and has not got it back. Some of the width is one straggler, but less than it used to be. Drop Mistral and the remaining seven span 24.1 points, against 12.8 before September 2026, and both ends moved to get there: Anthropic's new flagship took the lead to 60.9, GPT-6 Astra entered 5.2 points above the retired GPT-5.6 Sol at 54.7, and DeepSeek then swapped a 42.7 for its successor's 36.8 — the only flagship swap on this page to lower a lab's own HLE figure. The shape that produced is a ladder rather than a cluster: leader, Astra 6.2 points back, a four-model band inside 5.6 points, DeepSeek a further 6.3 below that, then Mistral far under everything. Three of the eight values are lab claims; the rest come from Artificial Analysis, which runs the full exam on every current flagship it charts. Three rows carry both halves, and all three agree well inside the divergence threshold — Anthropic 60.9 against 59.1, Alibaba 43.6 against 43.1, DeepSeek 36.8 against 39.2, the last being the rare case where the independent evaluator is the more generous of the two. A second independent measurement exists for three of the eight, and it comes from the benchmark's own authors: Scale AI's official Humanity's Last Exam leaderboard, run on every public question with a fixed judge model, lists Claude Fable 5.1 at 46.5, GPT-6 Astra at 54.8 and Gemini 3.8 Flash at 44.5, and Epoch AI republishes the same figures. Set against Artificial Analysis's numbers for the same models, those are gaps of 12.2, 0.1 and 3.3 points, so the disagreement is specific to one model rather than a systematic difference between the two harnesses. The Anthropic result — 12.2 points below Artificial Analysis and 14.4 below Anthropic's own figure — is still the widest evaluator-versus-evaluator gap anywhere on this page, nearly twice the 6.7 points separating the two GPQA Diamond boards on the retired Claude Fable 5, and neither board explains it. The authors released HLE-Diamond on September 22, 2026, a refined 1,000-question subset whose first results run from GPT-6 Astra's 60.6 down to Grok 4.7's 23.4; it is a different instrument and stays out of this column. So does Google's HLE-Verified (3.8 Flash 54.9), a filtered 1,811-item subset rather than the full exam — and the two instruments disagree about direction, since Google's subset shows a 1.3-point generational gain where the full-exam re-run has the new model a twentieth of a point behind its predecessor.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
SWE-Bench Pro
CodingDiscriminatingContamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases — 731 tasks in the public set, drawn from copyleft repositories to make training-data contamination legally awkward — that the model must patch end-to-end. Down to two of the eight current flagships, both lab claims run on that lab's own harness rather than on Scale's standardized leaderboard, and no longer the most widely reported column on this page — ARC-AGI-2 passed it in September 2026, briefly fell level at two apiece when xAI's flagship swap cost that column a cell, and pulled ahead again on September 24 when the ARC Prize board rebuilt and picked up Gemini 3.8 Flash. Claude Fable 5.1's 81.2 sits over Qwen3.8-Max's 67.7 and nothing else: a 13.5-point band on two cells, which is far too thin a sample to move a saturation label on, so the label below is the one this column carried when it had enough data to compute one. Anthropic is the only lab moving on this benchmark rather than away from it: Fable 5.1 raised its predecessor's 80 to 81.2 and added SWE-bench Multilingual and Multimodal alongside it, and two Anthropic tiers this page's roster does not track have since gone past it in their own system cards — Claude Opus 5.5 at 89.9 and Claude Sonnet 5.5 at 81.3 — so the top of this column is the top of the tracked flagships, not of the field. Everyone else moved away: OpenAI's GPT-6 Astra dropped the SWE-Bench Pro row GPT-5.6 Sol had claimed at 64.6 and reports Terminal-Bench 4.0, DeepSWE v1.1 and FrontierCode 1.1 instead, DeepSeek-V4.1-Flash dropped the 55.4 the retired V4-Pro claimed and reports DeepSWE v1.1, NL2Repo-Bench and ProgramBench, Google's Gemini 3.8 Flash reports no SWE-Bench variant of any kind for a second consecutive release, xAI's Grok 4.7 reports none and nor did Grok 4.6 before it — the last xAI figure in this column was the retired Grok 4.5's 64.7, two flagships back — and Meta and Mistral publish nothing here. Three of the six departures happened inside a single fortnight, and each of the three replaced SWE-Bench Pro with an agentic-terminal or repository-scale benchmark rather than with nothing — the substitution is toward longer-horizon coding work, not away from measuring it. The independent side is empty for a structural reason rather than a timing one: Scale AI stewards the benchmark, neither Artificial Analysis nor Epoch AI evaluates it at all, and Scale's own leaderboard for the original task set — one board for the public dataset and one for the private commercial subset — has no row for any of the eight. Both boards are topped by Muse Spark 1.1, two generations behind Meta's current flagship, and the page's summary prose still describes top models as scoring around 23% on the public set. In September 2026 Scale added SWE-Bench Pro V2, a refreshed public split with a modified task set and a locked evaluation protocol, which does carry three of this page's flagships — Claude Fable 5.1 at 99.1, GPT-6 Astra at 96.9 and Gemini 3.8 Flash at 94.9 on its full split, and 92.2, 90.2 and 58.8 on its hard split, each on its own lab's agent harness. Those figures come from a different instrument from the one every lab number in this column was measured on, so they stay out of the cells rather than being read as independent checks on the lab claims. Cross-lab corroboration is therefore this column's only check, and it thinned twice over in September 2026: OpenAI's and Alibaba's comparison tables had both carried the retired Claude Fable 5's 80, but the one competitor document that names Anthropic's newest flagship — OpenAI's GPT-6 Astra post — carries no SWE-Bench row of any variant, so 81.2 rests on one lab's scaffold until somebody publishes one. The same absence cost the column its cleanest worked example of corroboration, since Alibaba's open-weights card had independently reproduced OpenAI's own 64.6 on its own re-run and there is now no OpenAI figure left to check.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
AIME 2025
MathApproaching saturationAmerican Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for several refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes MathArena Apex and a Codeforces rating, Anthropic publishes ArXivMath and CritPT-Corrected, and Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. The labs' own framing has hardened over two releases. Anthropic's Fable 5 system card named this benchmark to explain the substitution, introducing USAMO as ‘the next step of the math olympiad track in the US after the AIME, which was a popular AI benchmark last year but is now saturated’; the Fable 5.1 card drops the olympiad track entirely for research-level mathematics, preferring a benchmark drawn monthly from recent arXiv abstracts because it is ‘more realistic and more closely connected to mathematical research than contest or Olympiad benchmarks’. The independent side is empty too, and that is checked rather than assumed: Artificial Analysis still maintains an AIME 2025 leaderboard, but it is both a full generation old — topped by GPT-5.2 at 99.0 — and steadily shrinking, down from seven charted rows to six and then to five across September 2026 as its last Anthropic entry and then its floor row dropped off, and it contains none of the eight — every one of the five survivors is a model Artificial Analysis itself now marks deprecated. It loses rows from the bottom while the top stays put, which is what a board nobody is submitting to looks like. Epoch AI publishes no AIME data of any year. That distinction matters for whether the column should survive. A leaderboard that exists but has not been refreshed could pick up a current flagship next month; a benchmark nobody maintains an evaluation page for will not. AIME is in the first category and MMMU is in the second, which is why the two zero-coverage columns on this page are not the same case. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
GPQA Diamond
KnowledgeApproaching saturationGraduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier. Seven of the eight current flagships carry a value; Grok 4.7 does not, and that gap is new — xAI published no GPQA figure with it, and Artificial Analysis, which had scored every previous Grok, has ingested the model without running this benchmark on it. So this column lost the all-eight coverage it shared with Humanity's Last Exam, and HLE now holds that alone. The frontier has effectively cleared what remains: six of the seven sit inside a 5.1-point band from DeepSeek-V4.1-Flash's 90.9 to GPT-6 Astra's 96.0, and Anthropic has acted on the position its Fable 5 system card set out — that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it — by publishing no GPQA figure at all for Claude Fable 5.1, where the benchmark survives only as one of three question sets used to measure chain-of-thought controllability. The full band reads 21.2 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a crossing of the 20-point threshold that is entirely an artifact of the floor. That arithmetic is recomputed rather than assumed on every refresh, and it has come out the same way every time. Three of the seven values are lab claims; the other four come from Artificial Analysis. This is also the one column where two independent evaluators can be compared directly, and the comparison is a caution against reading the top of it too closely: Epoch AI runs its own GPQA Diamond evaluations with published per-run logs, and where it and Artificial Analysis have measured the same model it has agreed to within a point on most of them while putting the retired Claude Fable 5 at 85.9 against AA's 92.6 — a 6.7-point disagreement between two careful evaluators on the same model and the same benchmark, larger than the entire band separating the other seven. That check was being restored from the top down, and it has stalled. Epoch runs GPT-6 Astra at 95.8 and has it leading its own 314-row board, against Artificial Analysis's 96.3 for the same model — a 0.5-point agreement, the closest the two evaluators have come on a current leader — and Gemini 3.8 Flash at 95.4, within a tenth of a point of Artificial Analysis's 95.3, so the two boards agree on both the first and the second model here rather than only on the leader. But the only row Epoch has added since early September 2026 is Claude Opus 5.5 at 90.6, a higher Anthropic tier this page does not track and one Artificial Analysis has not run on this benchmark at all. Below the top two the comparison runs out and is not being restored: Epoch has no row for Anthropic's current flagship, for Meta's, for xAI's or for DeepSeek's. So the two independent boards agree closely wherever both have measured the same model, the 6.7-point Fable 5 disagreement remains the outlier rather than the pattern, and the set of models where a comparison is possible at all is not growing.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
MMMU
MultimodalApproaching saturationMassive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: Google reports CharXiv Reasoning and LVBench, Anthropic reports Chartography and GDP.pdf, Meta publishes no multimodal figure of any kind, and no current flagship publishes a plain MMMU number. OpenAI's abandonment went a step further in September 2026 — where the GPT-5.6 post at least had a Multimodal table reporting MMMU-Pro, the GPT-6 Astra post has no multimodal table at all, and files its image work under computer use as ScreenSpot-Pro and OSWorld 2.0. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row, and DeepSeek's first natively multimodal flagship arrived in September 2026 with Chartography, BabyVision and ZeroBench in the visual slots and no MMMU row either — the only MMMU string anywhere on its card is an MMMU-Pro figure for the base checkpoint. The abandonment extends to the independent side, and that is the strongest evidence that this column has run its course: Artificial Analysis publishes an MMMU-Pro evaluation and no plain-MMMU one at all, and Epoch AI's benchmark corpus carries no MMMU file either. The benchmark's own authors are the third piece of the same picture, though not in the way the front page suggests. The leaderboard shows a ‘Last updated: 09/05/2025’ stamp and renders its table client-side, but the data file that table is drawn from is live: 211 rows, with entries as recent as July 2026. Every row added since November 2025 carries an MMMU-Pro score and leaves the plain-MMMU cell empty, and the last model to receive a plain-MMMU number there was Claude Opus 4.5. So the authors did not stop maintaining the board — they moved it to the harder variant, which is the same move the labs made. Five of the eight flagships are now charted on Artificial Analysis's MMMU-Pro board — GPT-6 Astra 86.9, Gemini 3.8 Flash 85.6, Qwen3.8-Max 82.8, DeepSeek-V4.1-Flash 77.0 and Mistral Medium 3.5 64.9 — but that is a different and harder instrument, so importing one would break the column's internal comparability. The board is not topped by any of them: Anthropic's Claude Opus 5.5, a tier head this page's roster does not track, leads it at 87.7. Muse Spark 1.3 is the one flagship Artificial Analysis has scored on MMMU-Pro without charting it there: the run it charts by default for the model is the max-effort one, which carries no MMMU-Pro score, and the 82.0 belongs to a non-default xhigh run. With no published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and the strongest removal candidate on the page.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
SWE-Bench Verified
CodingApproaching saturationThe 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025, and now down to a single one of the eight current flagships: Mistral Medium 3.5 at 77.6, and nothing else. Two labs left inside a fortnight. Anthropic's departure was the louder one, because it held the highest figure on the page at 95 and its replacement flagship's system card reports SWE-bench Pro, Multilingual and Multimodal with no Verified row of any kind; DeepSeek's was quieter and cost the column more, because the retired DeepSeek-V4-Pro carried both the higher of the two surviving values at 80.6 and the column's only independent measurement. OpenAI, Google, xAI, Meta and Alibaba had already gone. Alibaba's open-weights model card is the neatest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. One value defines no band at all, so the label below is the one this column carried when it last had enough data to compute one, and it should be read as a marker rather than as a measurement. The independent half is now empty too. Epoch AI runs the benchmark itself and publishes per-run results in its bulk data download, but its 35-row board's only DeepSeek entry is the retired April V4-Pro build at 77.6, and its newest entries from OpenAI, Anthropic, Google and Alibaba are all at least one generation behind those labs' current flagships. That pattern used to be the reason to read the labs' abandonment as substitution rather than neglect — the benchmark was still being run, just not led with. That reading is now harder to sustain: Epoch's board has added nothing since late June 2026, so the independent side is not quietly keeping the benchmark alive either. Aider's polyglot leaderboard, the other board that historically covered this benchmark, has added nothing since two DeepSeek V3.2-Exp runs in October 2025 and carries none of the eight. A column with one lab value, no independent value and no computable band is the strongest removal candidate on this page after MMMU.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
About this page
Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.
Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.
Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.
Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.
Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.
Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.
Last verified: September 29, 2026. 8 frontier models · 7 benchmarks · 8 labs.