AI benchmarks
Every frontier AI model on every major benchmark
Frontier AI benchmark scores as of August 21, 2026: on ARC-AGI-2, GPT-5.6 Sol leads at 92.5%; on GPQA, Grok 4.6 leads at 94.9%; on SWE Pro, Claude Fable 5 leads at 80%. Every score below cites the lab’s announcement post or an independent re-runner.
Last verified: August 21, 2026.
How to read this page
● 87.2 — lab-claimed score. Sourced from the lab’s own announcement post or model card. Click the number for the citation. Closed-weights labs report what they choose to report; treat as the upper bound.
□ 86.4 — independent re-run. Sourced from Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard. Click the number for the evaluator’s page.
⚠ diverge — lab and independent scores differ by more than 5 percentage points. Often signals a methodology gap (extended thinking enabled vs. not, tools on vs. off, different subset, leaked-test contamination).
— — the lab didn’t publish this score and no independent re-run has landed yet. Honest gap, not zero.
Headline matrix
Each row is a current frontier flagship from one lab; each column is a major benchmark, ordered from most-discriminating to most-saturated. Click any score for its primary-source citation; click any column header to jump to that benchmark’s section below.
By benchmark
Ordered by how much each benchmark currently discriminates between frontier labs. Discriminating benchmarks separate models by capability; approaching-saturation benchmarks separate by points within a tight band.
ARC-AGI-2
ReasoningDiscriminatingSuccessor to the original ARC-AGI prize. Visual-pattern abstraction puzzles designed to resist memorization. No frontier lab publishes an ARC-AGI-2 number in its own launch materials any more — OpenAI and Anthropic have both moved their abstract-reasoning slot to ARC-AGI-3 — so every value here comes from the ARC Prize Foundation's own verified leaderboard, recorded at each model's highest verified reasoning-effort variant. Coverage rose from three flagships to four this cycle: the payload was regenerated on August 20 and picked up both of August's new flagships at once. GPT-5.6 Sol still leads at 92.5, Claude Fable 5 lands at 89.2, Gemini 3.7 Flash enters at 84.6 and Grok 4.6 at 67.1 — a 25-point band, tighter than last cycle's 32 but still discriminating. Both newcomers are large generational gains rather than new entries on a blank column: Gemini 3.7 Flash's 84.6 is 24.2 points above the 3.6 Flash it replaced, on a post-training step with no new pre-train and at 41% of the cost per task, and Grok 4.6's 67.1 is 14.5 above the retired Grok 4.5. The remaining four flagships — Muse Spark 1.2, DeepSeek-V4-Pro, Mistral Medium 3.5 and Qwen3.8-Max — have no verified submission at all; the only DeepSeek model on the board is V4-Flash.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
Humanity's Last Exam
KnowledgeDiscriminatingCrowdsourced expert-level exam across 100+ subjects, designed to be the last academic benchmark needed before frontier models match expert humans. Reported with and without tools. One of two benchmarks on this page with a published value for all eight current flagships — GPQA Diamond is the other — and by far the wider band of the two: Claude Fable 5 tops the field at 55.5 and Mistral Medium 3.5 anchors it at 13.8, a 42-point spread with the middle six clustered between 41.0 and 49.5. It is also the most volatile column here, and three refresh cycles running have proved it. Artificial Analysis re-ran most of the board in the second week of August 2026 and revised five independent values upward at once (Fable 5 53.3 to 55.5, GPT-5.6 Sol 47.2 to 49.5, Gemini 38.3 to 40.8, DeepSeek 35.9 to 39.3, Mistral 12.8 to 13.8); DeepSeek separately restated its own claim from 37.7 to 42.7 when V4-Pro left preview; and the third week moved DeepSeek's independent value again, 39.3 to 41.0. Google's entry changed for a different reason — Gemini 3.7 Flash replaced 3.6 Flash on August 13 and enters at 47.9, seven points above its predecessor. A value carried forward from a prior cycle on this column is almost certainly wrong. One caveat on Google specifically: the only HLE-shaped number Google publishes is HLE-Verified (3.7 Flash 53.6), a filtered subset rather than the full exam, so it is recorded in the cell note rather than in the cell.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
SWE-Bench Pro
CodingDiscriminatingContamination-resistant, multi-language successor to SWE-bench Verified. Real GitHub issues from production codebases that the model must patch end-to-end. Still the most widely reported column on this page, but its coverage is now visibly eroding — four of the eight current flagships carry a value, down from five, and all four are lab claims run on the lab's own harness rather than Scale's standardized leaderboard. Claude Fable 5's 80 sits over a 55–68 cluster (Qwen3.8-Max 67.7, GPT-5.6 Sol 64.6, DeepSeek-V4-Pro 55.4) — a ~25pp band, so what remains still discriminates at the production frontier. Three consecutive August flagship swaps each took a cell off this column or replaced one: Alibaba added Qwen3.8-Max 67.7 in the open-weights card it posted on August 12; xAI's Grok 4.6 (August 12) reports no SWE-Bench number where the retired Grok 4.5 claimed 64.7; and Google's Gemini 3.7 Flash (August 13) does the same, dropping the 58.7 the 3.6 Flash card had published only three weeks earlier. Google's software-engineering slot moved to DeepSWE v1.1 and Terminal-bench instead. Meta and Mistral publish nothing here, and Scale AI's public leaderboard still has no row for any of the eight.
Saturation: Frontier scores still span a wide band — this benchmark separates the labs.
AIME 2025
MathApproaching saturationAmerican Invitational Mathematics Examination, 2025 edition. 15 integer-answer problems; the standard frontier math benchmark through 2025. No current frontier flagship reports it any more, and none has for two refresh cycles running — OpenAI publishes FrontierMath, DeepSeek publishes HMMT 2026 Feb, and Anthropic, Google, xAI, Meta, Mistral and Alibaba publish no competition-math benchmark at all. Anthropic's Fable 5 system card says why in passing: AIME ‘was a popular AI benchmark last year but is now saturated’. With zero published values the band is undefined, so the label below is the one this column carried when it last had data. The column is kept for continuity and is a candidate for removal.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
GPQA Diamond
KnowledgeApproaching saturationGraduate-level physics, chemistry, and biology multiple-choice questions written by domain experts and validated to be “google-proof”. The Diamond subset (~198 questions) is the hardest tier, and all eight current flagships carry a value. The frontier has effectively cleared it: seven of the eight sit inside a 4.8-point band from DeepSeek-V4-Pro's 90.1 to Grok 4.6's 94.9, and Anthropic's Claude Fable 5 / Mythos 5 system card states outright that it considers GPQA Diamond a saturated evaluation and plans to stop reporting it on future models. The full eight-model band technically reads 20.1 points wide, but only because Mistral Medium 3.5 sits alone at 74.8 — one straggler, not a re-opened frontier — so the saturation label is held at approaching-saturation rather than flipped on a hairline crossing of the 20-point threshold.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
MMMU
MultimodalApproaching saturationMassive Multi-discipline Multimodal Understanding & Reasoning — college-exam-level questions across 30 subjects mixing text with diagrams, charts, and images. The canonical multimodal benchmark, and now an abandoned one at the frontier: OpenAI reports MMMU-Pro, Google reports CharXiv Reasoning, Anthropic reports GDP.pdf vision, and no current flagship publishes a plain MMMU number. Alibaba's Qwen3.8-Max is natively vision-language and its August 2026 model card still has no MMMU row. With zero published values the band is undefined, so the label below is the one this column carried when it last had data. Kept for continuity and a candidate for removal.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
SWE-Bench Verified
CodingApproaching saturationThe 500-issue human-verified subset of SWE-bench. The canonical ‘can the model do real software work’ benchmark from 2024–2025. Three of the eight current flagships still report it — Claude Fable 5 at 95, then DeepSeek-V4-Pro 80.6 and Mistral Medium 3.5 77.6 — while OpenAI, Google, xAI, Meta and Alibaba have all dropped it from their launch materials in favor of the harder Pro variant. Alibaba's August 2026 open-weights model card is the clearest illustration: it publishes 31 benchmark rows for Qwen3.8-Max, including SWE-bench Pro, and no SWE-bench Verified row at all. With three values the band reads 17.4 points, but three cells out of eight is too thin a sample to move a saturation label on, so the prior cycle's label stands.
Saturation: Frontier scores cluster near the top — this benchmark separates labs by points, not by capability.
About this page
Cross-family comparison page in the /ai/ section. The roster is the current frontier flagship from every major lab on this site — Claude, GPT, Gemini, Grok, Llama / Muse, DeepSeek, Mistral, Qwen — matched to /ai/models/ so each row links back to the per-family version page for the full lineage.
Lab-claimed vs. independent. Each cell can carry two values. The lab-claimed score (filled circle) is what the lab published in its announcement post, system card, or model card — the lab chooses the configuration (extended thinking, tool use, eval subset). The independent re-run (open square) is what Epoch AI, Artificial Analysis, Aider, or the benchmark’s own public leaderboard reports under their own protocol. When the two diverge by more than 5 percentage points, the page flags it — the gap is the editorial signal that matters here. Where no independent re-run exists yet, the cell shows the lab number alone; closed-weights labs are harder to re-run, so independent coverage is concentrated on the open-weights side.
Benchmark selection. Seven benchmarks covering reasoning, knowledge, coding, math, and multimodal capability. Picked for citation volume (every frontier launch reports these), discrimination (the score band is wide enough to separate labs), and primary-source availability (the benchmark author publishes a leaderboard or the eval protocol is public). Vendor-only proprietary benchmarks that no other lab reports are excluded. LMArena’s Elo is widely cited but is a different measurement type (human preference voting, not standardized eval); the page omits it but the LMArena leaderboard covers that signal.
Saturation framing. A benchmark is treated as discriminating when frontier scores span at least 20 percentage points, approaching saturation when the band tightens below that, and saturated when all frontier models cluster within a few points of the ceiling. Saturated benchmarks (HumanEval, MMLU, HellaSwag, GSM8K) are intentionally omitted from the v1 matrix; they no longer separate labs by capability. The saturation labels are re-evaluated on every refresh.
Sources. Primary lab announcements: Anthropic at anthropic.com/news, OpenAI at openai.com/index, Google at blog.google/technology/google-deepmind, xAI at x.ai/news, Meta at ai.meta.com/blog, DeepSeek at api-docs.deepseek.com/news, Mistral at mistral.ai/news, Alibaba at qwenlm.github.io/blog. Independent re-runners: Epoch AI, Artificial Analysis, Aider polyglot, ARC Prize Foundation.
Refreshed on every major model launch and at least monthly between launches. The page’s job is to stay current within a release cycle; the worst failure mode is showing a stale lab number after that lab has shipped a newer flagship.
Last verified: August 21, 2026. 8 frontier models · 7 benchmarks · 8 labs.