AI context windows

Every model’s context window, in one place

The largest context window you can call today is 10,000,000 tokens — Meta Llama 4 Scout, as of September 15, 2026. Across 149 models from 8 providers; the historical chart below traces the running maximum from 2018 onward.

As of September 15, 2026.

Largest in production
10,000,000
Llama 4 Scout
Median in production
1,000,000
tokens, across 66 models in production
2020 baseline
2,048
tokens (GPT-3)
Growth since 2020
~4,883×
on the maximum

Maximum context window over time

One marker per public model release, color-coded by provider. The dashed line traces the running maximum across all providers — the “largest context window in production” at each point in time. Y-axis is log-scale so a 5,000× range still reads.

1K8K32K128K1M10M201820192020202120222023202420252026OpenAI · GPT-1 · 512 tokens · Jun 11, 2018OpenAI · GPT-2 · 1,024 tokens · Feb 14, 2019OpenAI · GPT-3 · 2,048 tokens · May 28, 2020OpenAI · InstructGPT (text-davinci-002) · 4,097 tokens · Jan 27, 2022OpenAI · GPT-3.5 / ChatGPT launch · 4,096 tokens · Nov 30, 2022Meta · LLaMA 1 · 2,048 tokens · Mar 3, 2023OpenAI · GPT-4 · 8,192 tokens · Mar 14, 2023Anthropic · Claude 1 · 9,000 tokens · Mar 14, 2023Google · Bard (LaMDA-based) · 8,192 tokens · Mar 21, 2023Google · PaLM 2 · 8,192 tokens · May 10, 2023OpenAI · GPT-3.5 Turbo 16K · 16,385 tokens · Jun 13, 2023OpenAI · GPT-4-32K · 32,768 tokens · Jun 13, 2023Anthropic · Claude 2 · 100,000 tokens · Jul 11, 2023Meta · Llama 2 · 4,096 tokens · Jul 18, 2023Alibaba · Qwen-7B · 8,192 tokens · Aug 3, 2023Anthropic · Claude Instant 1.2 · 100,000 tokens · Aug 9, 2023Meta · Code Llama · 16,384 tokens · Aug 24, 2023Mistral · Mistral 7B · 8,192 tokens · Sep 27, 2023DeepSeek · DeepSeek-Coder · 16,384 tokens · Nov 2, 2023xAI · Grok 1 · 8,192 tokens · Nov 4, 2023OpenAI · GPT-4 Turbo · 128,000 tokens · Nov 6, 2023Anthropic · Claude 2.1 · 200,000 tokens · Nov 21, 2023DeepSeek · DeepSeek-LLM 7B / 67B · 4,096 tokens · Nov 29, 2023Google · Gemini 1.0 Pro · 32,768 tokens · Dec 6, 2023Mistral · Mixtral 8x7B · 32,768 tokens · Dec 11, 2023Alibaba · Qwen 1.5 family · 32,768 tokens · Feb 4, 2024Google · Gemini 1.0 Ultra · 32,768 tokens · Feb 8, 2024Google · Gemini 1.5 Pro · 1,000,000 tokens · Feb 15, 2024Mistral · Mistral Large · 32,768 tokens · Feb 26, 2024Anthropic · Claude 3 Opus · 200,000 tokens · Mar 4, 2024Anthropic · Claude 3 Sonnet · 200,000 tokens · Mar 4, 2024Anthropic · Claude 3 Haiku · 200,000 tokens · Mar 13, 2024xAI · Grok 1.5 · 131,072 tokens · Mar 28, 2024Mistral · Mixtral 8x22B · 65,536 tokens · Apr 17, 2024Meta · Llama 3 (8B / 70B) · 8,192 tokens · Apr 18, 2024DeepSeek · DeepSeek-V2 · 128,000 tokens · May 6, 2024OpenAI · GPT-4o · 128,000 tokens · May 13, 2024Google · Gemini 1.5 Flash · 1,000,000 tokens · May 14, 2024Alibaba · Qwen2 family · 131,072 tokens · Jun 7, 2024Anthropic · Claude 3.5 Sonnet · 200,000 tokens · Jun 20, 2024OpenAI · GPT-4o mini · 128,000 tokens · Jul 18, 2024Meta · Llama 3.1 (8B / 70B / 405B) · 131,072 tokens · Jul 23, 2024Mistral · Mistral Large 2 · 128,000 tokens · Jul 24, 2024xAI · Grok 2 · 131,072 tokens · Aug 13, 2024OpenAI · o1-preview · 128,000 tokens · Sep 12, 2024OpenAI · o1-mini · 128,000 tokens · Sep 12, 2024Alibaba · Qwen2.5 family · 131,072 tokens · Sep 19, 2024Google · Gemini 1.5 (002 refresh, 2M) · 2,000,000 tokens · Sep 24, 2024Meta · Llama 3.2 (Vision + Edge) · 131,072 tokens · Sep 25, 2024Google · Gemini 1.5 Flash-8B · 1,000,000 tokens · Oct 3, 2024Anthropic · Claude 3.5 Sonnet (new) · 200,000 tokens · Oct 22, 2024Anthropic · Claude 3.5 Haiku · 200,000 tokens · Nov 4, 2024OpenAI · o1 · 200,000 tokens · Dec 5, 2024Meta · Llama 3.3 70B Instruct · 131,072 tokens · Dec 6, 2024Google · Gemini 2.0 Flash · 1,048,576 tokens · Dec 11, 2024DeepSeek · DeepSeek-V3 · 128,000 tokens · Dec 26, 2024DeepSeek · DeepSeek-R1 · 128,000 tokens · Jan 20, 2025Mistral · Mistral Small 3 · 32,768 tokens · Jan 30, 2025OpenAI · o3-mini · 200,000 tokens · Jan 31, 2025Google · Gemini 2.0 Pro Experimental · 2,000,000 tokens · Feb 5, 2025xAI · Grok 3 · 131,072 tokens · Feb 17, 2025Anthropic · Claude 3.7 Sonnet · 200,000 tokens · Feb 24, 2025OpenAI · GPT-4.5 (Orion) · 128,000 tokens · Feb 27, 2025Google · Gemini 2.5 Pro · 1,048,576 tokens · Mar 25, 2025Meta · Llama 4 Scout · 10,000,000 tokens · Apr 5, 2025Meta · Llama 4 Maverick · 1,048,576 tokens · Apr 5, 2025OpenAI · GPT-4.1 · 1,047,576 tokens · Apr 14, 2025OpenAI · o3 · 200,000 tokens · Apr 16, 2025OpenAI · o4-mini · 200,000 tokens · Apr 16, 2025Alibaba · Qwen3 family · 32,768 tokens · Apr 28, 2025Anthropic · Claude Opus 4 · 200,000 tokens · May 22, 2025Anthropic · Claude Sonnet 4 · 200,000 tokens · May 22, 2025OpenAI · o3-pro · 200,000 tokens · Jun 10, 2025Google · Gemini 2.5 Flash · 1,048,576 tokens · Jun 17, 2025xAI · Grok 4 · 256,000 tokens · Jul 9, 2025Google · Gemini 2.5 Flash-Lite · 1,048,576 tokens · Jul 22, 2025Alibaba · Qwen3-Coder · 262,144 tokens · Jul 22, 2025Anthropic · Claude Opus 4.1 · 200,000 tokens · Aug 5, 2025OpenAI · GPT-5 · 400,000 tokens · Aug 7, 2025DeepSeek · DeepSeek-V3.1 · 128,000 tokens · Aug 21, 2025Alibaba · Qwen3-Next-80B-A3B · 262,144 tokens · Sep 11, 2025Alibaba · Qwen3-Max · 262,144 tokens · Sep 15, 2025xAI · Grok 4 Fast · 2,000,000 tokens · Sep 19, 2025Alibaba · Qwen3-VL family · 262,144 tokens · Sep 23, 2025Anthropic · Claude Sonnet 4.5 · 200,000 tokens · Sep 29, 2025Anthropic · Claude Haiku 4.5 · 200,000 tokens · Oct 15, 2025OpenAI · GPT-5.1 · 400,000 tokens · Nov 12, 2025xAI · Grok 4.1 · 256,000 tokens · Nov 17, 2025Google · Gemini 3 Pro · 1,048,576 tokens · Nov 18, 2025xAI · Grok 4.1 Fast · 2,000,000 tokens · Nov 19, 2025Anthropic · Claude Opus 4.5 · 200,000 tokens · Nov 24, 2025DeepSeek · DeepSeek-V3.2 (+ Speciale) · 128,000 tokens · Dec 1, 2025Mistral · Mistral Large 3 · 256,000 tokens · Dec 2, 2025OpenAI · GPT-5.2 · 400,000 tokens · Dec 11, 2025Google · Gemini 3 Flash · 1,048,576 tokens · Dec 17, 2025Alibaba · Qwen3-Coder-Next · 262,144 tokens · Feb 4, 2026OpenAI · GPT-5.3-Codex · 400,000 tokens · Feb 5, 2026Anthropic · Claude Opus 4.6 · 1,000,000 tokens · Feb 5, 2026Alibaba · Qwen3.5 family + Plus · 262,144 tokens · Feb 16, 2026Anthropic · Claude Sonnet 4.6 · 1,000,000 tokens · Feb 17, 2026Google · Gemini 3.1 Pro · 1,048,576 tokens · Feb 19, 2026OpenAI · GPT-5.3 Instant · 128,000 tokens · Mar 3, 2026Google · Gemini 3.1 Flash-Lite · 1,048,576 tokens · Mar 3, 2026OpenAI · GPT-5.4 · 1,050,000 tokens · Mar 5, 2026xAI · Grok 4.20 · 1,000,000 tokens · Mar 10, 2026Mistral · Mistral Small 4 · 256,000 tokens · Mar 16, 2026Alibaba · Qwen 3.6-Max-Preview · 262,144 tokens · Apr 2, 2026Meta · Muse Spark · 262,144 tokens · Apr 8, 2026Anthropic · Claude Opus 4.7 · 1,000,000 tokens · Apr 16, 2026Alibaba · Qwen3.6-35B-A3B · 262,144 tokens · Apr 16, 2026xAI · Grok 4.3 · 1,000,000 tokens · Apr 17, 2026Alibaba · Qwen3.6-27B · 262,144 tokens · Apr 22, 2026OpenAI · GPT-5.5 · 1,050,000 tokens · Apr 23, 2026DeepSeek · DeepSeek-V4-Pro · 1,000,000 tokens · Apr 24, 2026DeepSeek · DeepSeek-V4-Flash · 1,000,000 tokens · Apr 24, 2026Mistral · Mistral Medium 3.5 · 256,000 tokens · Apr 28, 2026Google · Gemini 3.5 Flash · 1,048,576 tokens · May 19, 2026xAI · Grok Build 0.1 · 256,000 tokens · May 19, 2026Alibaba · Qwen3.7-Max · 1,000,000 tokens · May 20, 2026Anthropic · Claude Opus 4.8 · 1,000,000 tokens · May 28, 2026Alibaba · Qwen3.7-Plus · 1,000,000 tokens · May 31, 2026Anthropic · Claude Fable 5 · 1,000,000 tokens · Jun 9, 2026Anthropic · Claude Mythos 5 · 1,000,000 tokens · Jun 9, 2026Anthropic · Claude Sonnet 5 · 1,000,000 tokens · Jun 30, 2026xAI · Grok 4.5 · 500,000 tokens · Jul 8, 2026OpenAI · GPT-5.6 Sol · 1,050,000 tokens · Jul 9, 2026OpenAI · GPT-5.6 Terra · 1,050,000 tokens · Jul 9, 2026OpenAI · GPT-5.6 Luna · 1,050,000 tokens · Jul 9, 2026Meta · Muse Spark 1.1 · 1,048,576 tokens · Jul 9, 2026Google · Gemini 3.5 Flash-Lite · 1,048,576 tokens · Jul 21, 2026Google · Gemini 3.6 Flash · 1,048,576 tokens · Jul 21, 2026Anthropic · Claude Opus 5 · 1,000,000 tokens · Jul 24, 2026Alibaba · Qwen3.8-Max · 1,000,000 tokens · Aug 3, 2026Meta · Muse Spark 1.2 · 1,048,576 tokens · Aug 5, 2026OpenAI · GPT-5.6 Cyber · 400,000 tokens · Aug 10, 2026Meta · Muse Glimmer 30B · 131,072 tokens · Aug 10, 2026xAI · Grok 4.6 · 500,000 tokens · Aug 12, 2026Alibaba · Qwen3.8-2.4T-A95B · 262,144 tokens · Aug 12, 2026Google · Gemini 3.7 Flash · 1,048,576 tokens · Aug 13, 2026Alibaba · Qwen3.8-27B · 262,144 tokens · Aug 14, 2026DeepSeek · DeepSeek-V4-Flash-Vision-Exp · 1,000,000 tokens · Aug 21, 2026Alibaba · Qwen3.8-Flash-Next · 262,144 tokens · Aug 26, 2026Anthropic · Claude Fable 5.1 · 1,000,000 tokens · Sep 1, 2026Anthropic · Claude Mythos 5.1 · 1,000,000 tokens · Sep 1, 2026Google · Gemini 3.8 Flash · 1,048,576 tokens · Sep 2, 2026Meta · Muse Spark 1.3 · 1,048,576 tokens · Sep 2, 2026OpenAI · GPT-6 Astra · 1,050,000 tokens · Sep 3, 2026DeepSeek · DeepSeek-V4.1-Flash · 1,000,000 tokens · Sep 10, 2026xAI · Grok 4.7 · 500,000 tokens · Sep 21, 2026tokens (log)
AnthropicOpenAIGooglexAIMetaDeepSeekMistralAlibaba

Every model

Sort by any column. Filter by provider or by minimum context window.

Min context:
Llama 4 Scout
Still the largest published window on this page. Meta publishes it as 10M; the released config sets max_position_embeddings to 10,485,760.
Meta(Llama)
10M
Apr 5, 2025
Available
Gemini 1.5 (002 refresh, 2M)
First mainstream 2M context.
Google(Gemini)
2M
8.2K
Sep 24, 2024
Legacy
Gemini 2.0 Pro Experimental
Experimental SKU only; never reached a stable gemini-2.0-pro id. Superseded by Gemini 2.5 Pro seven weeks later.
Google(Gemini)
2M
8.2K
Feb 5, 2025
Legacy
Grok 4 Fast
First Grok 2M context. Retired May 15, 2026.
xAI(Grok)
2M
Sep 19, 2025
Deprecated
Grok 4.1 Fast
Retired May 15, 2026.
xAI(Grok)
2M
Nov 19, 2025
Deprecated
GPT-5.4
1M-token context across the line.
OpenAI(ChatGPT)
1.1M
128K
Mar 5, 2026
Available
GPT-5.5
1.05M-token context; superseded as flagship by GPT-5.6 on Jul 9, 2026 and still served.
OpenAI(ChatGPT)
1.1M
128K
Apr 23, 2026
Available
GPT-5.6 Sol
Flagship tier of the generation+capability-tier naming; 1M context, 922K max input, 128K output. GA Jul 9, 2026; superseded as flagship by GPT-6 Astra on Sep 3, 2026 and still served.
OpenAI(ChatGPT)
1.1M
128K
Jul 9, 2026
Available
GPT-5.6 Terra
Balanced tier; 1M context, 922K max input, 128K output.
OpenAI(ChatGPT)
1.1M
128K
Jul 9, 2026
Current
GPT-5.6 Luna
Fastest / lowest-cost tier; 1M context, 922K max input, 128K output.
OpenAI(ChatGPT)
1.1M
128K
Jul 9, 2026
Current
GPT-6 Astra
First of the GPT-6 generation; same 1.05M window as the GPT-5.6 tiers, with a 922K max input and 128K max output, and an Apr 30, 2026 knowledge cutoff. OpenAI's model catalog names it the model to start with. Prompts above 272K input tokens bill at 2x input and cache rates and 1.5x output across the whole request.
OpenAI(ChatGPT)
1.1M
128K
Sep 3, 2026
Current
Gemini 2.0 Flash
Shut down Jun 1, 2026; Google recommends Gemini 3.6 Flash.
Google(Gemini)
1M
8.2K
Dec 11, 2024
Deprecated
Gemini 2.5 Pro
Stable 1M; the 2M window the Gemini 1.5 line reached is not documented for this model.
Google(Gemini)
1M
65.5K
Mar 25, 2025
Available
Llama 4 Maverick
Card rounds to "1M"; the released config sets the exact limit at 1,048,576.
Meta(Llama)
1M
Apr 5, 2025
Current
Gemini 2.5 Flash
Google(Gemini)
1M
65.5K
Jun 17, 2025
Available
Gemini 2.5 Flash-Lite
Google(Gemini)
1M
65.5K
Jul 22, 2025
Available
Gemini 3 Pro
Shut down Mar 9, 2026; migrate to Gemini 3.1 Pro.
Google(Gemini)
1M
65.5K
Nov 18, 2025
Deprecated
Gemini 3 Flash
Google(Gemini)
1M
65.5K
Dec 17, 2025
Available
Gemini 3.1 Pro
Replaced Gemini 3 Pro after Mar 9, 2026 shutdown.
Google(Gemini)
1M
65.5K
Feb 19, 2026
Current
Gemini 3.1 Flash-Lite
Stable GA May 7, 2026; earliest shutdown May 7, 2027, with Gemini 3.5 Flash-Lite the named replacement.
Google(Gemini)
1M
65.5K
Mar 3, 2026
Available
Gemini 3.5 Flash
Still Stable in the API; superseded as the workhorse Flash by Gemini 3.6 Flash on Jul 21, 2026.
Google(Gemini)
1M
65.5K
May 19, 2026
Available
Muse Spark 1.1
Multimodal reasoning / agentic model, superseded by Muse Spark 1.2 on Aug 5, 2026 but still served on the Meta Model API. Meta documents the window as 1,048,576 tokens rather than a rounded 1M. No separate max-output cap documented.
Meta(Llama)
1M
Jul 9, 2026
Available
Gemini 3.5 Flash-Lite
Stable GA Jul 21, 2026; the high-throughput tier replacing Gemini 3.1 Flash-Lite.
Google(Gemini)
1M
65.5K
Jul 21, 2026
Available
Gemini 3.6 Flash
Was the workhorse Flash for three weeks; superseded by Gemini 3.7 Flash on Aug 13, 2026. Same 1,048,576 / 65,536-token envelope.
Google(Gemini)
1M
65.5K
Jul 21, 2026
Available
Muse Spark 1.2
Coding-focused update to 1.1, co-trained with the Muse Code terminal agent. Window unchanged from 1.1 at 1,048,576 tokens; no separate max-output cap. Superseded by 1.3 on Sep 2, 2026, but still the model Meta's docs route audio input to.
Meta(Llama)
1M
Aug 5, 2026
Available
Gemini 3.7 Flash
Google's workhorse Flash from Aug 13 to Sep 2, 2026. Built on 3.6 Flash and inherits its envelope unchanged — 1,048,576-token input limit, 65,536-token output limit.
Google(Gemini)
1M
65.5K
Aug 13, 2026
Available
Gemini 3.8 Flash
Google's current workhorse Flash; Stable GA Sep 2, 2026. A post-training iteration on 3.7 Flash, so the envelope is unchanged for a third generation — 1,048,576-token input limit, 65,536-token output limit.
Google(Gemini)
1M
65.5K
Sep 2, 2026
Current
Muse Spark 1.3
Current Meta frontier model; window unchanged from 1.2 at 1,048,576 tokens, no separate max-output cap. Meta's docs flag audio understanding as not fully supported here and route audio work to Muse Spark 1.2; the max reasoning-effort level is offered on the Standard tier only.
Meta(Llama)
1M
Sep 2, 2026
Current
GPT-4.1
First OpenAI 1M-context production model; docs list the exact window as 1,047,576.
OpenAI(ChatGPT)
1M
32.8K
Apr 14, 2025
Legacy
Gemini 1.5 Pro
First mainstream 1M context.
Google(Gemini)
1M
8.2K
Feb 15, 2024
Legacy
Gemini 1.5 Flash
Google(Gemini)
1M
8.2K
May 14, 2024
Legacy
Gemini 1.5 Flash-8B
Google(Gemini)
1M
8.2K
Oct 3, 2024
Legacy
Claude Opus 4.6
1M GA Mar 13, 2026 (no long-context premium); 128K max output.
Anthropic(Claude)
1M
128K
Feb 5, 2026
Available
Claude Sonnet 4.6
1M context, 128K max output; superseded by Sonnet 5 as the default mid-tier and now listed under Anthropic's legacy models.
Anthropic(Claude)
1M
128K
Feb 17, 2026
Available
Grok 4.20
Three SKUs (multi-agent / reasoning / non-reasoning); 1M context across all SKUs per the current xAI pricing page (revised down from the 2M cited at the Mar 2026 launch).
xAI(Grok)
1M
Mar 10, 2026
Available
Claude Opus 4.7
1M default on Claude API / Bedrock / Vertex; 128K max output.
Anthropic(Claude)
1M
128K
Apr 16, 2026
Available
Grok 4.3
Prior xAI chat flagship; superseded by Grok 4.5 on Jul 8, 2026. Native video input; in-chat doc generation.
xAI(Grok)
1M
Apr 17, 2026
Available
DeepSeek-V4-Pro
First DeepSeek 1M; Thinking + Non-Thinking modes; 384K max output, no vision. Left its four-month public preview on Aug 13, 2026 as DeepSeek-V4-Pro-0813, with the window unchanged. DeepSeek's pricing page says API service continues past the September 14, 2026 end date it had signalled, billing unchanged.
DeepSeek(DeepSeek)
1M
384K
Apr 24, 2026
Available
DeepSeek-V4-Flash
1M context, 384K max output. Official API release (V4-Flash-0731) Jul 31, 2026 — re-post-trained, same architecture and limits. Retired Sep 10, 2026: the deepseek-v4-flash string is still accepted but is served by DeepSeek-V4.1-Flash, which carries the same window.
DeepSeek(DeepSeek)
1M
384K
Apr 24, 2026
Deprecated
Qwen3.7-Max
Closed-weights; native extended-thinking. The first Qwen Max-tier model at 1M — 991K max input, 131K max output. Superseded at the top of the Max tier by Qwen3.8-Max on Aug 3, 2026; still served.
Alibaba(Qwen)
1M
131.1K
May 20, 2026
Available
Claude Opus 4.8
1M context, 128K output — Anthropic now documents the 1M window on the Claude API, Bedrock, Google Cloud and Microsoft Foundry alike. Superseded by Opus 5 on Jul 24, 2026; now listed under Anthropic's legacy models.
Anthropic(Claude)
1M
128K
May 28, 2026
Available
Qwen3.7-Plus
Multimodal vision+language sibling to Qwen3.7-Max; 1M context, 991K max input, 131K max output — the 65K output cap belongs to Qwen3.5-Plus, not to this tier. DashScope only.
Alibaba(Qwen)
1M
131.1K
May 31, 2026
Current
Claude Fable 5
Mythos-class flagship until Sep 1, 2026; 1M context, 128K output. Export-control suspension (Jun 12) lifted Jun 30; redeployed globally Jul 1, 2026. Superseded by Claude Fable 5.1, which carries the same envelope.
Anthropic(Claude)
1M
128K
Jun 9, 2026
Available
Claude Mythos 5
Shares Fable 5 capabilities without cyber safeguards; invitation-only via Project Glasswing. 1M context, 128K output. Redeployed alongside Fable 5 Jul 1, 2026. Anthropic's model page now names Claude Mythos 5.1 as the current Mythos model.
Anthropic(Claude)
1M
128K
Jun 9, 2026
Available
Claude Sonnet 5
New default mid-tier workhorse; 1M context, 128K max output. Shipped Jun 30, 2026.
Anthropic(Claude)
1M
128K
Jun 30, 2026
Current
Claude Opus 5
Current Opus flagship; 1M context, 128K max output. Thinking on by default with a new max effort tier. 300K output available on the Batches API behind a beta header.
Anthropic(Claude)
1M
128K
Jul 24, 2026
Current
Qwen3.8-Max
2.4T-parameter sparse-MoE native vision-language flagship; generally available Aug 3, 2026 after a Jul 19 WAIC preview. Qwen Cloud documents the 1M window as 991K max input / 131K max output, with a separate 262K reasoning budget. The qwen3.8-max-0902 snapshot (Sep 2, 2026) keeps that envelope unchanged.
Alibaba(Qwen)
1M
131.1K
Aug 3, 2026
Current
DeepSeek-V4-Flash-Vision-Exp
Experimental multimodal V4-Flash build, API-only. Carried the same 1M context / 384K max output as the rest of the V4 line. Retired Sep 10, 2026, twenty days after it shipped; the model string is still accepted but is served by DeepSeek-V4.1-Flash.
DeepSeek(DeepSeek)
1M
384K
Aug 21, 2026
Deprecated
Claude Fable 5.1
Mythos-class flagship; 1M context, 128K output — envelope unchanged from Fable 5. GA on every platform on day one, with always-on adaptive thinking and cache reads cut 75%.
Anthropic(Claude)
1M
128K
Sep 1, 2026
Current
Claude Mythos 5.1
Claude Fable 5.1 with more permissive safeguards, offered by invitation only through Project Glasswing. Anthropic documents the same envelope as Fable 5.1 — 1M context, 128K max output — and the same $10 / $50 pricing.
Anthropic(Claude)
1M
128K
Sep 1, 2026
Current
DeepSeek-V4.1-Flash
Called as deepseek-flash; DeepSeek's pricing page gives the model version as DeepSeek-V4.1-Flash. 1M context, 384K max output, and the first DeepSeek model whose pricing table marks vision as supported. A 552B Causal Encoder-Decoder MoE holding the global KV cache at 890 bytes per token, about a quarter of V4-Flash's, and the concurrency limit rises from 500 to 2,500.
DeepSeek(DeepSeek)
1M
384K
Sep 10, 2026
Current
Grok 4.5
500K context; configurable reasoning effort (low/medium/high). Superseded as xAI's flagship by Grok 4.6 on Aug 12, 2026; still in the API price list.
xAI(Grok)
500K
Jul 8, 2026
Available
Grok 4.6
500K context, no documented text-output limit; added an xhigh reasoning-effort level above Grok 4.5's low/medium/high. Superseded as xAI's flagship by Grok 4.7 on Sep 21, 2026; still in the API price list and still the model xAI's consumer plans name.
xAI(Grok)
500K
Aug 12, 2026
Available
Grok 4.7
Current xAI flagship; 500K context and no documented text-output limit, both unchanged from Grok 4.6, as is the low/medium/high/xhigh reasoning-effort dial. Knowledge cutoff moves to May 2026. A Grok 4.7 Fast variant runs at twice the output speed for twice the price.
xAI(Grok)
500K
Sep 21, 2026
Current
GPT-5
Unified router across reasoning + chat; 400K context, 272K max input, 128K output. Deprecated Jun 11, 2026; API shutdown scheduled Dec 11, 2026.
OpenAI(ChatGPT)
400K
128K
Aug 7, 2025
Legacy
GPT-5.1
OpenAI(ChatGPT)
400K
128K
Nov 12, 2025
Legacy
GPT-5.2
Still in the API catalog; the gpt-5.2-chat-latest variant shut down Aug 10, 2026.
OpenAI(ChatGPT)
400K
128K
Dec 11, 2025
Available
GPT-5.3-Codex
Code-tuned variant; 400K context, 272K max input, 128K output.
OpenAI(ChatGPT)
400K
128K
Feb 5, 2026
Available
GPT-5.6 Cyber
Cybersecurity-specialized fork of Sol behind the Daybreak Red approval tier; 400K window (272K max input) rather than the 1.05M the Sol / Terra / Luna tiers carry.
OpenAI(ChatGPT)
400K
128K
Aug 10, 2026
Available
Qwen3-Coder
256K native, extrapolated to 1M; Apache 2.0.
Alibaba(Qwen)
262.1K
Jul 22, 2025
Available
Qwen3-Next-80B-A3B
Ultra-sparse MoE; 256K native, extensible to 1M.
Alibaba(Qwen)
262.1K
Sep 11, 2025
Available
Qwen3-Max
Trillion-parameter MoE; superseded by Qwen 3.6-Max-Preview. Alibaba documents a 262,144-token context (258K max input, 65K max output) across every qwen3-max snapshot; the 1M figure quoted for it elsewhere is the tokens-per-minute rate limit, not the window. The Max tier did not reach 1M until Qwen3.7-Max.
Alibaba(Qwen)
262.1K
65.5K
Sep 15, 2025
Legacy
Qwen3-VL family
Vision-language; 256K native interleaved context.
Alibaba(Qwen)
262.1K
Sep 23, 2025
Available
Qwen3-Coder-Next
262,144 native; the model card documents no YaRN extension beyond it. Coding-agent MoE.
Alibaba(Qwen)
262.1K
Feb 4, 2026
Available
Qwen3.5 family + Plus
262K native on the open weights, 1M extended via YaRN; the hosted qwen3.5-plus SKU is documented at a 1M window with a 65K max output.
Alibaba(Qwen)
262.1K
Feb 16, 2026
Available
Qwen 3.6-Max-Preview
Proprietary preview, DashScope only; superseded by Qwen3.7-Max. Qwen Cloud documents 262,144 context / 245K max input / 65K max output — the Max tier did not reach 1M until Qwen3.7-Max.
Alibaba(Qwen)
262.1K
65.5K
Apr 2, 2026
Legacy
Muse Spark
262K context; closed weights. Superseded by Muse Spark 1.1 (Jul 9, 2026) as Meta's frontier model, and no longer listed among the available Muse Spark models on Meta's developer docs, which carry 1.1, 1.2 and 1.3 only.
Meta(Llama)
262.1K
Apr 8, 2026
Legacy
Qwen3.6-35B-A3B
First Qwen3.6 open-weights release; 35B-total / 3B-active MoE.
Alibaba(Qwen)
262.1K
Apr 16, 2026
Current
Qwen3.6-27B
Hybrid Gated DeltaNet + self-attention; 1M extensible.
Alibaba(Qwen)
262.1K
Apr 22, 2026
Current
Qwen3.8-2.4T-A95B
Alibaba's first open Max-class weights. Text-only and thinking-only, and 262K native (1.01M with YaRN) against the hosted Qwen3.8-Max's 1M default — the same base model with a smaller shipped window.
Alibaba(Qwen)
262.1K
Aug 12, 2026
Current
Qwen3.8-27B
27B dense, Apache 2.0, native vision-language. 262,144 native, extensible to 1,000,000 with YaRN; the hosted qwen3.8-27b SKU defaults to the full 1M.
Alibaba(Qwen)
262.1K
Aug 14, 2026
Current
Qwen3.8-Flash-Next
125B-total / 6B-active preview of the Qwen4 architecture, shipped as open weights under the Qwen Community License 1.0. 262,144 native, extensible to 1,000,000; the hosted qwen3.8-flash SKU built on it is documented at 1M context / 991K max input / 131K max output.
Alibaba(Qwen)
262.1K
Aug 26, 2026
Current
Grok 4
Retired May 15, 2026.
xAI(Grok)
256K
Jul 9, 2025
Deprecated
Grok 4.1
Removed from the xAI docs: no longer on the pricing table, and its per-model page 404s. Not named in the May 15, 2026 retirement list, which covered the grok-4-1-fast SKUs rather than grok-4.1 itself.
xAI(Grok)
256K
Nov 17, 2025
Deprecated
Mistral Large 3
Granular MoE flagship, Apache 2.0.
Mistral(Mistral)
256K
Dec 2, 2025
Current
Mistral Small 4
Mistral(Mistral)
256K
Mar 16, 2026
Current
Mistral Medium 3.5
Largest dense open-weights; merged chat/reasoning/code.
Mistral(Mistral)
256K
Apr 28, 2026
Current
Grok Build 0.1
Purpose-built agentic coding model.
xAI(Grok)
256K
May 19, 2026
Current
Claude 2.1
First 200K window. Retired Jul 21, 2025.
Anthropic(Claude)
200K
Nov 21, 2023
Deprecated
Claude 3 Opus
Retired Jan 5, 2026.
Anthropic(Claude)
200K
4.1K
Mar 4, 2024
Deprecated
Claude 3 Sonnet
Retired Jul 21, 2025.
Anthropic(Claude)
200K
4.1K
Mar 4, 2024
Deprecated
Claude 3 Haiku
Retired Apr 20, 2026.
Anthropic(Claude)
200K
4.1K
Mar 13, 2024
Deprecated
Claude 3.5 Sonnet
Retired Oct 28, 2025.
Anthropic(Claude)
200K
8.2K
Jun 20, 2024
Deprecated
Claude 3.5 Sonnet (new)
Computer-use beta. Retired Oct 28, 2025.
Anthropic(Claude)
200K
8.2K
Oct 22, 2024
Deprecated
Claude 3.5 Haiku
Retired Feb 19, 2026.
Anthropic(Claude)
200K
8.2K
Nov 4, 2024
Deprecated
o1
First production o-series. Still in OpenAI's model catalog; API shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
200K
100K
Dec 5, 2024
Legacy
o3-mini
Still in OpenAI's model catalog; API shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
200K
100K
Jan 31, 2025
Legacy
Claude 3.7 Sonnet
Extended thinking. Retired Feb 19, 2026.
Anthropic(Claude)
200K
64K
Feb 24, 2025
Deprecated
o3
Deprecated Jun 11, 2026; API shutdown scheduled Dec 11, 2026.
OpenAI(ChatGPT)
200K
100K
Apr 16, 2025
Legacy
o4-mini
API snapshot shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
200K
100K
Apr 16, 2025
Legacy
Claude Opus 4
Retired Jun 15, 2026.
Anthropic(Claude)
200K
32K
May 22, 2025
Deprecated
Claude Sonnet 4
Retired Jun 15, 2026.
Anthropic(Claude)
200K
64K
May 22, 2025
Deprecated
o3-pro
Deprecated Jun 11, 2026; API shutdown scheduled Dec 11, 2026.
OpenAI(ChatGPT)
200K
100K
Jun 10, 2025
Legacy
Claude Opus 4.1
Deprecated Jun 5, 2026; retired Aug 5, 2026.
Anthropic(Claude)
200K
32K
Aug 5, 2025
Deprecated
Claude Sonnet 4.5
Superseded by Sonnet 4.6; still served via the API (retirement no sooner than Sep 29, 2026).
Anthropic(Claude)
200K
64K
Sep 29, 2025
Available
Claude Haiku 4.5
Still the fast tier in Anthropic's latest-models table; 200K context, 64K max output.
Anthropic(Claude)
200K
64K
Oct 15, 2025
Current
Claude Opus 4.5
Anthropic(Claude)
200K
64K
Nov 24, 2025
Available
Grok 1.5
xAI(Grok)
131.1K
Mar 28, 2024
Legacy
Qwen2 family
Apache 2.0 turn.
Alibaba(Qwen)
131.1K
Jun 7, 2024
Legacy
Llama 3.1 (8B / 70B / 405B)
First Llama 128K.
Meta(Llama)
131.1K
Jul 23, 2024
Available
Grok 2
xAI(Grok)
131.1K
Aug 13, 2024
Legacy
Qwen2.5 family
Alibaba(Qwen)
131.1K
Sep 19, 2024
Legacy
Llama 3.2 (Vision + Edge)
Meta(Llama)
131.1K
Sep 25, 2024
Available
Llama 3.3 70B Instruct
Meta(Llama)
131.1K
Dec 6, 2024
Available
Grok 3
131K context across grok-3 / grok-3-fast / grok-3-mini / grok-3-mini-fast. Retired May 15, 2026; requests redirect to grok-4.3.
xAI(Grok)
131.1K
Feb 17, 2025
Deprecated
Muse Glimmer 30B
First Meta open-weights release since Llama 4 and the first under Apache 2.0; published to the meta-models HuggingFace org, not meta-llama. A 30B local agent model, so 131K rather than the flagship 1M.
Meta(Llama)
131.1K
Aug 10, 2026
Current
GPT-4 Turbo
DevDay 2023; first OpenAI 128K. Still in OpenAI's model catalog; API shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
128K
4.1K
Nov 6, 2023
Legacy
DeepSeek-V2
DeepSeek(DeepSeek)
128K
May 6, 2024
Legacy
GPT-4o
Multimodal flagship; replaced as ChatGPT default by GPT-5.
OpenAI(ChatGPT)
128K
16.4K
May 13, 2024
Legacy
GPT-4o mini
Replaced GPT-3.5 Turbo on the mini tier.
OpenAI(ChatGPT)
128K
16.4K
Jul 18, 2024
Legacy
Mistral Large 2
Retired Mar 30, 2025.
Mistral(Mistral)
128K
Jul 24, 2024
Deprecated
o1-preview
Reasoning track preview.
OpenAI(ChatGPT)
128K
32.8K
Sep 12, 2024
Deprecated
o1-mini
Smaller, cheaper reasoning companion to o1-preview.
OpenAI(ChatGPT)
128K
65.5K
Sep 12, 2024
Deprecated
DeepSeek-V3
DeepSeek(DeepSeek)
128K
Dec 26, 2024
Legacy
DeepSeek-R1
Reasoning track flagship.
DeepSeek(DeepSeek)
128K
Jan 20, 2025
Legacy
GPT-4.5 (Orion)
Largest pretraining run; research preview. API access ended Jul 14, 2025.
OpenAI(ChatGPT)
128K
16.4K
Feb 27, 2025
Legacy
DeepSeek-V3.1
DeepSeek(DeepSeek)
128K
Aug 21, 2025
Legacy
DeepSeek-V3.2 (+ Speciale)
The deepseek-chat / deepseek-reasoner API names that served it were discontinued Jul 24, 2026; weights stay MIT-licensed.
DeepSeek(DeepSeek)
128K
Dec 1, 2025
Legacy
GPT-5.3 Instant
Served as gpt-5.3-chat-latest, which OpenAI documents at 128K context / 16K output — the chat-latest tier never carried the 400K window of the base GPT-5.x models. Shut down in the API Aug 10, 2026.
OpenAI(ChatGPT)
128K
16.4K
Mar 3, 2026
Deprecated
Claude 2
First mainstream 100K window. Retired Jul 21, 2025.
Anthropic(Claude)
100K
Jul 11, 2023
Deprecated
Claude Instant 1.2
Retired Nov 6, 2024.
Anthropic(Claude)
100K
Aug 9, 2023
Deprecated
Mixtral 8x22B
Sparse MoE, 39B active of 141B total, Apache 2.0; Mistral's model card documents a 64k window. Retired Mar 30, 2025.
Mistral(Mistral)
65.5K
Apr 17, 2024
Deprecated
GPT-4-32K
Extended GPT-4 variant.
OpenAI(ChatGPT)
32.8K
Jun 13, 2023
Legacy
Gemini 1.0 Pro
Google(Gemini)
32.8K
8.2K
Dec 6, 2023
Legacy
Mixtral 8x7B
Retired Mar 30, 2025.
Mistral(Mistral)
32.8K
Dec 11, 2023
Deprecated
Qwen 1.5 family
Stable 32K across every size; first Qwen MoE followed Mar 2024.
Alibaba(Qwen)
32.8K
Feb 4, 2024
Legacy
Gemini 1.0 Ultra
Gemini Advanced launch.
Google(Gemini)
32.8K
8.2K
Feb 8, 2024
Legacy
Mistral Large
Retired Jun 16, 2025.
Mistral(Mistral)
32.8K
Feb 26, 2024
Deprecated
Mistral Small 3
Superseded by Mistral Small 4; retired Nov 30, 2025.
Mistral(Mistral)
32.8K
Jan 30, 2025
Deprecated
Qwen3 family
Hybrid Thinking / Non-Thinking. 32,768 tokens natively, 131,072 with YaRN — the first Qwen card to state the native and extended figures separately.
Alibaba(Qwen)
32.8K
Apr 28, 2025
Legacy
GPT-3.5 Turbo 16K
First mainstream 16K model. API shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
16.4K
4.1K
Jun 13, 2023
Legacy
Code Llama
Code-specialized.
Meta(Llama)
16.4K
Aug 24, 2023
Legacy
DeepSeek-Coder
DeepSeek(DeepSeek)
16.4K
Nov 2, 2023
Legacy
Claude 1
Retired Nov 6, 2024.
Anthropic(Claude)
9K
Mar 14, 2023
Deprecated
GPT-4
32K variant later in 2023. API shutdown scheduled Oct 23, 2026.
OpenAI(ChatGPT)
8.2K
8.2K
Mar 14, 2023
Legacy
Bard (LaMDA-based)
Pre-Gemini chat product.
Google(Gemini)
8.2K
Mar 21, 2023
Legacy
PaLM 2
Google(Gemini)
8.2K
May 10, 2023
Legacy
Qwen-7B
Lab debut.
Alibaba(Qwen)
8.2K
Aug 3, 2023
Legacy
Mistral 7B
Sliding-window attention. Retired Mar 30, 2025.
Mistral(Mistral)
8.2K
Sep 27, 2023
Deprecated
Grok 1
xAI(Grok)
8.2K
Nov 4, 2023
Legacy
Llama 3 (8B / 70B)
Meta(Llama)
8.2K
Apr 18, 2024
Legacy
InstructGPT (text-davinci-002)
RLHF-tuned.
OpenAI(ChatGPT)
4.1K
Jan 27, 2022
Legacy
GPT-3.5 / ChatGPT launch
ChatGPT launch; later 16K variant.
OpenAI(ChatGPT)
4.1K
Nov 30, 2022
Legacy
Llama 2
Meta(Llama)
4.1K
Jul 18, 2023
Legacy
DeepSeek-LLM 7B / 67B
DeepSeek(DeepSeek)
4.1K
Nov 29, 2023
Legacy
GPT-3
The 2020 baseline.
OpenAI(ChatGPT)
2K
May 28, 2020
Legacy
LLaMA 1
Meta(Llama)
2K
Mar 3, 2023
Legacy
GPT-2
Doubled GPT-1.
OpenAI(ChatGPT)
1K
Feb 14, 2019
Legacy
GPT-1
Original transformer paper baseline.
OpenAI(ChatGPT)
512
Jun 11, 2018
Legacy

Showing all 149 models.

By provider

Each provider’s context-window history at a glance. “Current” is that lab’s current flagship as listed on the frontier-model roster; it can sit below the lifetime maximum when a provider has rolled an early extended-context experiment back into a smaller production window.

Anthropic
Claude family
Versions →
Current
1M
Claude Fable 5.1
Lifetime max
1M
First model
9K
Mar 14, 2023
Most recent
Sep 1, 2026Claude Fable 5.11M
Sep 1, 2026Claude Mythos 5.11M
Jul 24, 2026Claude Opus 51M
Jun 30, 2026Claude Sonnet 51M
Jun 9, 2026Claude Fable 51M
Jun 9, 2026Claude Mythos 51M
OpenAI
ChatGPT family
Versions →
Current
1.1M
GPT-6 Astra
Lifetime max
1.1M
First model
512
Jun 11, 2018
Most recent
Sep 3, 2026GPT-6 Astra1.1M
Aug 10, 2026GPT-5.6 Cyber400K
Jul 9, 2026GPT-5.6 Sol1.1M
Jul 9, 2026GPT-5.6 Terra1.1M
Jul 9, 2026GPT-5.6 Luna1.1M
Apr 23, 2026GPT-5.51.1M
Google
Gemini family
Versions →
Current
1M
Gemini 3.8 Flash
Lifetime max
2M
First model
8.2K
Mar 21, 2023
Most recent
Sep 2, 2026Gemini 3.8 Flash1M
Aug 13, 2026Gemini 3.7 Flash1M
Jul 21, 2026Gemini 3.5 Flash-Lite1M
Jul 21, 2026Gemini 3.6 Flash1M
May 19, 2026Gemini 3.5 Flash1M
xAI
Grok family
Versions →
Current
500K
Grok 4.7
Lifetime max
2M
First model
8.2K
Nov 4, 2023
Most recent
Sep 21, 2026Grok 4.7500K
Aug 12, 2026Grok 4.6500K
Jul 8, 2026Grok 4.5500K
May 19, 2026Grok Build 0.1256K
Apr 17, 2026Grok 4.31M
Mar 10, 2026Grok 4.201M
Meta
Llama family
Versions →
Current
1M
Muse Spark 1.3
Lifetime max
10M
First model
2K
Mar 3, 2023
Most recent
Sep 2, 2026Muse Spark 1.31M
Aug 10, 2026Muse Glimmer 30B131.1K
Aug 5, 2026Muse Spark 1.21M
Jul 9, 2026Muse Spark 1.11M
Apr 8, 2026Muse Spark262.1K
Apr 5, 2025Llama 4 Scout10M
DeepSeek
DeepSeek family
Versions →
Current
1M
DeepSeek-V4.1-Flash
Lifetime max
1M
First model
16.4K
Nov 2, 2023
Most recent
Sep 10, 2026DeepSeek-V4.1-Flash1M
Apr 24, 2026DeepSeek-V4-Pro1M
Apr 24, 2026DeepSeek-V4-Flash1M
Aug 21, 2025DeepSeek-V3.1128K
Mistral
Mistral family
Versions →
Current
256K
Mistral Medium 3.5
Lifetime max
256K
First model
8.2K
Sep 27, 2023
Most recent
Apr 28, 2026Mistral Medium 3.5256K
Mar 16, 2026Mistral Small 4256K
Dec 2, 2025Mistral Large 3256K
Jan 30, 2025Mistral Small 332.8K
Jul 24, 2024Mistral Large 2128K
Apr 17, 2024Mixtral 8x22B65.5K
Alibaba
Qwen family
Versions →
Current
1M
Qwen3.8-Max
Lifetime max
1M
First model
8.2K
Aug 3, 2023
Most recent
Aug 26, 2026Qwen3.8-Flash-Next262.1K
Aug 14, 2026Qwen3.8-27B262.1K
Aug 12, 2026Qwen3.8-2.4T-A95B262.1K
Aug 3, 2026Qwen3.8-Max1M
May 31, 2026Qwen3.7-Plus1M
May 20, 2026Qwen3.7-Max1M

Notes and caveats

Effective vs. nominal context. Several long-context models advertise large windows but degrade past a certain length on real-world tasks — this page records the documented nominal capacity, not benchmark-measured effective length. The latter is benchmark-dependent and out of scope per the section’s no-benchmarks rule.

Input vs. output limits. Some providers document a single “context window” that includes both prompt and response tokens; others publish a separate max-output cap alongside it. Which is which varies by lab, by generation, and by whether the model is hosted or open-weights — a lab that documents one number for its open checkpoints often documents two for the hosted SKU built on them. Where a separate cap is documented, the Output column shows it; where it isn’t, the column reads — and the row treats the input window as the whole budget.

Beta and tier-gated context. Some providers ship a default context size for the standard API and a larger one behind a beta flag, batch endpoint, or paid tier. The headline number on this page is the standard-API value documented as generally available; the per-row notes call out when a beta or tier-gated extended window exists.

Open-weights inference. For open-weights models (Llama, DeepSeek, Mistral, Qwen) the “context window” is the value the model card claims; serving infrastructure (vLLM, Together, Fireworks, Hugging Face Inference) often caps the deployed window lower for memory reasons. Always check the specific endpoint’s docs before relying on the full nominal window.

Tokenizer differences. One token is not a fixed unit across providers. OpenAI’s o200k tokenizer, Anthropic’s tokenizer, Google’s SentencePiece, and Meta’s tiktoken-derived tokenizers all produce different token counts for identical text. Compare context windows in tokens, not in characters or pages, but treat them as a same-provider apples-to-apples comparison rather than a strict cross-provider one.

About this page

Cross-family comparison page in the /ai/ section. Each row’s context-window value is sourced from the provider’s own model documentation — OpenAI’s platform.openai.com/docs/models, Anthropic’s platform.claude.com, Google’s ai.google.dev, xAI’s docs.x.ai, Meta’s llama.com, huggingface.co/meta-llama and huggingface.co/meta-models, DeepSeek’s api-docs.deepseek.com, Mistral’s docs.mistral.ai, and Alibaba’s help.aliyun.com/zh/model-studio.

The model roster mirrors the per-family pages already on this site — Claude, ChatGPT, Gemini, Grok, Llama, DeepSeek, Mistral, Qwen — so each row links back to the matching version-page entry for the full per-release context.

Refreshed daily. Each refresh re-verifies every row against the provider’s current documentation; values that changed since the previous run are updated and the row’s “as of” date is bumped. See release cadence for the cross-family ship-cadence picture this page complements.

Last updated: September 15, 2026. 149 models · 8 providers.