AI API pricing history

Every frontier AI model’s per-token price, current and historical

The cheapest frontier-tier input price you can call today is $0.18 per million tokens — Meta Llama 4 Scout (Together), as of September 1, 2026. That’s a 167× compression against GPT-4’s March 2023 launch price of $30 per million. Across 51 models in production from 8 providers, with 127 distinct price points charted from 2020 onward.

As of September 1, 2026. Input pricing only in the headline; output, cached, and batch tiers all live in the table below.

Cheapest frontier in production
$0.18 / $0.59
Llama 4 Scout (Together) (in / out per M)
Cheapest small tier in production
$0.10
Gemini 2.5 Flash-Lite input / M
2023 baseline (GPT-4)
$30.00
input / M at launch
Compression vs GPT-4
~167×
cheaper today

Per-token input price over time

One marker per published price point, color-coded by provider. The dashed line traces the running minimum across the frontier tier — the cheapest input-token price for a flagship-class model at each point in time. Y-axis is log-scale to fit the 6,000× range between $60/M davinci tokens and the sub-penny cache-hit floor.

$0.10$1$10$602020202120222023202420252026OpenAI · GPT-3 (text-davinci-003) · $60.00 / M input · Jun 11, 2020OpenAI · GPT-3.5 Turbo · $2.00 / M input · Mar 1, 2023OpenAI · GPT-4 · $30.00 / M input · Mar 14, 2023Anthropic · Claude 1 · $11.02 / M input · May 15, 2023Anthropic · Claude Instant · $1.63 / M input · May 15, 2023OpenAI · GPT-3.5 Turbo · $1.50 / M input · Jun 13, 2023OpenAI · GPT-4 32K · $60.00 / M input · Jun 13, 2023Anthropic · Claude 2 · $11.02 / M input · Jul 11, 2023Meta · Llama 2 70B (Together) · $0.90 / M input · Oct 1, 2023OpenAI · GPT-3.5 Turbo · $1.00 / M input · Nov 6, 2023OpenAI · GPT-4 Turbo · $10.00 / M input · Nov 6, 2023Anthropic · Claude Instant · $0.80 / M input · Nov 21, 2023Anthropic · Claude 2 · $8.00 / M input · Nov 21, 2023Mistral · Mistral Medium (original) · $2.50 / M input · Dec 11, 2023OpenAI · GPT-3.5 Turbo · $0.50 / M input · Jan 25, 2024Google · Gemini 1.0 Pro · $0.50 / M input · Feb 15, 2024Mistral · Mistral Large · $8.00 / M input · Feb 26, 2024Anthropic · Claude 3 Opus · $15.00 / M input · Mar 4, 2024Anthropic · Claude 3 Sonnet · $3.00 / M input · Mar 4, 2024Anthropic · Claude 3 Haiku · $0.25 / M input · Mar 13, 2024Meta · Llama 3 70B (Together) · $0.90 / M input · Apr 18, 2024DeepSeek · DeepSeek-V2 · $0.14 / M input · May 6, 2024OpenAI · GPT-4o · $5.00 / M input · May 13, 2024Google · Gemini 1.5 Pro · $7.00 / M input · May 14, 2024Google · Gemini 1.5 Flash · $0.075 / M input · May 14, 2024Alibaba · Qwen Plus · $0.26 / M input · Jun 6, 2024Alibaba · Qwen Max · $0.78 / M input · Jun 6, 2024Anthropic · Claude 3.5 Sonnet · $3.00 / M input · Jun 20, 2024OpenAI · GPT-4o mini · $0.15 / M input · Jul 18, 2024Meta · Llama 3.1 405B (Together) · $3.50 / M input · Jul 23, 2024Mistral · Mistral Large 2 · $3.00 / M input · Jul 24, 2024OpenAI · o1-mini · $3.00 / M input · Sep 12, 2024Google · Gemini 1.5 Pro · $1.25 / M input · Sep 24, 2024OpenAI · GPT-4o · $2.50 / M input · Oct 1, 2024Anthropic · Claude 3.5 Haiku · $1.00 / M input · Nov 4, 2024OpenAI · o1 · $15.00 / M input · Dec 5, 2024Meta · Llama 3.3 70B (Together) · $0.88 / M input · Dec 6, 2024Google · Gemini 2.0 Flash · $0.10 / M input · Dec 11, 2024xAI · Grok 2 · $2.00 / M input · Dec 12, 2024DeepSeek · DeepSeek-V3 · $0.14 / M input · Dec 26, 2024Alibaba · Qwen Plus · $0.40 / M input · Jan 25, 2025Mistral · Mistral Small 3 · $0.20 / M input · Jan 30, 2025OpenAI · o3-mini · $1.10 / M input · Jan 31, 2025DeepSeek · DeepSeek-V3 · $0.27 / M input · Feb 8, 2025DeepSeek · DeepSeek R1 · $0.55 / M input · Feb 8, 2025xAI · Grok 3 · $3.00 / M input · Feb 17, 2025xAI · Grok 3 Mini · $0.30 / M input · Feb 17, 2025Anthropic · Claude 3.7 Sonnet · $3.00 / M input · Feb 24, 2025Google · Gemini 2.5 Pro · $1.25 / M input · Mar 25, 2025Meta · Llama 4 Scout (Together) · $0.18 / M input · Apr 5, 2025Meta · Llama 4 Maverick (Together) · $0.27 / M input · Apr 5, 2025OpenAI · GPT-4.1 · $2.00 / M input · Apr 14, 2025OpenAI · o3 · $10.00 / M input · Apr 16, 2025OpenAI · o4-mini · $1.10 / M input · Apr 16, 2025Mistral · Mistral Medium 3 · $0.40 / M input · May 7, 2025Anthropic · Claude Opus 4 · $15.00 / M input · May 22, 2025Anthropic · Claude Sonnet 4 · $3.00 / M input · May 22, 2025OpenAI · o3 · $2.00 / M input · Jun 10, 2025OpenAI · o3-pro · $20.00 / M input · Jun 10, 2025Google · Gemini 2.5 Flash · $0.30 / M input · Jun 17, 2025xAI · Grok 4 · $3.00 / M input · Jul 9, 2025Google · Gemini 2.5 Flash-Lite · $0.10 / M input · Jul 22, 2025Anthropic · Claude Opus 4.1 · $15.00 / M input · Aug 5, 2025OpenAI · GPT-5 · $1.25 / M input · Aug 7, 2025Anthropic · Claude Sonnet 4.5 · $3.00 / M input · Sep 29, 2025Anthropic · Claude Haiku 4.5 · $1.00 / M input · Oct 15, 2025OpenAI · GPT-5.1 · $1.25 / M input · Nov 13, 2025Google · Gemini 3 Pro · $1.25 / M input · Nov 18, 2025Anthropic · Claude Opus 4.5 · $5.00 / M input · Nov 24, 2025Mistral · Mistral Large 3 · $2.00 / M input · Dec 2, 2025OpenAI · GPT-5.2 · $1.75 / M input · Dec 11, 2025Google · Gemini 3 Flash · $0.50 / M input · Dec 17, 2025OpenAI · GPT-5.3 · $1.75 / M input · Feb 5, 2026Anthropic · Claude Opus 4.6 · $5.00 / M input · Feb 5, 2026Alibaba · Qwen3.6 Plus · $0.33 / M input · Feb 12, 2026Anthropic · Claude Sonnet 4.6 · $3.00 / M input · Feb 17, 2026Google · Gemini 3.1 Pro · $1.25 / M input · Feb 19, 2026Google · Gemini 3.1 Flash-Lite · $0.25 / M input · Mar 3, 2026OpenAI · GPT-5.4 · $2.50 / M input · Mar 5, 2026xAI · Grok 4.20 · $2.00 / M input · Mar 10, 2026Mistral · Mistral Small 3.1 · $0.20 / M input · Mar 16, 2026Mistral · Mistral Small 4 · $0.10 / M input · Mar 16, 2026Alibaba · Qwen3.6 Plus · $0.50 / M input · Apr 2, 2026Anthropic · Claude Opus 4.7 · $5.00 / M input · Apr 16, 2026OpenAI · GPT-5.5 · $5.00 / M input · Apr 23, 2026DeepSeek · DeepSeek-V4-Flash · $0.14 / M input · Apr 24, 2026DeepSeek · DeepSeek-V4-Pro · $1.74 / M input · Apr 24, 2026Mistral · Mistral Medium 3.5 · $1.50 / M input · Apr 26, 2026xAI · Grok 4.3 · $1.25 / M input · Apr 30, 2026xAI · Grok Build 0.1 · $1.00 / M input · May 14, 2026Alibaba · Qwen3.7 Max · $2.50 / M input · May 15, 2026Google · Gemini 3.1 Pro · $2.00 / M input · May 19, 2026Google · Gemini 3.5 Flash · $1.50 / M input · May 19, 2026Anthropic · Claude Opus 4.8 · $5.00 / M input · May 28, 2026Alibaba · Qwen3.7 Plus · $0.40 / M input · May 31, 2026DeepSeek · DeepSeek-V4-Pro · $0.43 / M input · Jun 1, 2026xAI · Grok Composer 2.5 · $0.50 / M input · Jun 1, 2026Mistral · Mistral Large 3 · $0.50 / M input · Jun 4, 2026OpenAI · GPT-5.4 · $2.50 / M input · Jun 5, 2026Anthropic · Claude Fable 5 · $10.00 / M input · Jun 9, 2026xAI · Grok 4.3 · $1.25 / M input · Jun 10, 2026xAI · Grok 4.20 · $1.25 / M input · Jun 14, 2026Meta · Llama 3.3 70B (Together) · $1.04 / M input · Jun 18, 2026Anthropic · Claude Sonnet 5 · $2.00 / M input · Jun 30, 2026Mistral · Mistral Small 4 · $0.15 / M input · Jun 30, 2026xAI · Grok 4.5 · $2.00 / M input · Jul 8, 2026OpenAI · GPT-5.6 Sol · $5.00 / M input · Jul 9, 2026OpenAI · GPT-5.6 Terra · $2.50 / M input · Jul 9, 2026OpenAI · GPT-5.6 Luna · $1.00 / M input · Jul 9, 2026Meta · Muse Spark 1.1 · $1.25 / M input · Jul 9, 2026Google · Gemini 3.5 Flash-Lite · $0.30 / M input · Jul 21, 2026Google · Gemini 3.6 Flash · $1.50 / M input · Jul 21, 2026Anthropic · Claude Opus 5 · $5.00 / M input · Jul 24, 2026OpenAI · GPT-5.6 Terra · $2.00 / M input · Jul 30, 2026OpenAI · GPT-5.6 Luna · $0.20 / M input · Jul 30, 2026Alibaba · Qwen Max · $1.60 / M input · Jul 31, 2026Alibaba · Qwen3.8 Max · $2.00 / M input · Aug 3, 2026Meta · Muse Spark 1.2 · $1.25 / M input · Aug 5, 2026Meta · Muse Glimmer 30B (Together) · $0.35 / M input · Aug 10, 2026xAI · Grok 4.6 · $2.00 / M input · Aug 12, 2026Google · Gemini 3.6 Flash · $0.75 / M input · Aug 13, 2026Google · Gemini 3.7 Flash · $0.75 / M input · Aug 13, 2026DeepSeek · DeepSeek-V4-Flash · $0.22 / M input · Aug 16, 2026DeepSeek · DeepSeek-V4-Pro · $0.66 / M input · Aug 16, 2026Alibaba · Qwen3.8-27B · $0.50 / M input · Aug 19, 2026OpenAI · GPT-5.6 Sol · $4.00 / M input · Aug 21, 2026Alibaba · Qwen3.8-Flash · $0.15 / M input · Aug 26, 2026$ / M input (log)
AnthropicOpenAIGooglexAIMetaDeepSeekMistralAlibaba

Hover or tap any dot for the model, the full price (input, output, cached), the effective date, lifecycle status, and the source link.

The frontier price-floor in fourteen moves

The releases that reset what a flagship-class model costs per million input tokens. These are the moves worth knowing about, not a mechanical reading of the chart’s dashed line: that line is a strict all-time running minimum, so it stops stepping at $0.14 on May 6, 2024 — DeepSeek-V2 — and nothing charted since has gone lower. Most of the entries below sit above that floor and still mattered, because they moved the price of the model people actually reached for. Each links into the model’s row on the relevant per-family versions page for the full release context.

Every model’s current price

Sort by any column. Filter by provider, by tier, or to current-only. The Cached column shows the prompt-caching read rate where the provider documents one. Every lab on this page now publishes one for its current models, but coverage still varies by SKU — and where a cell is empty it is because the provider prints nothing there, not because the rate went unchecked.

Gemini 1.5 Flash
Cheap tier launch. Removed from Google's API pricing page by June 2026 — flipped legacy → deprecated; historical entry preserved as the original $0.075/M cheap-tier floor.
Google(Small / cheap)
$0.075
$0.30
May 14, 2024
Deprecated
Gemini 2.0 Flash
Held the $0.10 / $0.40 cheap-tier price through 2.5 Flash-Lite. Shut down June 1, 2026.
Google(Small / cheap)
$0.10
$0.40
Dec 11, 2024
Deprecated
Gemini 2.5 Flash-Lite
Held the 2.0 Flash cheap-tier price. The cached-input rate has held at $0.01 on Google's pricing page since launch and was added to this table on 2026-06-20.
Google(Small / cheap)
$0.10
$0.40
$0.01
Jul 22, 2025
Available
DeepSeek-V2
The price point that triggered the China-side AI price war.
DeepSeek(Frontier)
$0.14
$0.28
May 6, 2024
Deprecated
GPT-4o mini
60% cheaper than GPT-3.5 Turbo; replaced the cheap tier.
OpenAI(Small / cheap)
$0.15
$0.60
$0.075
Jul 18, 2024
Legacy
Mistral Small 4
Small 4 repriced up to $0.15 / $0.60 — a genuine post-launch raise above the $0.10 / $0.30 documented through 2026-06-20, confirmed against both the launch announcement's embedded model card and Mistral's pricing page. Cached input is $0.015, from the per-model Cached input column Mistral's inference docs publish; no effective date is given for it, so it sits on this entry rather than starting a new one.
Mistral(Small / cheap)
$0.15
$0.60
$0.015
Jun 30, 2026
Current
Qwen3.8-Flash
Hosted counterpart to Qwen3.8-Flash-Next, the open-weights preview of the architecture Qwen4 will be built on — 125B total with 6B active, released and hosted the same day. $0.15 / $0.47 list, implicit-cache input $0.016 (explicit cache creation $0.20, explicit read $0.016), no promotional discount running when the page was read. International (Singapore) USD basis. This is Alibaba's cheapest hosted text tier on this table and the family's first entry below the twenty-cent line.
Alibaba(Small / cheap)
$0.15
$0.47
$0.016
Aug 26, 2026
Current
Llama 4 Scout (Together)
10M-context open-weights model; Groq publishes the same weights at $0.11 / $0.34. Together lists this endpoint as Dedicated-only — it dropped the Llama 4 pair from its serverless table in June 2026 — and prints no cached rate, so the Cached cell is empty because Together publishes nothing there. Maverick is the current open-weights flagship in the Llama line; Scout is its smaller companion.
Meta(Frontier)
$0.18
$0.59
Apr 5, 2025
Available
Mistral Small 3
Re-Apache-2.0 era kickoff.
Mistral(Small / cheap)
$0.20
$0.60
Jan 30, 2025
Legacy
Mistral Small 3.1
Small tier held at the January 2025 reset price; superseded by Small 4. As of September 2026 the SKU no longer appears on Mistral's inference-pricing table at all — which now lists only Large 3, Medium 3.5, Small 4 and the Ministral 3 sizes, so neither do Small 3 or Medium 3 — so $0.20 / $0.60 is its last documented rate rather than a currently-quotable one. Left at legacy rather than deprecated: absence from a pricing table is not a deprecation notice, and Mistral has published none.
Mistral(Small / cheap)
$0.20
$0.60
Mar 16, 2026
Legacy
GPT-5.6 Luna
80% off both legs — the steepest single OpenAI price move charted here, matched only by o3's 2025 reset. OpenAI: "$0.20 per million input tokens and $1.20 per million output tokens for Luna." Cache reads $0.02; cache writes $0.25. Long-context (>272K) requests bill at 2x.
OpenAI(Small / cheap)
$0.20
$1.20
$0.02
Jul 30, 2026
Current
DeepSeek-V4-Flash
Flat-rate billing ends. From 16:00 UTC on August 16, 2026 DeepSeek bills peak / off-peak, and the figures here are the OFF-PEAK leg — the majority case and the lower of the two, recorded the way the xAI rows record the under-200K leg. Peak is exactly double on every leg: $0.014 cache-hit / $0.44 input / $1.32 output, in force 01:00-04:00 and 06:00-10:00 UTC on Monday through Friday only; every other hour, and all of Saturday and Sunday, is off-peak. That works out to 133 of the 168 hours in a week at the rate shown. Even the off-peak leg is a raise on the flat rate it replaces: input +57%, output +136%. An experimental multimodal sibling, deepseek-v4-flash-vision-exp, went API-only on August 21, 2026 at identical text-token rates, so it shares this row rather than duplicating the same price point.
DeepSeek(Small / cheap)
$0.22
$0.66
$0.007
Aug 16, 2026
Current
Claude 3 Haiku
The cheapest model from a frontier lab when it shipped — half GPT-3.5 Turbo's then-current $0.50 input, and an eighth of the $2.00 GPT-3.5 Turbo had launched at a year earlier.
Anthropic(Small / cheap)
$0.25
$1.25
Mar 13, 2024
Deprecated
Gemini 3.1 Flash-Lite
Cheap tier for the Gemini 3 series — replaces 2.5 Flash-Lite as the cost floor.
Google(Small / cheap)
$0.25
$1.50
$0.025
Mar 3, 2026
Available
DeepSeek-V3
Promotion ended; new permanent pricing. Off-peak (16:30–00:30 UTC): 50% off.
DeepSeek(Frontier)
$0.27
$1.10
$0.07
Feb 8, 2025
Deprecated
Llama 4 Maverick (Together)
Groq lists the same weights at $0.20 / $0.60. Together lists this endpoint as Dedicated-only and prints no cached rate, so the Cached cell is empty because Together publishes nothing there. Meta's newest open-weights release is Muse Glimmer 30B (Apache 2.0, August 2026), priced separately below.
Meta(Frontier)
$0.27
$0.85
Apr 5, 2025
Current
Grok 3 Mini
Cheapest xAI tier.
xAI(Small / cheap)
$0.30
$0.50
Feb 17, 2025
Deprecated
Gemini 2.5 Flash
Output stepped up vs 2.0 Flash to reflect added thinking capacity. The cached-input rate has held at $0.03 on Google's pricing page since launch and was added to this table on 2026-06-20.
Google(Small / cheap)
$0.30
$2.50
$0.03
Jun 17, 2025
Available
Gemini 3.5 Flash-Lite
Google's cheapest GA model in the 3.5 family — the same $0.30 / $2.50 tier 2.5 Flash carries, one step above 3.1 Flash-Lite. Input rate covers text / image / video / audio alike.
Google(Small / cheap)
$0.30
$2.50
$0.03
Jul 21, 2026
Available
Muse Glimmer 30B (Together)
Meta's first open-weights frontier-lab release since Llama 4 and the first ever under Apache 2.0 rather than a bespoke Llama license — weights at the new meta-models HuggingFace org, not meta-llama. Meta describes it as a local agent model distilled from Muse Spark, not a frontier model, hence the small tier. There is no Meta first-party API rate, so this row uses the same hosted-provider proxy the Llama rows use: Together AI's serverless rate for meta-models/Muse-Glimmer-30B, $0.35 input / $0.04 cached / $1.50 output. Other hosts run cheaper — resellers list it near $0.30 / $1.20 — so treat the Together figure as the reference point rather than the floor.
Meta(Small / cheap)
$0.35
$1.50
$0.04
Aug 10, 2026
Current
Qwen Plus
First point on the International (Singapore) USD basis — the same basis the Qwen3.7 rows use, and the basis every current Qwen rate on this page is quoted from. Δ is suppressed against the 2024 predecessor because that point is on Alibaba's Chinese-mainland table; the difference between the two is a change of reference table, not a price move. The `qwen-plus` alias resolves to qwen-plus-2025-12-01, and every International checkpoint back to qwen-plus-2025-01-25 (the oldest International one still published) carries this rate, so 2025-01-25 is the earliest date the current price can be sourced to. Output is the non-thinking rate; thinking mode bills output at $4.00. Prompts above 256K tokens bill at $1.20 / $3.60. Cached input is the implicit-cache rate of $0.08 (explicit cache creation $0.50, explicit cache read $0.04); Batch File rates are $0.20 / $0.60, a flat 50%.
Alibaba(Frontier)
$0.40
$1.20
$0.08
Jan 25, 2025
Available
Mistral Medium 3
Launch — Mistral's announcement explicitly cites '$0.4 / M input, $2 / M output' as the 'medium is the new large' mid-tier positioning. Superseded by Medium 3.5 on April 26, 2026, and no longer listed on Mistral's inference-pricing table.
Mistral(Frontier)
$0.40
$2.00
May 7, 2025
Legacy
Qwen3.7 Plus
Multimodal agent flagship; ~1/6 the per-token price of text-only Qwen3.7-Max. Closed-weights on Alibaba Cloud Bailian / Model Studio (DashScope `qwen3.7-plus`). Alibaba's model page lists $0.40 / $1.60 for prompts up to 256K in both thinking and non-thinking mode, with a limited-time 20% off shown alongside ($0.32 / $1.28); above 256K bills at $1.20 / $4.80. List is the figure shown here. Cached input is Alibaba's implicit-cache rate of $0.08 — the same field every current Qwen row uses; the separate explicit-cache product reads at $0.04 and creates at $0.50.
Alibaba(Frontier)
$0.40
$1.60
$0.08
May 31, 2026
Current
GPT-3.5 Turbo
Further 50% cut on input, 25% cut on output. Still billable at this rate: OpenAI's deprecation schedule shuts the gpt-3.5-turbo snapshots down on October 23, 2026 and names GPT-5.6 Terra as the substitute.
OpenAI(Small / cheap)
$0.50
$1.50
Jan 25, 2024
Legacy
Gemini 1.0 Pro
Paid tier opened after a free preview window.
Google(Frontier)
$0.50
$1.50
Feb 15, 2024
Deprecated
Mistral Large 3
75% input cut, 75% output cut — pricing-page reset; Large positioned as the lower-cost open-weights tier alongside Medium 3.5. Mistral's own FAQ block still quotes the pre-cut $2 in / $6 out for Mistral Large; the model card on the same site reads $0.50 / $1.50, and the card is the one to trust. Cached input is $0.05, from the per-model Cached input column Mistral's inference docs publish — Mistral gives no effective date for it, so it sits on this entry rather than starting a new one.
Mistral(Frontier)
$0.50
$1.50
$0.05
Jun 4, 2026
Current
Gemini 3 Flash
Flash tier stepped up from $0.30 / $2.50 (2.5 Flash); audio input $1.00.
Google(Small / cheap)
$0.50
$3.00
$0.05
Dec 17, 2025
Available
Qwen3.6 Plus
First point on the International (Singapore) USD basis — the same basis the Qwen3.7 rows use. Δ is suppressed against the 2026-02 predecessor because that point is on Alibaba's Chinese-mainland table; the difference between the two is a change of reference table, not a price move. The `qwen3.6-plus` alias resolves to qwen3.6-plus-2026-04-02, the only published checkpoint, and it carries $0.50 / $3.00 for prompts up to 256K (thinking and non-thinking alike); prompts above 256K bill at $2.00 / $6.00. Alibaba publishes no implicit-cache rate for this model — only explicit cache creation ($0.625) and explicit cache read ($0.05) — so the Cached figure here is the explicit-read rate rather than the implicit one the Qwen3.7 and 3.8 rows use.
Alibaba(Frontier)
$0.50
$3.00
$0.05
Apr 2, 2026
Available
Grok Composer 2.5
Agentic-coding model selectable inside Grok Build. xAI publishes no live per-token rate for Composer 2.5 — it is absent from the Text API pricing table that lists every other xAI text SKU, including Grok Build 0.1 and all three Grok 4.20 variants, and it has no model-detail page. The launch post says only that it is available to SuperGrok and X Premium+ subscribers inside Grok Build, and the $0.50 / $2.50 here is the figure that post quoted. Read it as a subscription-bundled model rather than one you can buy per token.
xAI(Small / cheap)
$0.50
$2.50
Jun 1, 2026
Current
Qwen3.8-27B
Alibaba's newest open-weights release — 27B dense, Apache 2.0, natively vision-language — weights published August 14, 2026 and hosted on Qwen Cloud from August 19, which is the date this entry records because it is the first date the model had a per-token price. $0.50 / $3.00 list with implicit-cache input at $0.10; explicit cache creation $0.625, explicit cache read $0.05. No promotional discount was running when the page was read, so list and charged rate are the same figure. Read on the International (Singapore) USD basis, the basis every current Qwen row on this page uses. Small tier despite the Versions page's Flagship chip: at a quarter of Qwen3.8-Max's rate this is Alibaba's cheap tier by price, and the Versions chip tracks family lineage rather than price class.
Alibaba(Small / cheap)
$0.50
$3.00
$0.10
Aug 19, 2026
Current
DeepSeek R1
Reasoning-track companion to V3. Off-peak: 75% off.
DeepSeek(Reasoning)
$0.55
$2.19
$0.14
Feb 8, 2025
Deprecated
DeepSeek-V4-Pro
The end of DeepSeek's flat-rate era, and the largest single price increase on DeepSeek's frontier tier since the V2-to-V3 reset. From 16:00 UTC on August 16, 2026 the API bills peak / off-peak; the figures here are the OFF-PEAK leg, the same way the xAI rows show the under-200K leg — the majority case and the lower of the two. Peak is exactly double on every leg: $0.044 cache-hit / $1.32 input / $3.96 output, in force 01:00-04:00 and 06:00-10:00 UTC on Monday through Friday only; every other hour, and all of Saturday and Sunday, is off-peak, which is 133 of the 168 hours in a week. Off-peak input is +52% and off-peak output +128% on the flat rate it replaces; peak is +203% and +355%.
DeepSeek(Frontier)
$0.66
$1.98
$0.022
Aug 16, 2026
Current
Gemini 3.6 Flash
Halved retroactively — an unusual move for a superseded model. When Gemini 3.7 Flash shipped at an introductory $0.75 / $3.75, Google applied the same cut to 3.6 Flash rather than leaving the older model priced above its successor: 'We're also applying this new rate to 3.6 Flash' (ai.google.dev/gemini-api/docs/latest-model, August 13, 2026). Both models revert to $1.50 / $7.50 on January 1, 2027, and cached input reverts $0.075 to $0.15 with them. Verified on the pricing page 2026-08-21.
Google(Frontier)
$0.75
$3.75
$0.075
Aug 13, 2026
Available
Gemini 3.7 Flash
Google's current frontier-class Flash, and the first Gemini flagship to launch on an explicitly time-boxed introductory rate: half of 3.6 Flash's launch price through December 31, 2026, reverting to $1.50 / $7.50 on January 1, 2027 (cached input $0.075 to $0.15). Batch halves both legs again ($0.375 / $1.875). The figure shown is the rate in force today; the January reversion is a scheduled announcement rather than a price anyone has been charged, so it stays in this note instead of on the chart.
Google(Frontier)
$0.75
$3.75
$0.075
Aug 13, 2026
Current
Claude Instant
Cut alongside Claude 2.1 launch.
Anthropic(Small / cheap)
$0.80
$2.40
Nov 21, 2023
Deprecated
Llama 2 70B (Together)
Reference hosted price via Together AI; Llama is open-weights.
Meta(Frontier)
$0.90
$0.90
Oct 1, 2023
Deprecated
Llama 3 70B (Together)
Reference Together AI launch price.
Meta(Frontier)
$0.90
$0.90
Apr 18, 2024
Deprecated
Claude 3.5 Haiku
Raised Haiku from $0.25 / $1.25 — the only mid-generation increase the Haiku tier has taken. Anthropic has since retired it from the first-party API; it still bills on Bedrock and Google Cloud, and Anthropic's table now carries it there at a lower $0.80 / $0.08 cached / $4.00.
Anthropic(Small / cheap)
$1.00
$5.00
$0.10
Nov 4, 2024
Deprecated
Claude Haiku 4.5
Haiku held the $1 / $5 tier Claude 3.5 Haiku launched at. Anthropic's table now carries the retired 3.5 Haiku lower, at $0.80 / $4.00 on Bedrock and Google Cloud, so 4.5 Haiku is the more expensive of the two on paper — the Haiku line has never been cut on Anthropic's own API. Cache reads $0.10, 5m cache writes $1.25, 1h cache writes $2; Batch halves it to $0.50 / $2.50.
Anthropic(Small / cheap)
$1.00
$5.00
$0.10
Oct 15, 2025
Current
Grok Build 0.1
Coding-focused agentic model — successor to grok-code-fast-1. 256K context.
xAI(Small / cheap)
$1.00
$2.00
$0.20
May 14, 2026
Current
Llama 3.3 70B (Together)
Together AI repriced the endpoint upward to $1.04 / $1.04 from the $0.88 / $0.88 it launched at; Together's own model page carries both figures across the history and aggregators still quote the older one. Together's model-specifications block for this endpoint lists an input price and an output price only — no cached figure — so the Cached cell is empty because Together publishes nothing there, not because it went unchecked.
Meta(Frontier)
$1.04
$1.04
Jun 18, 2026
Available
o3-mini
Replaced o1-mini at one-third the cost. Still billable at this rate; shutdown scheduled for October 23, 2026, with GPT-5.6 Sol as the substitute.
OpenAI(Reasoning)
$1.10
$4.40
$0.55
Jan 31, 2025
Legacy
o4-mini
Same shelf price as o3-mini. Still billable at this rate; shutdown scheduled for October 23, 2026, with GPT-5.6 Terra as the substitute.
OpenAI(Reasoning)
$1.10
$4.40
$0.28
Apr 16, 2025
Legacy
Gemini 1.5 Pro
82% input cut, 76% output cut. >128K: $2.50 / $10.
Google(Frontier)
$1.25
$5.00
Sep 24, 2024
Deprecated
Gemini 2.5 Pro
≤200K context; >200K tier at $2.50 / $15 with cached at $0.25. The ≤200K cached-input rate has held at $0.125 on Google's pricing page since launch and was added to this table on 2026-06-20.
Google(Frontier)
$1.25
$10.00
$0.12
Mar 25, 2025
Available
GPT-5
Unified router across reasoning and chat; first OpenAI frontier under $2 input.
OpenAI(Frontier)
$1.25
$10.00
$0.12
Aug 7, 2025
Legacy
GPT-5.1
First post-GPT-5 update; Instant and Thinking modes. Pricing held at GPT-5 tier.
OpenAI(Frontier)
$1.25
$10.00
$0.12
Nov 13, 2025
Legacy
Gemini 3 Pro
Held the 2.5 Pro price.
Google(Frontier)
$1.25
$10.00
Nov 18, 2025
Deprecated
Grok 4.20
xAI's pricing table lists all three 4.20 SKUs — reasoning, non-reasoning, and multi-agent — at $1.25 in / $0.20 cached / $2.50 out for prompts under 200K, and exactly double for prompts at or above it. The $2.00 / $6.00 shown for March 2026 was the introductory rate; xAI cut it without an announcement, so the cut is dated to when the lower figure first appeared on the pricing table rather than to a launch post. Aggregators still quote the introductory rate.
xAI(Frontier)
$1.25
$2.50
$0.20
Jun 14, 2026
Available
Grok 4.3
$1.25 in / $0.20 cached / $2.50 out for prompts under 200K, and exactly double for prompts at or above it. The $0.125 cached rate shown at launch was never a figure xAI published; $0.20 is what its pricing table carries.
xAI(Frontier)
$1.25
$2.50
$0.20
Jun 10, 2026
Available
Muse Spark 1.1
Meta's first paid API — the Meta Model API public preview, launched with Muse Spark 1.1 GA (Jul 9, 2026). $1.25 / $4.25 per M is Meta's own Meta Model API list price, ~one-quarter of the Anthropic / OpenAI frontier tier. Reasoning (chain-of-thought) tokens bill at the $4.25 output rate. Direct Meta list price, not an open-weights hosting proxy like the Llama rows. Superseded as Meta's flagship by Muse Spark 1.2 on August 5, 2026 at the same rate; still served. Meta prices prompt-cache reads at $0.15 across both Spark releases.
Meta(Frontier)
$1.25
$4.25
$0.15
Jul 9, 2026
Available
Muse Spark 1.2
Meta's current frontier flagship, GA on the Meta Model API the day it was announced. The rate holds at Muse Spark 1.1's $1.25 / $4.25 per M — a generation change with no price move — and prompt-cache reads at $0.15. A second SKU, muse-spark-1.2-contributor, lists at $0.10 in / $0.002 cached / $0.20 out, roughly 12x and 21x cheaper, in exchange for permission to train future Meta models on your prompts and completions. That tier is not a row here: this page compares list prices for equivalent service, and a rate conditioned on surrendering data rights is not the same product. Its $0.20 output would undercut every output rate this page records, historical ones included — which is exactly why letting it in would distort the comparison.
Meta(Frontier)
$1.25
$4.25
$0.15
Aug 5, 2026
Current
Mistral Medium 3.5
128B dense open-weights flagship under a modified MIT license; current Mistral frontier. Cached input is $0.15, from the per-model Cached input column Mistral's inference docs publish across every family; the figures land exactly on 0.1x input. Mistral gives no effective date for the cached rates, so they sit on this entry rather than starting a new one. Batch is a flat 50%.
Mistral(Frontier)
$1.50
$7.50
$0.15
Apr 26, 2026
Current
Gemini 3.5 Flash
Frontier-class Flash priced above the 3.x Flash line — 3x the 3 Flash rate. Output (including thinking tokens) $9 / M. Superseded as Google's workhorse Flash by Gemini 3.6 Flash on July 21, 2026, and still listed at this rate.
Google(Frontier)
$1.50
$9.00
$0.15
May 19, 2026
Available
Qwen Max
First point on the International (Singapore) USD basis — the same basis the Qwen3.7 rows use. Alibaba made NO price change on this date: the undated `qwen-max` alias row reads $1.60 / $6.40, no tiered pricing, non-thinking mode only, and Alibaba publishes no effective date and no dated qwen-max checkpoint for it, so the entry is dated to the day the rate was observed. Δ is suppressed against the 2024 predecessor because that point is on Alibaba's Chinese-mainland table; the difference between the two is a change of reference table, not a price move. Cached input is the implicit-cache rate of $0.32; Batch File rates are $0.80 / $3.20, a flat 50%. 32K context.
Alibaba(Frontier)
$1.60
$6.40
$0.32
Jul 31, 2026
Available
GPT-5.2
Three modes (Instant / Thinking / Pro). 40% input bump and 40% output bump vs GPT-5.1; SOTA on FrontierMath.
OpenAI(Frontier)
$1.75
$14.00
$0.17
Dec 11, 2025
Available
GPT-5.3
Codex-specialized 5.3 shipped Feb 5; 5.3 Instant chat-latest two days later. Pricing held at the 5.2 tier. gpt-5.3-chat-latest shut down on August 10, 2026, and the general 5.3 SKU no longer appears anywhere on OpenAI's pricing table (5.2, 5.1 and 5 all still do). Only gpt-5.3-codex survives, still at this $1.75 / $14 tier.
OpenAI(Frontier)
$1.75
$14.00
$0.17
Feb 5, 2026
Deprecated
Grok 2
First paid xAI API model; vision variant same price.
xAI(Frontier)
$2.00
$10.00
Dec 12, 2024
Deprecated
GPT-4.1
First OpenAI 1M-context production model; 75% prompt-cache discount.
OpenAI(Frontier)
$2.00
$8.00
$0.50
Apr 14, 2025
Legacy
o3
80% price cut three months post-launch.
OpenAI(Reasoning)
$2.00
$8.00
$0.50
Jun 10, 2025
Legacy
Gemini 3.1 Pro
Repriced alongside the Gemini 3.1 Preview cycle; >200K tier at $4 / $18.
Google(Frontier)
$2.00
$12.00
$0.20
May 19, 2026
Current
Claude Sonnet 5
New default mid-tier workhorse; 1M context, 128K output. $2 / $10 with cache reads at $0.20 (5m cache writes $2.50, 1h $4.00); Batch halves it to $1 / $5. The scheduled step-up has been cancelled: as of 2026-08-13 Anthropic's pricing page states that the $2 / $10 launch rate, originally introductory pricing through August 31, 2026, "is now the standard price" and that "the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur." That date has now passed with no increase.
Anthropic(Frontier)
$2.00
$10.00
$0.20
Jun 30, 2026
Current
Grok 4.5
500K context, configurable reasoning effort (low/medium/high). xAI's model docs carry a two-row tiered table: prompts under 200K bill at $2.00 / $0.30 cached / $6.00, prompts at or above 200K bill the whole request at $4.00 / $0.60 / $12.00. The under-200K leg is the one shown here. Grok 4.6 took the flagship slot on August 12, 2026; xAI still lists 4.5 on its pricing table at the same rate.
xAI(Frontier)
$2.00
$6.00
$0.30
Jul 8, 2026
Available
GPT-5.6 Terra
20% cut. OpenAI: "Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra." Cache reads hold the 90% discount ($0.20); cache writes $2.50. Long-context (>272K) requests bill at 2x.
OpenAI(Frontier)
$2.00
$12.00
$0.20
Jul 30, 2026
Current
Qwen3.8 Max
Alibaba's current flagship, GA 2026-08-03. $2.00 / $6.00 per M is a 20% cut on both sides against Qwen3.7-Max's $2.50 / $7.50 — the newer, more capable Max is also the cheaper one. Read off Alibaba's own model page, which became the canonical pricing surface once the older Model Studio price roster froze in July 2026. The open-weights artifact released alongside it, Qwen3.8-2.4T-A95B, is hosted at the identical $2.00 / $6.00 / $0.25, so it shares this row. Cached input is the implicit-cache rate of $0.25 per M; explicit cache creation lists at $2.50 and explicit cache read at $0.17. Alibaba advertises a Batch API for this model but publishes no discount percentage for it, unlike the flat 50% it prints for Qwen-Plus and Qwen-Max.
Alibaba(Frontier)
$2.00
$6.00
$0.25
Aug 3, 2026
Current
Grok 4.6
Current xAI flagship, announced August 12, 2026 — xAI's docs now name Grok 4.6 as the model to use "for everything else, including code." 500K context. Same prompt-length-tiered shape as Grok 4.5: prompts under 200K bill at $2.00 / $0.50 cached / $6.00, prompts at or above 200K bill the whole request at $4.00 / $1.00 / $12.00; the under-200K leg is the one shown here. Headline input and output hold at Grok 4.5's $2.00 / $6.00 — the only rate that moved is cached input, up from $0.30 to $0.50 (a 75% discount rather than 85%). Ship date is August 12, not the August 7 several aggregators give; August 7 was a pre-announced target with no matching entry in xAI's changelog.
xAI(Frontier)
$2.00
$6.00
$0.50
Aug 12, 2026
Current
Mistral Medium (original)
Mistral's original mid-tier on la Plateforme.
Mistral(Frontier)
$2.50
$7.50
Dec 11, 2023
Deprecated
GPT-4o
50% input cut, 33% output cut; prompt caching introduced.
OpenAI(Frontier)
$2.50
$10.00
$1.25
Oct 1, 2024
Legacy
GPT-5.4
OpenAI's pricing table carries GPT-5.4 at $2.50 / $15.00 for context lengths under 272K, with cache reads at $0.25.
OpenAI(Frontier)
$2.50
$15.00
$0.25
Jun 5, 2026
Available
Qwen3.7 Max
Reasoning-agent flagship; 1M context with native extended thinking. Superseded at the top of the Max tier by Qwen3.8-Max on 2026-08-03, which is both more capable on Alibaba's own ordering and cheaper — leaving this row priced above its own successor. Alibaba's list block reads $2.50 input / $7.50 output / $0.50 implicit cache, and a limited-time discount has run alongside it at half those figures; list is what is shown here, which is how this page treats Alibaba's undated temporary discounts. Explicit cache creation lists at $3.125 and explicit cache read at $0.25.
Alibaba(Frontier)
$2.50
$7.50
$0.50
May 15, 2026
Available
Claude 3 Sonnet
Established the Sonnet $3/$15 price that has held through 4.x.
Anthropic(Frontier)
$3.00
$15.00
Mar 4, 2024
Deprecated
Claude 3.5 Sonnet
Prompt caching introduced; 90% cache discount.
Anthropic(Frontier)
$3.00
$15.00
$0.30
Jun 20, 2024
Deprecated
Mistral Large 2
62% input cut, 62% output cut vs Large.
Mistral(Frontier)
$3.00
$9.00
Jul 24, 2024
Deprecated
o1-mini
Cheap reasoning track companion to o1-preview.
OpenAI(Reasoning)
$3.00
$12.00
$1.50
Sep 12, 2024
Deprecated
Grok 3
First xAI 1M-context model.
xAI(Frontier)
$3.00
$15.00
$0.75
Feb 17, 2025
Deprecated
Claude 3.7 Sonnet
Extended-thinking mode; Sonnet price unchanged.
Anthropic(Frontier)
$3.00
$15.00
$0.30
Feb 24, 2025
Deprecated
Claude Sonnet 4
Sonnet price unchanged into the 4.x line.
Anthropic(Frontier)
$3.00
$15.00
$0.30
May 22, 2025
Deprecated
Grok 4
Held Grok 3 pricing.
xAI(Frontier)
$3.00
$15.00
$0.75
Jul 9, 2025
Deprecated
Claude Sonnet 4.5
1M-context beta on Vertex / Bedrock; standard pricing for ≤200K.
Anthropic(Frontier)
$3.00
$15.00
$0.30
Sep 29, 2025
Available
Claude Sonnet 4.6
1M-context default; superseded by Sonnet 5 as the default mid-tier on June 30, 2026.
Anthropic(Frontier)
$3.00
$15.00
$0.30
Feb 17, 2026
Available
Llama 3.1 405B (Together)
405B-parameter open-weights flagship hosted price via Together AI. Together Serverless API removed the model from its lineup by June 15, 2026 (model page now reads "This model is not available on Together's Serverless API").
Meta(Frontier)
$3.50
$3.50
Jul 23, 2024
Deprecated
GPT-5.6 Sol
A 20% input cut and a 33% output cut on OpenAI's flagship — and the first Sol move since launch. This one carries an expiry: OpenAI's pricing page states the promotional rate "is available at least through November 21, 2026," so a caller quoting $4 / $20 in December may be quoting a rate that no longer exists. Cache reads hold the 90% discount ($0.40); cache writes bill at 1.25x input ($5.00). Prompts above the 272K short-context boundary bill on the long-context tier at $8.00 input / $0.80 cached / $10.00 cache write / $30.00 output — note the output leg there did not move, so the cut narrows as prompts grow. Fast mode still bills at 2x standard. A cybersecurity-specialized fork, gpt-5.6-cyber (alias daybreak-red-latest), lists at $12.50 / $1.25 cached / $75 per M with 400K context, but is not a row here: OpenAI files it under a separate "Cyber models" table and it is reachable only after separate Daybreak Red approval and provisioning rather than ordinary API access, while the rest of this comparison is rates any caller can buy at will.
OpenAI(Frontier)
$4.00
$20.00
$0.40
Aug 21, 2026
Current
Claude Opus 4.5
First Opus price cut — $15 / $75 reset to $5 / $25. Third member of the 4.5 generation after Sonnet and Haiku.
Anthropic(Frontier)
$5.00
$25.00
$0.50
Nov 24, 2025
Available
Claude Opus 4.6
Opus tier held at $5 / $25 from 4.5.
Anthropic(Frontier)
$5.00
$25.00
$0.50
Feb 5, 2026
Available
Claude Opus 4.7
Opus tier held at $5 / $25 — established on Opus 4.5 in November 2025.
Anthropic(Frontier)
$5.00
$25.00
$0.50
Apr 16, 2026
Available
GPT-5.5
GA ChatGPT-default tier; 90% cache discount. Superseded as OpenAI's frontier flagship by GPT-5.6 on July 9, 2026, and still listed at this rate for context lengths under 272K. The higher-compute gpt-5.5-pro variant lists separately at $30 / $180.
OpenAI(Frontier)
$5.00
$30.00
$0.50
Apr 23, 2026
Available
Claude Opus 4.8
Same $5 / $25 tier as 4.7; 5m cache writes $6.25, 1h cache writes $10. Optional fast mode at $10 / $50. Claude Opus 5 took the Opus slot on July 24, 2026.
Anthropic(Frontier)
$5.00
$25.00
$0.50
May 28, 2026
Available
Claude Opus 5
New Opus generation at the $5 / $25 tier Opus 4.5 established in November 2025 — a step-change in capability at a held price. 1M context. 5m cache writes $6.25, 1h cache writes $10, cache reads $0.50. Fast mode bills at $10 / $50.
Anthropic(Frontier)
$5.00
$25.00
$0.50
Jul 24, 2026
Current
Claude 2
Reduced alongside Claude 2.1 launch.
Anthropic(Frontier)
$8.00
$24.00
Nov 21, 2023
Deprecated
Mistral Large
Top-tier proprietary launch.
Mistral(Frontier)
$8.00
$24.00
Feb 26, 2024
Deprecated
GPT-4 Turbo
DevDay 2023; first 128K-context OpenAI model; 3x cheaper input vs GPT-4. Still billable at this rate; shutdown scheduled for October 23, 2026, with GPT-5.6 Sol as the substitute.
OpenAI(Frontier)
$10.00
$30.00
Nov 6, 2023
Legacy
Claude Fable 5
First generally-available Mythos-class model — a tier above Opus. Half the rate of Claude Mythos Preview; 5m cache writes $12.50, 1h cache writes $20. Access was suspended June 12, 2026 under a US export-control directive (pricing intact); controls lifted June 30 and the model redeployed globally July 1, 2026 at the same $10 / $50. Batch halves it to $5 / $25. Its invitation-only sibling Claude Mythos 5 — the same underlying model with safeguards lifted in some areas, offered to a small group of cyberdefenders under Project Glasswing — lists at the identical $10 / $1 cached / $50 but is not a separate row: Anthropic marks it "limited availability" rather than generally available, and this page compares rates any caller can buy.
Anthropic(Frontier)
$10.00
$50.00
$1.00
Jun 9, 2026
Current
Claude 1
Original Claude API pricing.
Anthropic(Frontier)
$11.02
$32.68
May 15, 2023
Deprecated
Claude 3 Opus
Top-tier launch pricing established the Opus tier.
Anthropic(Frontier)
$15.00
$75.00
Mar 4, 2024
Deprecated
o1
First production reasoning model; reasoning tokens billed as output. Still billable at this rate; shutdown scheduled for October 23, 2026, with GPT-5.6 Sol as the substitute.
OpenAI(Reasoning)
$15.00
$60.00
$7.50
Dec 5, 2024
Legacy
Claude Opus 4
Opus tier carried through from Claude 3.
Anthropic(Frontier)
$15.00
$75.00
$1.50
May 22, 2025
Deprecated
Claude Opus 4.1
Opus pricing held at $15 / $75 for the life of the model. Retired from Anthropic's first-party API on August 5, 2026; Anthropic's pricing table now labels the row "retired, except on Bedrock and Google Cloud", where it still bills at this rate.
Anthropic(Frontier)
$15.00
$75.00
$1.50
Aug 5, 2025
Deprecated
o3-pro
Higher-compute o3 variant; Responses API only. OpenAI's pricing table publishes no cached-input rate for o3-pro, so the Cached cell is empty because the provider leaves it empty.
OpenAI(Reasoning)
$20.00
$80.00
Jun 10, 2025
Legacy
GPT-4
Launch pricing for the 8K context model, and still what gpt-4-0613 bills three years on. OpenAI's deprecation schedule shuts it down on October 23, 2026 and names GPT-5.6 Sol as the substitute.
OpenAI(Frontier)
$30.00
$60.00
Mar 14, 2023
Legacy
GPT-3 (text-davinci-003)
Davinci tier at launch; $0.06 per 1K tokens, no input/output split.
OpenAI(Frontier)
$60.00
$60.00
Jun 11, 2020
Deprecated
GPT-4 32K
Extended-context GPT-4.
OpenAI(Frontier)
$60.00
$120.00
Jun 13, 2023
Deprecated

Showing all 101 models.

By provider

Each provider’s current flagship — as listed on the frontier-model roster — plus its cheapest input tier in production and its most-recent price-change log. Δ column shows the per-million input-price change from the same model’s prior price point; “tier hold” means the price held against the predecessor in the same family, and “rebased” means the two points are quoted on different pricing bases, so no price move can be read from the difference.

Anthropic
Claude family
Versions →
Current flagship
$10.00 / $50.00
Claude Fable 5
Cheapest input in production
$1.00 / M
Claude Haiku 4.5
Recent price-change log
EffectiveModelIn / OutΔ
Jul 24, 2026Claude Opus 5$5.00 / $25.00launch
Jun 30, 2026Claude Sonnet 5$2.00 / $10.00launch
Jun 9, 2026Claude Fable 5$10.00 / $50.00launch
May 28, 2026Claude Opus 4.8$5.00 / $25.00launch
Apr 16, 2026Claude Opus 4.7$5.00 / $25.00launch
Feb 17, 2026Claude Sonnet 4.6$3.00 / $15.00launch
Feb 5, 2026Claude Opus 4.6$5.00 / $25.00launch
OpenAI
ChatGPT family
Versions →
Current flagship
$4.00 / $20.00
GPT-5.6 Sol
Cheapest input in production
$0.20 / M
GPT-5.6 Luna
Recent price-change log
EffectiveModelIn / OutΔ
Aug 21, 2026GPT-5.6 Sol$4.00 / $20.00↓ 20%
Jul 30, 2026GPT-5.6 Terra$2.00 / $12.00↓ 20%
Jul 30, 2026GPT-5.6 Luna$0.20 / $1.20↓ 80%
Jul 9, 2026GPT-5.6 Sol$5.00 / $30.00launch
Jul 9, 2026GPT-5.6 Terra$2.50 / $15.00launch
Jul 9, 2026GPT-5.6 Luna$1.00 / $6.00launch
Jun 5, 2026GPT-5.4$2.50 / $15.00tier hold
Google
Gemini family
Versions →
Current flagship
$0.75 / $3.75
Gemini 3.7 Flash
Cheapest input in production
$0.10 / M
Gemini 2.5 Flash-Lite
Recent price-change log
EffectiveModelIn / OutΔ
Aug 13, 2026Gemini 3.6 Flash$0.75 / $3.75↓ 50%
Aug 13, 2026Gemini 3.7 Flash$0.75 / $3.75launch
Jul 21, 2026Gemini 3.5 Flash-Lite$0.30 / $2.50launch
Jul 21, 2026Gemini 3.6 Flash$1.50 / $7.50launch
May 19, 2026Gemini 3.1 Pro$2.00 / $12.00↑ 60%
May 19, 2026Gemini 3.5 Flash$1.50 / $9.00launch
Mar 3, 2026Gemini 3.1 Flash-Lite$0.25 / $1.50launch
xAI
Grok family
Versions →
Current flagship
$2.00 / $6.00
Grok 4.6
Cheapest input in production
$0.50 / M
Grok Composer 2.5
Recent price-change log
EffectiveModelIn / OutΔ
Aug 12, 2026Grok 4.6$2.00 / $6.00launch
Jul 8, 2026Grok 4.5$2.00 / $6.00launch
Jun 14, 2026Grok 4.20$1.25 / $2.50↓ 38%
Jun 10, 2026Grok 4.3$1.25 / $2.50tier hold
Jun 1, 2026Grok Composer 2.5$0.50 / $2.50launch
May 14, 2026Grok Build 0.1$1.00 / $2.00launch
Apr 30, 2026Grok 4.3$1.25 / $2.50launch
Meta
Llama family
Versions →
Current flagship
$1.25 / $4.25
Muse Spark 1.2
Cheapest input in production
$0.18 / M
Llama 4 Scout (Together)
Recent price-change log
EffectiveModelIn / OutΔ
Aug 10, 2026Muse Glimmer 30B (Together)$0.35 / $1.50launch
Aug 5, 2026Muse Spark 1.2$1.25 / $4.25launch
Jul 9, 2026Muse Spark 1.1$1.25 / $4.25launch
Jun 18, 2026Llama 3.3 70B (Together)$1.04 / $1.04↑ 18%
Apr 5, 2025Llama 4 Scout (Together)$0.18 / $0.59launch
Apr 5, 2025Llama 4 Maverick (Together)$0.27 / $0.85launch
Dec 6, 2024Llama 3.3 70B (Together)$0.88 / $0.88launch
DeepSeek
DeepSeek family
Versions →
Current flagship
$0.66 / $1.98
DeepSeek-V4-Pro
Cheapest input in production
$0.22 / M
DeepSeek-V4-Flash
Recent price-change log
EffectiveModelIn / OutΔ
Aug 16, 2026DeepSeek-V4-Flash$0.22 / $0.66↑ 57%
Aug 16, 2026DeepSeek-V4-Pro$0.66 / $1.98↑ 52%
Jun 1, 2026DeepSeek-V4-Pro$0.43 / $0.87↓ 75%
Apr 24, 2026DeepSeek-V4-Flash$0.14 / $0.28launch
Apr 24, 2026DeepSeek-V4-Pro$1.74 / $3.48launch
Feb 8, 2025DeepSeek-V3$0.27 / $1.10↑ 93%
Feb 8, 2025DeepSeek R1$0.55 / $2.19launch
Mistral
Mistral family
Versions →
Current flagship
$1.50 / $7.50
Mistral Medium 3.5
Cheapest input in production
$0.15 / M
Mistral Small 4
Recent price-change log
EffectiveModelIn / OutΔ
Jun 30, 2026Mistral Small 4$0.15 / $0.60↑ 50%
Jun 4, 2026Mistral Large 3$0.50 / $1.50↓ 75%
Apr 26, 2026Mistral Medium 3.5$1.50 / $7.50launch
Mar 16, 2026Mistral Small 3.1$0.20 / $0.60launch
Mar 16, 2026Mistral Small 4$0.10 / $0.30launch
Dec 2, 2025Mistral Large 3$2.00 / $6.00launch
May 7, 2025Mistral Medium 3$0.40 / $2.00launch
Alibaba
Qwen family
Versions →
Current flagship
$2.00 / $6.00
Qwen3.8 Max
Cheapest input in production
$0.15 / M
Qwen3.8-Flash
Recent price-change log
EffectiveModelIn / OutΔ
Aug 26, 2026Qwen3.8-Flash$0.15 / $0.47launch
Aug 19, 2026Qwen3.8-27B$0.50 / $3.00launch
Aug 3, 2026Qwen3.8 Max$2.00 / $6.00launch
Jul 31, 2026Qwen Max$1.60 / $6.40rebased
May 31, 2026Qwen3.7 Plus$0.40 / $1.60launch
May 15, 2026Qwen3.7 Max$2.50 / $7.50launch
Apr 2, 2026Qwen3.6 Plus$0.50 / $3.00rebased

Notes and caveats

Headline tier is API list price. Every figure on this page is the public per-million-token rate the provider documents on its own pricing page. Negotiated enterprise pricing, committed-spend discounts, and special programs (OpenAI startup credits, Anthropic enterprise volume, Google Vertex AI region rates) are out of scope — this is a list-price reference.

Cached input pricing is a third axis that has grown more important since OpenAI introduced automatic prompt caching in October 2024 and Anthropic shipped explicit prompt caching in mid-2024. For workloads with a long stable prefix (agents replaying the same system prompt; RAG against a fixed corpus), the cached input rate dominates the effective per-call cost. The page surfaces it as a column rather than the headline number because non-cached workloads still see the headline rate.

Batch endpoints typically halve both input and output rates (OpenAI Batch, Anthropic Message Batches, Google Batch API) in exchange for asynchronous turnaround up to 24 hours. The row note flags batch availability per model; the headline number is the synchronous-API rate.

Open-weights pricing is a hosting-provider proxy. Meta’s Llama, DeepSeek, Mistral’s Apache-2.0-licensed releases, and Alibaba’s open-weights Qwen sizes have no “official” per-token price — the model is free to download. The row records the Together AI or Groq published rate as a reference. Cheaper rates exist on self-hosted inference (vLLM, sglang) or on cost-optimized providers (Fireworks, DeepInfra); more expensive ones exist on the hyperscalers (AWS Bedrock, Azure AI Foundry).

Reasoning tokens count as output. The o-series, Anthropic’s extended-thinking models, and Gemini’s 2.5 Flash thinking budget all bill reasoning tokens at the output rate. A “cheap-looking” reasoning model can be expensive in practice because each call generates more reasoning than visible output. The page records the per-token rate; the effective per-call cost is workload-dependent.

Tokenizer differences cross-provider. $1 / M tokens does not buy the same amount of text from OpenAI’s o200k tokenizer as it does from Anthropic’s claude-3 tokenizer or Google’s SentencePiece. It does not even buy the same amount within one provider: Anthropic’s pricing page states that Claude 4.7 and later, plus the Mythos line, use a newer tokenizer producing approximately 30% more tokens for the same text, while Claude Sonnet 4.6 and earlier use the previous one — so a per-token rate that looks flat across those two groups is a real increase in what a page of text costs. Compare numbers in tokens, not in characters or pages, and treat the table as a same-provider apples-to-apples rather than a strict cross-provider one.

Historical pricing is sourced from Wayback Machine and launch posts. Every effective-date entry links to a primary source — provider blog post, archived pricing page, or release announcement. Where the Wayback snapshot is the only viable cite (the provider has rewritten the page since), the URL points there directly.

Qwen prices are the International (Singapore) list, not the Chinese-mainland one. Alibaba publishes two USD price tables for the same Qwen model — an International deployment scope and a Chinese-mainland one — and they differ by roughly 3–5×. Every current Qwen rate here is the International figure, because the rest of this table is US-hosted APIs and that is the comparable basis. Some pre-2026 Qwen points predate that basis and are quoted on the mainland table instead; those are kept as historical record and their Δ cell reads “rebased” rather than a percentage, because the gap between the two tables is a change of reference, not a price the provider moved.

About this page

Cross-family comparison page in the /ai/ section. Each row’s price is sourced from the provider’s own pricing page — OpenAI’s openai.com/api/pricing, Anthropic’s anthropic.com/pricing, Google’s ai.google.dev/gemini-api/docs/pricing, xAI’s docs.x.ai, DeepSeek’s api-docs.deepseek.com, Mistral’s docs.mistral.ai, and Alibaba’s per-model pages on Qwen Cloud. Meta’s Llama models are open-weights; this page cites Together AI’s published rate as the reference hosting-provider price.

The model roster mirrors the per-family pages already on this site — Claude, ChatGPT, Gemini, Grok, Llama, DeepSeek, Mistral, Qwen — so each row links back to the matching version-page entry for the full per-release context. The cross-family context-windows, release-cadence, and benchmarks pages cover the other comparison axes.

Refreshed daily as part of the cross-family /ai/ sweep. Each refresh re-verifies every active row against the provider’s current pricing page; values that changed since the previous run get a new entry appended to the model’s price history and the row’s effective date is bumped. Stale rows for deprecated models are kept as historical record — the price history is the chart, not just the current rate.

Last updated: September 1, 2026. 101 models · 8 providers · 127 price points.