2023 – 2026

Gemini Versions

Google's current Gemini frontier is Gemini 3.6 Flash, Stable GA since July 21, 2026 — the workhorse successor to 3.5 Flash, leading Google's line on coding, knowledge-work, and computer-use benchmarks at 17% fewer output tokens and a lower price. The Pro slot is held by Gemini 3.1 Pro (February 19, 2026 preview; the announced Gemini 3.5 Pro remains delayed) and the cost-efficient slot by Gemini 3.5 Flash-Lite (Stable GA July 21, 2026). I track every Gemini release here — from Bard's March 2023 launch and Gemini 1.0 in December 2023 onward — with API model strings, ship dates, and the major changes per version. Below the table: the Bard–to–Gemini rename, the Brain × DeepMind merger, the February 2024 image-generation incident, the DOJ search-monopoly remedies that explicitly cover the Gemini app, the $40B Anthropic investment, and the Project Astra / Live / Mariner agentic surfaces.

Family & status

Family

Pro — flagship reasoning models, including the original Gemini 1.0 Ultra tier
Flash — speed- and cost-optimized chat models, including Flash-Lite and Flash-8B
Specialized — on-device (Nano) and dedicated image / multimodal output models
Pre-Gemini — PaLM 2 and the original LaMDA-based Bard, before the December 2023 Gemini rebrand

Status

Current — actively recommended; the latest in its family
Available — still served via API but superseded
Legacy — deprecated or sunset; no longer the recommended surface

Gemini version table

Model
Gemini 3.6 Flash
gemini-3.6-flash
Flash
Current
Jul 21, 2026
Google's new workhorse Flash. 17% fewer output tokens than 3.5 Flash at a lower price ($1.50 / $7.50 per 1M). Leads the line on coding, knowledge-work, and computer-use benchmarks.
  • Released July 21, 2026 as the Stable GA gemini-3.6-flash, announced alongside Gemini 3.5 Flash-Lite and the limited-access Gemini 3.5 Flash Cyber; the announcement is at blog.google and the Stable listing is on the API models page.
  • 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index (up to 65% on DeepSWE), with fewer reasoning steps and tool calls per multi-step workflow — priced lower at $1.50 / $7.50 per million input / output tokens (3.5 Flash output was $9).
  • Vendor-stated gains over 3.5 Flash: DeepSWE 49% vs. 37%; MLE-Bench 63.9% vs. 49.7%; GDPval-AA v2 1421 vs. 1349; OSWorld-Verified 83.0% vs. 78.4%. Computer use is a built-in tool; the knowledge cutoff advances to March 2026. 1,048,576-token input context; 65,536-token output cap.
  • Now the default model powering the Antigravity agent; available day one in the Gemini app, AI Studio, Android Studio, Antigravity, and the Gemini Enterprise Agent Platform. Ships with strengthened CBRN / cyber-offense Frontier Safety safeguards; the model card is on the DeepMind site.
  • Starting with 3.6 Flash and 3.5 Flash-Lite, the API deprecates the temperature / top_p / top_k sampling parameters and rejects prefilled model turns, per the latest-models developer guide.
  • Gemini 3.5 Pro is still absent. Bloomberg reported July 16, 2026 that the rollout was delayed after the model fell short of internal performance goals; the launch post says it is “currently testing with partners” and will ship “as soon as it's ready.” The same post says Google has “started our most ambitious pre-training run yet, for Gemini 4.”
Model
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite
Flash
Available
Jul 21, 2026
Fastest, cheapest 3.5-family model — 350 output tokens/s, $0.30 / $2.50 per 1M. Replaces 3.1 Flash-Lite as the recommended high-throughput tier. Rolling out in Google Search.
  • Released July 21, 2026 as the Stable GA gemini-3.5-flash-lite, in the same announcement as Gemini 3.6 Flash; the post is at blog.google and the model page at ai.google.dev.
  • 350 output tokens per second (Artificial Analysis) at $0.30 / $2.50 per million input / output tokens — the fastest model in the 3.5 series, positioned for agentic search, document processing, and high-volume subagent execution.
  • Vendor-stated gains over 3.1 Flash-Lite: Terminal-Bench 2.1 54% vs. 31%; GDM-MRCR v2 long-context 72.2% vs. 60.1%; GDPval-AA v2 1140 vs. 642. It also beats Gemini 3 Flash on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%).
  • Computer use built in as a native tool; configurable thinking levels with a minimal default for maximum throughput. Available in the Gemini app, AI Studio, Android Studio, and the Gemini Enterprise Agent Platform, and rolling out in Google Search.
  • The model card is on the DeepMind site.
Model
Gemini 3.5 Flash Cyber
no public API id — limited-access pilot via CodeMender
Specialized
Available
Jul 21, 2026
Security-specialized 3.5 Flash fine-tune for finding, validating, and patching code vulnerabilities. Limited-access pilot for governments and trusted partners via CodeMender.
  • Announced July 21, 2026 alongside Gemini 3.6 Flash and 3.5 Flash-Lite; the announcement is at blog.google, with a DeepMind companion post at deepmind.google.
  • Built on Gemini 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models; competitive at the frontier on the CyberGym benchmark (vendor-stated) when run as multiple agents inside CodeMender.
  • CodeMender — DeepMind's code-security agent — orchestrates multiple 3.5 Flash Cyber agents into a single combined vulnerability report.
  • Access is deliberately gated. Google says the model will be “exclusively available to governments and trusted partners” via CodeMender as a limited-access pilot program, citing the dual-use nature of the capability. It has no public API model id and is absent from the API models page; this row records the announcement, not a generally available surface.
Model
Gemini 3.5 Live Translate
gemini-3.5-live-translate-preview
Specialized
Available
Jun 9, 2026
Low-latency audio-to-audio speech translation in 70+ languages. Shipping to Google Meet and the Google Translate app on Android / iOS.
  • Released June 9, 2026 in preview as gemini-3.5-live-translate-preview; the announcement is at blog.google. Documented at ai.google.dev.
  • Near real-time speech-to-speech translation with 70+ source and target languages, enabling 2,000+ bidirectional language combinations with high accuracy and natural voice output.
  • Uses the Gemini Live API for low-latency audio-to-audio streaming; designed for conversational translation use cases including video calls and in-person real-time translation.
  • Rolling out to Google Meet for spoken-conversation translation and to the Google Translate app (Android and iOS) for the Live translate feature; all audio output watermarked with SynthID.
  • A Specialized sibling to gemini-3.1-flash-live-preview and gemini-3.1-flash-tts-preview in the 3.x audio stack, purpose-built for the translation use case rather than general dialogue.
Model
Gemini 3.5 Flash
gemini-3.5-flash
Flash
Available
May 19, 2026
Frontier-class Flash that outperforms Gemini 3.1 Pro on hard coding and agentic benchmarks. Kicked off the 3.5 family at Google I/O. Superseded as the workhorse by Gemini 3.6 Flash on July 21, 2026.
  • Released May 19, 2026 at Google I/O as the generally available (GA) gemini-3.5-flash; the announcement is at blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5 and the Stable label is on the API models page.
  • Outperforms Gemini 3.1 Pro on hard coding and agentic benchmarks — the first Flash tier to beat the prior generation's Pro flagship at frontier evaluation suites.
  • Roughly 4× faster output tokens per second than other frontier models (vendor-stated); the default agentic coding model on Google surfaces until Gemini 3.6 Flash took over the Antigravity-agent default on July 21, 2026.
  • Powers Gemini Spark, Google's new always-on personal AI agent, and is the default behind the Gemini app and Search AI Mode.
  • Computer use became a built-in tool in 3.5 Flash on June 24, 2026 — the standalone Gemini 2.5 computer-use model folded natively into the main Flash model, letting developers build agents that see, reason, and act across browser, mobile, and desktop; the announcement is at blog.google.
  • Superseded as Google's workhorse Flash by gemini-3.6-flash on July 21, 2026; gemini-3.5-flash remains a served Stable id with no announced shutdown date per the deprecations page.
  • Gemini 3.5 Pro was announced at the same event but has not shipped. The I/O post said Google was “hard at work on 3.5 Pro,” that it was “already being used internally,” and that Google looked forward to “rolling it out next month” — i.e. June 2026. As of July 24, 2026 it carries no public model id and is absent from the API models page, the Vertex AI roster, and the DeepMind model list, so it gets no row here; Gemini 3.1 Pro remains the Pro tier (the delay is detailed in the Gemini 3.6 Flash row above). Gemini Omni Flash — a multimodal “any-input, any-output” model anchored on video — was announced alongside and reached developers as gemini-omni-flash-preview on June 30, 2026 (covered in the Imagen / Veo section below).
Model
Gemini 3.1 Flash TTS
gemini-3.1-flash-tts-preview
Specialized
Available
Apr 15, 2026
Cost-efficient, expressive, steerable text-to-speech. Adds new expressive audio tags for precise narration control over Gemini 2.5 Flash TTS.
  • Released April 15, 2026 in preview; the launch is recorded in the Gemini API changelog entry for that date.
  • Powerful, low-latency speech generation with natural outputs, steerable prompts, and new expressive audio tags for precise narration control — the 3.1-family successor to gemini-2.5-flash-preview-tts.
  • Documented alongside the broader speech-generation stack at ai.google.dev/gemini-api/docs/speech-generation; sits as a Specialized sibling to the Flash-Live audio-to-audio model.
  • Available via the Gemini API in Google AI Studio.
Model
Gemini 3.1 Flash Live
gemini-3.1-flash-live-preview
Specialized
Available
Mar 26, 2026
High-quality, low-latency audio-to-audio (A2A) model for real-time dialogue and voice-first AI apps. Backs the Live API on Gemini 3.1.
  • Released March 26, 2026 in preview as the latest audio-to-audio (A2A) model; the launch is recorded in the Gemini API changelog entry for that date.
  • Designed for real-time dialogue and voice-first AI applications — the 3.1-family successor to gemini-2.5-flash-native-audio-preview-12-2025, sitting under the Live API.
  • Documented alongside the broader Live API surface at ai.google.dev/gemini-api/docs/live-api; sits as a Specialized sibling to Flash TTS in the audio-output stack.
  • Available via the Gemini API in Google AI Studio.
Model
Gemini 3.1 Flash-Lite
gemini-3.1-flash-lite
Flash
Available
Mar 3, 2026
First Flash-Lite in the Gemini 3 series. Cost-efficient tier below 3 Flash for high-throughput / low-latency. GA on May 7, 2026.
  • Released March 3, 2026 in preview as gemini-3.1-flash-lite-preview; the announcement is at blog.google.
  • GA on May 7, 2026 as gemini-3.1-flash-lite; the preview snapshot gemini-3.1-flash-lite-preview was shut down on May 25, 2026 per the Gemini API changelog.
  • First Flash-Lite tier in the Gemini 3 series — the cost-efficient companion to Gemini 3 Flash, sitting below it on price/performance and replacing the 2.5 Flash-Lite lineage as the cheapest production tier.
  • Available via the Gemini API in Google AI Studio and via Vertex AI for enterprise customers.
  • The Vertex AI model card sits at docs.cloud.google.com.
Model
Gemini 3.1 Pro
gemini-3.1-pro-preview
Pro
Current
Feb 19, 2026
Reasoning + agentic refresh of Gemini 3 Pro. ARC-AGI-2 jumped from 31.1% to 77.1%. 1M input / 65K output. Current Google flagship.
  • Released February 19, 2026; the announcement is at blog.google. Replaced gemini-3-pro-preview, which was deprecated on March 9, 2026.
  • ARC-AGI-2 jumped from 31.1% to 77.1% over the 3 Pro baseline; agentic-benchmark gains in the 45–80% range across coding, browser, and tool-use evals (vendor-stated).
  • 1,048,576-token input context; 65,536-token output cap. Multimodal across text, audio, images, video, and full code repositories.
  • Available across the Gemini app (AI Pro and AI Ultra tiers), NotebookLM, Google AI Studio, the Antigravity agent platform, Vertex AI, Gemini Enterprise, the Gemini CLI, and Android Studio at launch.
  • The model card is published on the DeepMind site.
Model
Gemini 3 Flash
gemini-3-flash-preview
Flash
Available
Dec 17, 2025
Was the default model in the Gemini app and Search AI Mode until Gemini 3.5 Flash GA in May 2026. “Pro-grade reasoning at Flash speed.” $0.50 / $3 per 1M tokens.
  • Released December 17, 2025; the announcement is at blog.google/products/gemini/gemini-3-flash.
  • Replaced Gemini 2.5 Flash as the default in the Gemini app and in Search's AI Mode.
  • Pricing $0.50 / $3.00 per million input / output tokens at launch.
  • Available simultaneously across the Gemini API, AI Studio, Antigravity, the Gemini CLI, Android Studio, Vertex AI, and Gemini Enterprise.
  • Marketed as “Pro-grade reasoning at Flash-level speed”; the smaller, cheaper companion to Gemini 3 Pro and (later) 3.1 Pro.
Model
Gemini 3 Pro
gemini-3-pro-preview — shut down 2026-03-09
Pro
Legacy
Nov 18, 2025
First Google model launched simultaneously into Search, Gemini app, AI Studio, Vertex AI, and the Gemini CLI on day one. Top of LMArena at 1501 Elo at launch.
  • Released November 18, 2025; the announcement is at blog.google/products/gemini/gemini-3.
  • Day-one rollout across Search, Gemini app, AI Studio, Vertex AI, and the Gemini CLI — Google's first model launched simultaneously into Search.
  • Topped the LMArena Text Arena at 1501 Elo at launch; 37.5% on Humanity's Last Exam (no tools); 91.9% GPQA Diamond.
  • Introduced the Gemini 3 Deep Think mode, gated to AI Ultra subscribers and safety-tester cohorts. The mode generalized the parallel-thinking pattern from Gemini 2.5 Deep Think.
  • Superseded by Gemini 3.1 Pro on February 19, 2026; gemini-3-pro-preview was shut down on March 9, 2026 and the id now redirects to gemini-3.1-pro-preview per the Gemini API deprecations page.
Model
Gemini 2.5 Flash Image (“Nano Banana”)
gemini-2.5-flash-image
Specialized
Available
Aug 26, 2025
Image generation and natural-language editing inside the Gemini API. Character consistency, scene blending, local edits. SynthID watermarked.
  • Released August 26, 2025; the announcement is at developers.googleblog.com. Documented at ai.google.dev.
  • Character consistency, scene blending, and natural-language local edits — remove people, change pose, colorize, restyle — without the round-trip to a dedicated image-generation app.
  • Codenamed Nano Banana in pre-release testing; the codename stuck in much of the developer-community framing of the launch.
  • All output watermarked with SynthID; pricing approximately $30 per million output tokens (~$0.039 / image).
  • Distinct from the Imagen model line; covered briefly in the Imagen / Veo prose section below.
Model
Gemini 2.5 Deep Think
mode toggle within gemini-2.5-pro — AI Ultra only
Pro
Available
Aug 1, 2025
Parallel-thinking mode on top of Gemini 2.5 Pro. Generates and cross-evaluates many candidate solutions. AI Ultra subscribers only at $249.99/mo.
  • Released August 1, 2025; the announcement is at blog.google/products/gemini/gemini-2-5-deep-think. Model card at storage.googleapis.com (PDF).
  • Parallel thinking — the model generates many candidate ideas, considers them simultaneously, and revises or combines them before answering, rather than producing a single chain-of-thought trace.
  • Available only on the Google AI Ultra tier ($249.99/month) at launch, with a fixed daily prompt cap; first teased at Google I/O 2025.
  • State-of-the-art on LiveCodeBench V6 at launch; the same pattern was generalized into the gated Gemini 3 Deep Think mode on Gemini 3 Pro.
Model
Gemini 2.5 Flash-Lite
gemini-2.5-flash-lite
Flash
Available
Jul 22, 2025
Cheapest, fastest member of the 2.5 family. $0.10 / $0.40 per 1M tokens. 1M context. Routine recommendation for high-volume cheap-tier work.
  • Preview shipped June 17, 2025 alongside the 2.5 Pro / Flash GA; stable GA on July 22, 2025. The announcement is at developers.googleblog.com.
  • $0.10 / $0.40 per million input / output tokens at launch — the cheapest tier in the 2.5 family.
  • 1M-token context window; optimized for high-volume, low-latency workloads (classification, routing, summarization, agent steps).
  • Updated preview gemini-2.5-flash-lite-preview-09-2025 shipped September 25, 2025 with quality improvements; the stable id remained gemini-2.5-flash-lite.
Model
Gemini 2.5 Flash
gemini-2.5-flash
Flash
Available
Jun 17, 2025
The Flash sibling to 2.5 Pro. Default Gemini-app model from Google I/O 2025 until 3 Flash took over in December. Native audio output. Project Mariner computer use.
  • Preview unveiled at Google I/O 2025 on May 20, 2025; stable GA on June 17, 2025. The I/O announcement is at blog.google; thinking-mode updates at developers.googleblog.com.
  • Default Gemini-app model from I/O 2025 until Gemini 3 Flash took over the slot in December 2025.
  • 20–30% fewer tokens used in evals vs. the prior preview SKU; native audio output and “thought summaries” for enterprise auditability.
  • First Flash-tier model with Project Mariner computer-use capabilities folded in.
Model
Gemini 2.5 Pro
gemini-2.5-pro
Pro
Available
Mar 25, 2025
First Gemini family member to ship as a thinking model. Built-in chain-of-thought reasoning with a controllable thinking budget. 1M context.
  • Experimental released March 25, 2025; stable GA on June 17, 2025 after preview snapshots gemini-2.5-pro-preview-03-25, -05-06, and -06-05. The announcement is at blog.google.
  • First Gemini family member to ship reasoning as a default capability rather than a separate Thinking-Experimental SKU — built-in chain-of-thought, with a controllable thinkingBudget parameter.
  • 63.8% on SWE-Bench Verified at launch; 18.8% on Humanity's Last Exam (no tools).
  • 1M-token context window at launch; the 2M expansion that 1.5 Pro pioneered remained on Google's roadmap rather than shipping immediately.
  • The Deep Think mode on top of 2.5 Pro shipped four months later (August 1, 2025) on the AI Ultra tier; the page row above covers it.

The reasoning era — March 25, 2025. Above this line: the 2.5 and 3.x generations, where reasoning, multi-agent collaboration, native image generation, and agentic surfaces (Astra, Mariner, Antigravity) became first-class capabilities and the Gemini app moved to a one-launch-everywhere release pattern. Below: the 2.0 family's Thinking-Experimental separate-SKU era, and before that the 1.5 long-context era, the original 1.0 Ultra / Pro / Nano line, and the pre-Gemini PaLM 2 / Bard era.

Model
Gemini 2.0 Pro Experimental
gemini-2.0-pro-exp-02-05
Pro
Legacy
Feb 5, 2025
Largest context window of the era at 2M tokens. Top coding and world-knowledge model when shipped.
  • Released February 5, 2025 alongside 2.0 Flash GA and 2.0 Flash-Lite. The announcement is at blog.google.
  • 2,000,000-token context window — matched the 1.5 Pro 2M expansion of June 2024 and remained the largest Gemini context for several months.
  • Top of Google's coding and world-knowledge benchmarks for the 2.x family at launch; superseded by Gemini 2.5 Pro seven weeks later.
  • Shipped only as an Experimental SKU; never reached a stable gemini-2.0-pro id.
Model
Gemini 2.0 Flash + Flash-Lite
gemini-2.0-flash, gemini-2.0-flash-lite — shut down Jun 1, 2026
Flash
Legacy
Feb 5, 2025
2.0 Flash GA. 2.0 Flash-Lite as the cost-optimized sibling. Outperformed 1.5 Pro at twice the speed.
  • Both stable ids shipped February 5, 2025; the developer-side announcement is at developers.googleblog.com.
  • Outperformed Gemini 1.5 Pro on most benchmarks at twice the speed by Google's accounting at launch.
  • 1M-token context window across both SKUs.
  • 2.0 Flash-Lite was positioned as the cost-optimized variant for high-volume work; later superseded by the 2.5 Flash-Lite family in July 2025.
  • gemini-2.0-flash and gemini-2.0-flash-lite (including -001 aliases) were shut down June 1, 2026 per the Gemini API changelog; use gemini-3.5-flash or gemini-3.1-flash-lite instead.
Model
Gemini 2.0 Flash Thinking Experimental
gemini-2.0-flash-thinking-exp-1219
Flash
Legacy
Dec 19, 2024
Google's first reasoning model. Visible “thinking process” in the response. The OpenAI o1 response, three months after o1-preview.
  • Released December 19, 2024; competitive with OpenAI's o1 family that had launched September–December 2024.
  • Visible “thinking process” — the response surface exposed the model's chain-of-thought, the same pattern that Gemini 2.5 Pro later folded into the default Pro tier.
  • 32,767-token context at launch — substantially smaller than the 1M-token 2.0 Flash family it sat alongside.
  • Free in Google AI Studio at launch with a 2 RPM / 50 RPD limit; superseded by the integrated reasoning in Gemini 2.5 Pro three months later.
Model
Gemini 2.0 Flash (Experimental)
gemini-2.0-flash-exp
Flash
Legacy
Dec 11, 2024
Project Astra, Project Mariner, and Jules debut alongside. Native multimodal output (image, multilingual TTS) and Multimodal Live API.
  • Released December 11, 2024 in the “agentic era” framing post; the announcement is at blog.google.
  • Project Astra (universal AI agent prototype), Project Mariner (browser-controlling agent), and Jules (developer agent) all debuted alongside this release.
  • Native multimodal output: image generation and steerable, multilingual text-to-speech in a single API call.
  • Native tool use including Google Search and code execution; first model to expose the Multimodal Live API for real-time audio and video conversations.
  • Superseded by the GA gemini-2.0-flash on February 5, 2025.
Model
Gemini 1.5 Flash-8B
gemini-1.5-flash-8b
Flash
Legacy
Oct 3, 2024
Lowest cost-per-intelligence in the 1.5 family. 8B-parameter Flash sibling. 1M context.
  • Experimental released September 2024; stable GA on October 3, 2024. The announcement is at developers.googleblog.com.
  • Smallest Gemini SKU to ship a public API id at launch; positioned for the highest-volume cheap-tier agent steps.
  • 1M-token context window despite the smaller parameter count; superseded by the 2.0 / 2.5 Flash-Lite line.
Model
Gemini 1.5 — 2M context, “002” refresh
gemini-1.5-pro-002, gemini-1.5-flash-002
Pro
Legacy
Sep 24, 2024
2M context for 1.5 Pro (June 27, 2024). Stable “002” refresh of Pro and Flash with reduced 1.5 Pro pricing (September 24, 2024).
  • Two consecutive milestones rolled into one row. 2,000,000-token context for 1.5 Pro shipped June 27, 2024 alongside code execution; the announcement is at developers.googleblog.com.
  • “002” stable refresh of both 1.5 Pro and 1.5 Flash on September 24, 2024 with reduced 1.5 Pro pricing and increased rate limits; developer-blog post.
  • The 2M context expansion was the first time any production frontier model went beyond 1M; the milestone became the routine framing for “long-context” through 2024 and most of 2025.
Model
Gemini 1.5 Flash
gemini-1.5-flash, gemini-1.5-flash-001
Flash
Legacy
May 14, 2024
First Flash-tier release. 1M context. Multimodal text+image+audio+video input. Debuted at Google I/O 2024.
  • Released May 14, 2024 at Google I/O; the announcement is at blog.google.
  • First Flash-tier Gemini — the cheaper, faster sibling pattern that has been the routine recommendation for high-volume work in every Gemini generation since.
  • 1M-token context window; multimodal text + image + audio + video input on the same surface as 1.5 Pro.
  • Project Astra was first demoed at the same I/O; covered in the prose history section below.
Model
Gemini 1.5 Pro
gemini-1.5-pro, gemini-1.5-pro-001
Pro
Legacy
Feb 15, 2024
First Mixture-of-Experts in the family. 128K standard context, 1M in private preview. Equal or better to 1.0 Ultra at substantially less compute.
  • Announced February 15, 2024; the post is at blog.google. Public preview on Vertex AI followed April 9, 2024.
  • First Mixture-of-Experts model in the Gemini family — the architecture choice that made the 1M / 2M context expansions practical.
  • 128,000-token standard context at launch with 1,000,000-token context in private preview — the first production frontier model to ship a 1M-token surface.
  • Equal or better to Gemini 1.0 Ultra on most evaluations at substantially less inference compute, by Google's accounting.
Model
Gemini 1.0 Ultra (“Gemini Advanced”)
gemini-ultra — primarily product-only
Pro
Legacy
Feb 8, 2024
The largest and most capable Gemini 1.0 tier. Launch of Gemini Advanced subscription. Bard renamed to Gemini the same day.
  • Announced December 6, 2023 alongside the rest of the Gemini 1.0 family; general availability on February 8, 2024 via the new Gemini Advanced subscription on the $19.99/month Google One AI Premium plan. The announcement is at blog.google.
  • Reported 90.0% on MMLU at launch — Google's framing was that Ultra was the “first to outperform human experts” on the benchmark.
  • 32,768-token context window; primarily exposed through the consumer Gemini Advanced surface rather than as a routine API model.
  • Gemini Advanced was rebranded to Google AI Pro at I/O 2025 alongside the higher-priced AI Ultra tier ($249.99/month). The Bard–to–Gemini rename, also February 8, 2024, is covered in the prose history below.
Model
Gemini 1.0 Pro
gemini-pro, gemini-1.0-pro
Pro
Legacy
Dec 6, 2023
First public Gemini API model. Powered Bard from day one (170 countries, English). API GA December 13, 2023.
  • Announced December 6, 2023; the announcement is at blog.google/technology/ai/google-gemini-ai. Public API GA on December 13, 2023.
  • Powered Bard from day one in 170 countries (English at launch) — the model behind Bard for the two months between the December 6 launch and the February 8 Bard–to–Gemini rename.
  • 32,768-token context window; the first Gemini SKU developers could call from the public Gemini API.
  • Superseded by the Mixture-of-Experts 1.5 line two months later; deprecated through 2024.
Model
Gemini 1.0 Nano (and successors)
gemini-nano — on-device, AICore / ML Kit GenAI APIs
Specialized
Available
Dec 6, 2023
First Gemini on-device. Pixel 8 Pro at launch (Smart Reply, Summarize). Expanded across Pixel 8 / 8a / 9 / 10 and Chrome through 2024–2025.
  • Released December 6, 2023 on the Pixel 8 Pro Feature Drop — the first phone to run a frontier-lab model on-device. The post is at blog.google/products/pixel.
  • Initial features: Smart Reply in Gboard (WhatsApp at launch) and Summarize in Recorder; on Tensor G3 silicon.
  • Developer preview to Pixel 8 in April 2024; experimental access generally available across Android via AICore + ML Kit GenAI APIs in October 2024.
  • Expanded across Pixel 9 (multimodal Nano, August 2024), Pixel 8a, and Pixel 10 (August 2025) with successive on-device feature additions; later wired into Chrome's Built-in AI Prompt API.
  • This row consolidates the full on-device Nano lineage rather than splitting per-device-feature drops; treat it as the open-ended Specialized track within the Gemini family.

The Gemini era — December 6, 2023. Above this line: the natively-multimodal Gemini family, Google's response to ChatGPT and to the Brain × DeepMind merger of April 2023. Below: the Pre-Gemini era — PaLM 2 in May 2023 as Bard's underlying model, and the original LaMDA-based Bard launched in March 2023 as Google's reactive answer to ChatGPT.

Model
PaLM 2
chat-bison, text-bison, code-bison — deprecated
Pre-Gemini
Legacy
May 10, 2023
Replaced LaMDA as Bard's underlying model. Four sizes (Gecko / Otter / Bison / Unicorn). 100+ languages. Med-PaLM 2 and Sec-PaLM specialty variants.
  • Announced May 10, 2023 at Google I/O 2023; the Bard-related framing is at blog.google.
  • Replaced LaMDA as the model underlying Bard; supported 100+ natural languages and 20+ programming languages.
  • Four internal size codenames — Gecko, Otter, Bison, Unicorn — exposed via the PaLM API as chat-bison, text-bison, code-bison, etc.
  • Foundation for the Med-PaLM 2 medical and Sec-PaLM security specialty models.
  • Deprecated entirely once Gemini 1.0 shipped in December 2023; the PaLM API was sunset alongside.
Model
Bard (LaMDA-based)
never had a public API; product-only
Pre-Gemini
Legacy
Mar 21, 2023
Google's initial ChatGPT response. LaMDA-based at launch. Renamed to Gemini on February 8, 2024.
  • Released March 21, 2023 in early access to US and UK users; the post is at blog.google/innovation-and-ai/products/try-bard.
  • LaMDA-based at launch; framed as “a complement to Search,” not a replacement.
  • Followed Google's February 2023 ChatGPT-response announcement and the widely-reported Bard demo gaffe (the “James Webb Telescope first exoplanet image” error in the launch demo video).
  • Initial rollout blocked in the EU over GDPR concerns; came online there in July 2023.
  • Renamed Gemini on February 8, 2024 alongside the Gemini Advanced subscription launch — bard.google.com began redirecting to gemini.google.com that day.

Click any row to expand. Each row has a stable id for sharing — e.g. /ai/gemini/versions/#gemini-3-1-pro, #gemini-2-5-pro, #bard. The current model list is at ai.google.dev/gemini-api/docs/models; release notes at ai.google.dev/gemini-api/docs/changelog; Vertex AI roster at cloud.google.com/vertex-ai.

The Bard launch (March 2023) and the Bard–to–Gemini rename (February 2024)

Google's reactive answer to ChatGPT was announced on February 6, 2023 and shipped to early-access US and UK users on March 21, 2023 as Bard, an experimental conversational AI service backed by LaMDA. The launch demo's well-publicized error — the model misattributed the first exoplanet image to the James Webb Space Telescope — helped frame the early-2023 narrative that Google was caught flat-footed by OpenAI's November 2022 ChatGPT release. EU access was blocked initially over GDPR concerns and came online in July 2023.

At Google I/O on May 10, 2023, Bard's underlying model swapped from LaMDA to PaLM 2. Specialty variants — Med-PaLM 2 for medical question answering, Sec-PaLM for security — followed. Duet AI for Workspace shipped at $30/user/month on August 29, 2023 as the productivity-suite analog. The PaLM API was deprecated entirely once Gemini 1.0 shipped in December 2023.

The Bard–to–Gemini rename happened on February 8, 2024, the same day Gemini 1.0 Ultra reached general availability via the new Gemini Advanced subscription on the $19.99/month Google One AI Premium plan. bard.google.com began redirecting to gemini.google.com; Duet AI for Workspace was rebranded Gemini for Workspace; a new Android Gemini app shipped that day. The post is at blog.google. Gemini Advanced was itself rebranded Google AI Pro at I/O 2025 on May 20, 2025 alongside the higher-priced Google AI Ultra tier ($249.99/month).

The rebrand pattern continued on July 16, 2026, when Google renamed NotebookLM to Gemini Notebook — folding the research tool that launched as Project Tailwind at I/O 2023 into the Gemini brand. Google's post says it “remains a standalone product focused on being your premier research tool” while doing more across the Google ecosystem, and cites more than 30 million people and over 600,000 organizations using it. Shipping alongside the rename: a secure cloud computer per notebook that lets Gemini Notebook write and execute code natively against your sources (available at launch to Google AI Ultra users and Workspace customers with AI Ultra / AI Expanded Access, rolling out to all Pro users on the web), full cross-app syncing between the Gemini app and the standalone notebook experience, and a stated plan to bring notebooks into AI Mode in Search. The announcement, by Google Labs / Gemini app VP Josh Woodward, is at blog.google.

The Brain × DeepMind merger (April 2023)

On April 20, 2023, Sundar Pichai announced the merger of Google Brain and DeepMind into a single unit, Google DeepMind, with Demis Hassabis (DeepMind cofounder, then CEO) named CEO of the combined entity. The announcement is at blog.google/technology/ai/april-ai-update; DeepMind's side is at deepmind.google/blog/announcing-google-deepmind.

Jeff Dean, previously the Google Brain lead, was named Chief Scientist of Google + Alphabet, reporting directly to Pichai. The combined unit's accomplishments cited in the announcement: AlphaGo, the original Transformer paper, word2vec, WaveNet, AlphaFold, seq2seq, deep reinforcement learning, TensorFlow, and JAX. The merger is the corporate-side throughline that explains why Gemini was framed in December 2023 as “natively multimodal trained from the ground up” — Google's first model assembled by the unified post-merger team.

Demis Hassabis and DeepMind's John Jumper were jointly awarded the 2024 Nobel Prize in Chemistry with David Baker for AlphaFold's protein-structure-prediction work. The DeepMind announcement is at deepmind.google. Hassabis remains CEO of Google DeepMind through April 2026.

The February 2024 Gemini image-generation incident

Within days of the Bard–to–Gemini rename, users posted historically inaccurate images Gemini's image generator produced — racially diverse Nazi soldiers, multi-ethnic US Founding Fathers, Black female popes — that drew sharp criticism. Google paused people-image generation in Gemini on February 22, 2024. The explainer post by then-SVP for Search Prabhakar Raghavan said directly: “we got it wrong.”

On February 28, 2024 Pichai's internal memo — reported in NPR — called the outputs “completely unacceptable” and committed to structural changes, updated guidelines, and red-teaming. People-image generation was restored on August 28, 2024 with Imagen 3, gradually rolled out starting with Gemini Advanced / Business / Enterprise; TechCrunch coverage.

The episode is recorded here factually because it bears on the Gemini-1.0 / Imagen-2 rollout window and on Google's subsequent moderation framing for image generation. This page does not litigate the underlying decisions or their handling.

The Anthropic investments

Google has been one of Anthropic's two major cloud and capital partners alongside Amazon, which remains Anthropic's primary cloud provider. The October 27, 2023 $2 billion commitment ($500M upfront, $1.5B over time) was reported in CNBC, alongside Anthropic's $3B Google Cloud commitment over four years and adoption of TPU v5e for Claude inference.

On April 24, 2026, Google announced an additional commitment of up to $40 billion in Anthropic — $10B initial, with the remaining $30B contingent on performance milestones — at Anthropic's then-latest valuation of $380 billion. Coverage in TechCrunch, CNBC, and Bloomberg, which first reported the deal.

On the compute side, Anthropic announced on April 6, 2026 a new agreement with Google and chip partner Broadcom for multiple gigawatts of next-generation TPU capacity expected to come online starting in 2027 — 5 gigawatts over five years per CNBC — which Anthropic called “our most significant compute commitment to date” in its own announcement. That post also notes that Amazon remains Anthropic's primary cloud provider and training partner; Google is a capital partner and one of three platforms Claude runs on. A widely-circulated figure putting Anthropic's Google Cloud spend at $200 billion over five years originates with The Information (May 5, 2026) and was relayed by Reuters, which noted it could not verify the report and that neither company commented; neither Google nor Anthropic has confirmed it, so it is recorded here as reported, not confirmed. The scale of the bet surfaced in Alphabet's Q2 2026 results, reported July 22, 2026: roughly $99 billion in quarterly gains on equity investments — primarily the Anthropic and SpaceX stakes, after Anthropic's May 2026 raise at a $965 billion valuation — drove a record $112.1 billion net-income quarter, per Fortune; Bloomberg reported July 23 that Alphabet's SEC filing valued its private-company investments — primarily the Anthropic stake — at roughly $124.3 billion as of June 30, 2026. The Anthropic relationship is the corporate-strategy counterweight to the in-house Gemini line; for the Anthropic side, see Anthropic Leadership.

Project Astra, Gemini Live, and Project Mariner

Project Astra — the “universal AI agent” prototype with the smart-glasses demo — was first unveiled at Google I/O on May 14, 2024 as a research preview. The DeepMind page is at deepmind.google/models/project-astra. Astra's multimodal real-time understanding capabilities folded into Gemini Live (consumer) and into Search's AI Mode (May 2025).

Gemini Live, the conversational voice surface (10 voices, interruptible, runs in background), launched alongside the Pixel 9 series on August 13, 2024, initially as a Gemini Advanced subscriber feature. Camera and screen-sharing capabilities followed in 2025 as Astra's surface migrated into the Gemini app.

Project Mariner — the Gemini 2.0–powered research-prototype Chrome extension that controls the browser (cursor, clicks, forms) — was announced December 11, 2024 alongside the Gemini 2.0 Flash Experimental release. Single-tab and foreground-only at launch; 83.5% on the WebVoyager benchmark. Mariner's computer-use capability was folded into Gemini 2.5 Flash by I/O 2025 and into the Antigravity agent platform that launched alongside Gemini 3 Pro.

At Google I/O 2026 on May 19, 2026, Google launched Gemini Spark — a 24/7 personal AI agent that runs on Gemini 3.5 Flash via the Antigravity platform. Spark operates in the background and is designed to take actions on the user's behalf (scheduling, web tasks, custom sub-agents, authorized payments) while surfacing confirmations before major actions. The initial rollout is to trusted testers / AI Ultra subscribers in the US; the announcement is at blog.google. Gemini Spark is the consumer-facing instantiation of the agentic direction that Astra and Mariner prototyped in 2024–2025.

Imagen and Veo

The image and video lines under the Gemini umbrella ship on a parallel cadence to the chat models. Imagen 1 (May 2022) was a Google Brain research model; Imagen 2 (December 2023) powered Bard's image generation and was the model behind the February 2024 incident; Imagen 3 (announced May 14, 2024 at I/O; ImageFX rollout August 15, 2024; people-image return August 28, 2024) introduced the photorealistic-rendering jump. Imagen 4 shipped at I/O 2025 (May 20, 2025) in Standard, Fast, and Ultra variants with major text-rendering improvements and up to 2K resolution; Workspace Docs / Slides / Vids integration followed. The Imagen 4 line is now deprecated — imagen-4.0-generate-001, -fast-generate-001, and -ultra-generate-001 are scheduled to shut down August 17, 2026 per the June 15, 2026 Gemini API changelog entry, with the Nano Banana lineage (gemini-3-pro-image / gemini-2.5-flash-image) as the recommended replacement for image generation.

Nano Banana Pro (November 20, 2025; preview id gemini-3-pro-image-preview, stable GA id gemini-3-pro-image released May 28, 2026) followed Gemini 2.5 Flash Image as a professional design engine with a reasoning core — studio-quality 4K visuals, complex layouts, and precise text rendering, framed as the high-end successor to the original Nano Banana model. Nano Banana 2 (February 26, 2026; preview id gemini-3.1-flash-image-preview, stable GA id gemini-3.1-flash-image released May 28, 2026) shipped as the speed-and-volume sibling on the 3.1 base — high-efficiency production-scale visual creation, optimized for the rapid-iteration use cases the original Nano Banana shape was named for. Both preview ids are deprecated and shut down June 25, 2026 per the Gemini API deprecations page; the stable ids are documented at ai.google.dev/gemini-api/docs/image-generation. Nano Banana 2 Lite (gemini-3.1-flash-lite-image, stable GA released June 30, 2026) followed as the efficiency specialist of the family — ultra-low latency (text-to-image in ~4 seconds), $0.034 per 1K-resolution image, and 1K output only — and is now Google’s recommended replacement for the original Nano Banana (gemini-2.5-flash-image) on high-volume, cost-sensitive pipelines; the announcement is at blog.google. The Nano Banana shape now spans four models — the 2.5 Flash Image table row above plus Nano Banana Pro, Nano Banana 2, and Nano Banana 2 Lite in this prose section — covering the original codename, the professional tier, the balanced workhorse, and the speed/cost tier.

Veo 1 debuted as a Vertex research preview at I/O 2024. Veo 2 (December 16, 2024) brought 4K resolution and clips up to two minutes. Veo 3 at I/O 2025 was the first Veo with native synchronized audio — dialogue, sound effects, and ambient sound generated in the same pass — available initially through the AI Ultra tier and via the Google Flow filmmaking surface. Veo 3.1 Lite Preview (veo-3.1-lite-generate-preview, March 31, 2026) is the cost-efficient, developer-first addition to the Veo 3.1 family for rapid iteration and high-volume applications, per the Gemini API changelog. The earlier veo-2.0-generate-001, veo-3.0-generate-001, and veo-3.0-fast-generate-001 ids are deprecated and scheduled to shut down June 30, 2026 (June 15, 2026 changelog entry); the Veo 3.1 preview ids and the GA models on the Gemini Enterprise Agent Platform are the recommended replacements.

At Google I/O 2026 on May 19, 2026, Google introduced Gemini Omni Flash, a multimodal “any-input, any-output” video generation and editing model that accepts text, image, video, and audio inputs and generates high-quality video output. The I/O announcement is at blog.google. Omni Flash is designed for creative workflows — scene transformation, action reimagining, multi-turn conversational video refinement, and avatar-based video creation — and is built on Gemini’s world knowledge for physically plausible scene generation. On June 30, 2026, Google brought Omni Flash to developers for the first time as gemini-omni-flash-preview (public preview) via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform, alongside the Gemini app and Google Flow; it is priced at $0.10 per second of video output (matching Veo 3.1 Fast) and generates 10-second clips at launch, with SynthID watermarking on all output. The joint developer-availability post is at blog.google; the developer docs are at ai.google.dev/gemini-api/docs/omni. Image and video models are deliberately not enumerated as separate rows in the version table above (Gemini 2.5 Flash Image is the one row exception); if either line grows substantially, a dedicated /ai/gemini/image-models/ sibling page is in scope.

People who shaped Gemini

Sundar Pichai — CEO of Google since 2015, CEO of Alphabet since December 2019. The 2023 ChatGPT-response framing, the Brain × DeepMind merger, the December 2023 Gemini 1.0 launch, the February 2024 image-generation memo, and the strategy around the DOJ search remedies all run through Pichai's office.

Demis Hassabis — CEO of Google DeepMind since the April 2023 merger. DeepMind cofounder. 2024 Nobel laureate in Chemistry (jointly with then-DeepMind colleague John Jumper and David Baker) for AlphaFold. The technical-leadership face of the post-merger Gemini program.

Two senior departures landed in June 2026. Noam Shazeer — VP of engineering and a co-lead of the Gemini models, co-author of the 2017 “Attention Is All You Need” paper that introduced the Transformer, and cofounder of Character.AI — announced on June 17, 2026 that he was leaving for OpenAI, less than two years after Google brought him back in the reported $2.7B Character.AI licensing deal of August 2024; CNBC covered the move. Days later, John Jumper — the AlphaFold lead who shared the 2024 Nobel with Hassabis — announced he was leaving Google DeepMind for Anthropic, per Fortune (June 23, 2026). Neither has publicly stated a reason. Recorded here because both sat at the technical center of the Gemini and AlphaFold programs; this page does not handicap what the departures mean for the lab.

Jeff Dean — Chief Scientist, Google + Alphabet (since the April 2023 merger), reporting directly to Pichai. Previously co-led Google Brain. Sissie Hsiao led the Bard product through the rename. Eli Collins (VP, Google DeepMind product) co-launched Bard alongside Hsiao. Prabhakar Raghavan (then SVP, Search) authored the February 2024 image-generation explainer; he transitioned out of Search leadership in October 2024.

Current leadership beyond Pichai / Hassabis / Dean shifts more frequently and should be re-verified at each refresh run; the Musk v. Altman–style governance episode that has played out at OpenAI does not have a Google analog.

The competitive landscape

Gemini competes most directly with OpenAI's ChatGPT (the line that Bard was Google's reactive response to in 2023; see ChatGPT Versions), Anthropic's Claude (Google's primary capital and cloud partner alongside Amazon; see Claude Versions), xAI's Grok (see Grok Versions), and Meta's Llama. Gemini's distinguishing positioning has been the integration into Search, Workspace, Pixel devices, and Chrome alongside the standalone Gemini app, and the on-device Nano line that no other major frontier lab has matched at scale. This page does not attempt a benchmark roundup or a ranking.

Use Gemini

The browser cannot detect which Gemini model you've used or are using — there's no fingerprint or header that exposes it. The block below carries the practical information instead: the current model strings, a copy-paste API call, and the surfaces where Gemini is available.

Current model strings

Use these in the model field of an API request. Verify against ai.google.dev/gemini-api/docs/models for the freshest list.

# Latest Flash — Stable GA workhorse (Jul 21, 2026); 17% fewer output tokens than 3.5 Flash
gemini-3.6-flash

# Prior Flash flagship — Stable GA (May 19, 2026); still served
gemini-3.5-flash

# Flagship Pro (preview) — still the recommended Pro tier; 3.5 Pro delayed
gemini-3.1-pro-preview

# Cost-efficient Lite — Stable GA (Jul 21, 2026); fastest 3.5-family model
gemini-3.5-flash-lite

# Prior Lite — Stable GA (May 7, 2026); still served
gemini-3.1-flash-lite

# Flash 3 (preview) — still served; recommended replacement is gemini-3.6-flash
gemini-3-flash-preview

# Specialized 3.5 audio — speech-to-speech translation (70+ languages)
gemini-3.5-live-translate-preview

# Specialized 3.1 audio — real-time voice + steerable TTS
gemini-3.1-flash-live-preview
gemini-3.1-flash-tts-preview

# Specialized 3.x image — Nano Banana lineage (image generation / editing)
gemini-3-pro-image                  # Nano Banana Pro (stable GA May 28, 2026; preview shut down Jun 25)
gemini-3.1-flash-image              # Nano Banana 2 (stable GA May 28, 2026; preview shut down Jun 25)
gemini-3.1-flash-lite-image         # Nano Banana 2 Lite (stable GA Jun 30, 2026; fastest / cheapest)
gemini-2.5-flash-image              # original Nano Banana (legacy; upgrade to 2 Lite)

# Specialized 3.x video — Omni Flash (video gen + conversational editing)
gemini-omni-flash-preview           # public preview (Jun 30, 2026); Interactions API

# 2.5 family — stable GA, still routinely served
gemini-2.5-pro
gemini-2.5-flash
gemini-2.5-flash-lite

# On-device
gemini-nano   # on-device via AICore / ML Kit GenAI APIs

Quick API call

Drop in your GEMINI_API_KEY and run. The generateContent endpoint is the canonical entry point.

$ curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.6-flash:generateContent" \
    -H "x-goog-api-key: $GEMINI_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "contents": [{
        "parts": [{ "text": "Hello, Gemini." }]
      }]
    }'

Where to access Gemini

Multiple surfaces, same models underneath. Pick whichever fits the task.

# Standalone consumer app — free + Google AI Pro / Ultra
https://gemini.google.com/

# Developer surfaces
https://ai.google.dev/                         # Gemini API + AI Studio
https://aistudio.google.com/
https://cloud.google.com/vertex-ai/             # Vertex AI (enterprise)

# Built-in surfaces
Search (AI Mode + AI Overviews)
Workspace (Gmail, Docs, Sheets, Slides, Meet)   # formerly Duet AI
Pixel + Android (on-device Gemini Nano)
Chrome (Built-in AI Prompt API powered by Nano)
Gemini Notebook (research tool)                 # renamed from NotebookLM Jul 16, 2026

# Native apps
Gemini for iOS, Android

Model lifecycle

Google publishes a changelog alongside each new model id; preview ids are eventually replaced with stable equivalents. Pin a dated id only when you need bit-for-bit reproducibility; otherwise prefer the un-dated alias so you migrate forward automatically.

# Changelog and lifecycle
https://ai.google.dev/gemini-api/docs/changelog

# Stable alias — rolls forward to the latest snapshot
"model": "gemini-2.5-pro"

# Pinned snapshot — freezes the exact training cut
"model": "gemini-2.5-pro-preview-06-05"

Sources: Gemini API model docs; Gemini API changelog; Vertex AI model docs; per-release announcement posts at blog.google and developers.googleblog.com; US v. Google (D.D.C. 1:20-cv-03010) and the ad-tech case (E.D. Va. 1:23-cv-00108) on CourtListener; contemporaneous reporting in NYT, WSJ, Bloomberg, The Information, CNBC, NPR, NBC News, TechCrunch, and 9to5Google. Last updated July 2026.

Mungomash LLC · More AI pages