AI model lineage
How the major AI models are actually related
The major AI model lineages span 151 models across 8 per-provider family trees, connected by 146 typed edges — successions, post-trainings, fine-tunes, distillations, retrainings, and the GPT line / o-series merge into GPT-5, as of September 4, 2026. Each per-provider tree below cites the provider documentation that establishes each edge.
As of September 4, 2026.
Current production flagships
One node per provider, color-coded. Click any flagship to jump into the per-provider family tree below.
How to read the trees
Each box is a model. Box fill indicates lifecycle status (Current emerald, Available sky, Legacy amber, Deprecated rose). Each box has a left-edge stripe in the provider color. Each arrow is a typed edge:
Hover any node or edge for its full label, ship date, and (for edges) the source the relationship was sourced to. The "Lineage as text" block under each tree contains the same information for screen readers and crawlers.
Anthropic · Claude
Versions page →Lineage as text (24 edges) ↓
Every edge in the Claude tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- Claude 2 (Jul 11, 2023) — succession of Claude 1
- Claude Instant 1.2 (Aug 9, 2023) — succession of Claude 1 — Anthropic has never documented Claude Instant as a distillation of the Claude flagship — only as the faster, lower-cost model in the lineup. The Claude 2 announcement (anthropic.com/news/claude-2, July 11, 2023) does not mention Claude Instant anywhere, so nothing stronger than chronological succession is disclosed for this step.
- Claude 2.1 (Nov 21, 2023) — succession of Claude 2 — Anthropic discloses no base-model relationship between Claude 2 and Claude 2.1. The launch post (anthropic.com/news/claude-2-1, November 21, 2023) says only that 'Claude 2.1 delivers advancements in key capabilities for enterprises—including an industry-leading 200K token context window, significant reductions in rates of model hallucination, system prompts and our new beta feature: tool use.' That is a capability list, not a statement about whether 2.1 shares Claude 2's base, so the edge is typed succession.
- Claude 3 Opus (Mar 4, 2024) — new base of Claude 2.1 — Anthropic's launch post is headlined 'Introducing the next generation of Claude' and announces 'the Claude 3 model family … three state-of-the-art models in ascending order of capability: Claude 3 Haiku, Claude 3 Sonnet, and Claude 3 Opus' — a new generation of base models rather than a Claude 2 fine-tune. Per anthropic.com/news/claude-3-family (March 4, 2024).
- Claude 3 Sonnet (Mar 4, 2024) — succession of Claude 3 Opus — Anthropic shipped Opus / Sonnet / Haiku as the three tiers of the Claude 3 generation simultaneously.
- Claude 3 Haiku (Mar 13, 2024) — succession of Claude 3 Opus
- Claude 3.5 Sonnet (Jun 20, 2024) — succession of Claude 3 Sonnet
- 3.5 Sonnet (new) (Oct 22, 2024) — post-training of Claude 3.5 Sonnet — Anthropic announces 'an upgraded Claude 3.5 Sonnet, and a new model, Claude 3.5 Haiku' — the October 2024 release keeps the Claude 3.5 Sonnet name and is presented as an upgrade of it, not a new model. Per anthropic.com/news/3-5-models-and-computer-use (October 22, 2024).
- Claude 3.5 Haiku (Nov 4, 2024) — succession of 3.5 Sonnet (new) — Anthropic makes no distillation, teacher-student, or shared-base claim between the two 3.5-generation models. The launch announcement (anthropic.com/news/3-5-models-and-computer-use, October 22, 2024) introduces Claude 3.5 Haiku as 'a new model' that 'matches the performance of Claude 3 Opus' — a capability comparison against an older flagship, not a lineage statement about 3.5 Sonnet.
- Claude 3.7 Sonnet (Feb 24, 2025) — succession of 3.5 Sonnet (new)
- Claude Sonnet 4 (May 22, 2025) — new base of Claude 3.7 Sonnet — Anthropic: 'Today, we're introducing the next generation of Claude models: Claude Opus 4 and Claude Sonnet 4' — a new generation rather than fine-tunes of 3.7 Sonnet. Per anthropic.com/news/claude-4 (May 22, 2025).
- Claude Opus 4 (May 22, 2025) — new base of Claude 3.7 Sonnet
- Claude Opus 4.1 (Aug 5, 2025) — succession of Claude Opus 4
- Claude Sonnet 4.5 (Sep 29, 2025) — succession of Claude Sonnet 4
- Claude Haiku 4.5 (Oct 15, 2025) — succession of Claude 3.5 Haiku
- Claude Opus 4.5 (Nov 24, 2025) — succession of Claude Opus 4.1
- Claude Opus 4.6 (Feb 5, 2026) — succession of Claude Opus 4.5
- Claude Sonnet 4.6 (Feb 17, 2026) — succession of Claude Sonnet 4.5
- Claude Opus 4.7 (Apr 16, 2026) — succession of Claude Opus 4.6
- Claude Opus 4.8 (May 28, 2026) — succession of Claude Opus 4.7
- Claude Fable 5 (Jun 9, 2026) — succession of Claude Opus 4.8 — Anthropic introduced Claude Fable 5 (June 9, 2026) as the first generally-available model in the new Mythos class — a tier positioned above Opus in capability. Claude Mythos 5 (claude-mythos-5) is disclosed as the same underlying model as Fable 5 with safeguards lifted in some areas — restricted to Project Glasswing cyberdefenders and trusted-access partners, not a public GA model, so it is not drawn as a separate node. Opus 4.8 is described as the 'next-most-capable model' that Fable 5 falls back to on safeguard-triggered queries; Anthropic does not disclose Fable 5 as a fine-tune, post-training, or distillation of Opus 4.8. The announcement page also carries the release's two access notices — 'Claude Mythos 5 and Fable 5 access unavailable' (June 12, 2026: 'We are suspending access to Claude Fable 5 and Claude Mythos 5') and 'Claude Mythos 5 and Fable 5 redeployed' (July 1, 2026) — without stating a cause; the Claude Versions page carries the fuller account. Per anthropic.com/news/claude-fable-5-mythos-5 (June 9, 2026).
- Claude Sonnet 5 (Jun 30, 2026) — succession of Claude Sonnet 4.6
- Claude Opus 5 (Jul 24, 2026) — succession of Claude Opus 4.8 — Claude Opus 5 (July 24, 2026) is the new head of the Opus tier, described by Anthropic as coming 'close to the frontier intelligence of Claude Fable 5 at half the price' and as providing 'greatly improved performance for the same cost as its predecessor, Opus 4.8.' It is the new default model on Claude Max and the strongest model on Claude Pro, and it remains behind Mythos 5 on cybersecurity tasks. Anthropic does not disclose Opus 5 as a fine-tune, post-training, or distillation of Opus 4.8 — per the anthropic.com/news/claude-opus-5 announcement.
- Claude Fable 5.1 (Sep 1, 2026) — succession of Claude Fable 5 — Claude Fable 5.1 (September 1, 2026) is Anthropic's next Mythos-class flagship, generally available on every platform on launch day with no preview stage. The announcement frames the pair exactly as it framed Fable 5 and Mythos 5: 'Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs.' Because Anthropic calls Mythos 5.1 'identical to Fable 5.1' — one set of weights, two safeguard configurations, reached through the Cyber Verification Program and its life-sciences counterpart — it is not drawn as a separate node, the same call this tree made for Mythos 5. Anthropic discloses no base-model relationship to Fable 5 in either the launch post or the 212-page system card — no fine-tune, post-training or distillation claim — so this edge is a succession rather than the post-training step the Gemini Flash track carries. Fable 5 stays served at the same $10 / $50 per M; the release's price move is a 75% cut to cache reads, $1 to $0.25 per M — per anthropic.com/claude-fable-and-mythos-5-1 (September 1, 2026).
OpenAI · ChatGPT
Versions page →Lineage as text (21 edges) ↓
Every edge in the ChatGPT tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- GPT-3.5 (Nov 30, 2022) — post-training of GPT-3 — OpenAI's InstructGPT paper, the line GPT-3.5 descends from, states the base relationship outright: 'we collect a dataset of labeler demonstrations of the desired model behavior, which we use to fine-tune GPT-3 using supervised learning. We then collect a dataset of rankings of model outputs, which we use to further fine-tune this supervised model using reinforcement learning from human feedback.' Per arxiv.org/abs/2203.02155.
- GPT-4 (Mar 14, 2023) — new base of GPT-3.5 — OpenAI's GPT-4 Technical Report presents GPT-4 as its own pretrained model rather than a GPT-3.5 fine-tune: 'We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.' Per arxiv.org/abs/2303.08774.
- GPT-4 Turbo (Nov 6, 2023) — succession of GPT-4 — OpenAI never documented GPT-4 Turbo as sharing GPT-4's base. Its model-catalog entry says only that 'GPT-4 Turbo is the next generation of GPT-4, an older high-intelligence GPT model. It was designed to be a cheaper, better version of GPT-4' — product positioning, and the same page moves the knowledge cutoff to December 1, 2023 against GPT-4's September 2021, which implies fresh pre-training data rather than a post-training pass. Per developers.openai.com/api/docs/models/gpt-4-turbo; the November 2023 DevDay announcement at openai.com/index/ is bot-blocked from automated fetching and could not be re-read.
- GPT-4o (May 13, 2024) — new base of GPT-4 Turbo — The GPT-4o system card describes a single new natively-multimodal network rather than a pass over GPT-4 Turbo: GPT-4o is 'an autoregressive omni model, which accepts as input any combination of text, audio, image, and video', and 'It's trained end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network.' Per cdn.openai.com/gpt-4o-system-card.pdf.
- GPT-4o mini (Jul 18, 2024) — succession of GPT-4o — OpenAI makes no distillation or teacher-student claim between GPT-4o and GPT-4o mini. The launch post (openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence, July 18, 2024) introduces it as 'our most cost-efficient small model' and benchmarks it against Gemini Flash and Claude Haiku rather than against GPT-4o. The two things it does say the models share are a tokenizer — 'Thanks to the improved tokenizer shared with GPT-4o' — and safety work: 'GPT-4o mini has the same safety mitigations built-in as GPT-4o.' Neither is a base-model or teacher-student claim, so the edge is typed succession.
- o1-preview (Sep 12, 2024) — new base of GPT-4o — The OpenAI o1 system card (September 12, 2024) opens 'The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought', calls o1 'this new model family', and states that 'the two models were pre-trained on diverse datasets, including a mix of publicly available data, proprietary data accessed through partnerships, and custom datasets developed in-house' — their own pre-training run, with no claim of continuity from GPT-4o anywhere in the document. Per cdn.openai.com/o1-system-card.pdf.
- o1 (Dec 5, 2024) — succession of o1-preview
- o3-mini (Jan 31, 2025) — succession of o1
- GPT-4.1 (Apr 14, 2025) — succession of GPT-4o
- o4-mini (Apr 16, 2025) — succession of o3 — The joint launch post (openai.com/index/introducing-o3-and-o4-mini, April 16, 2025) says only that 'OpenAI o4-mini is a smaller model optimized for fast, cost-efficient reasoning.' Smaller-and-cheaper describes the product, not the training relationship, so no distillation claim attaches and the edge is typed succession.
- o3 (Apr 16, 2025) — succession of o3-mini
- GPT-5 (Aug 7, 2025) — lines merged of GPT-4.1 — The GPT-5 system card describes the merge in its first line: 'GPT-5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool need and explicit intent.' The chat half of that system is the GPT line this edge follows; the card's 'Model progressions' table maps 'GPT-4o → gpt-5-main' and 'GPT-4.1-nano → gpt-5-thinking-nano'. Per cdn.openai.com/gpt-5-system-card.pdf.
- GPT-5 (Aug 7, 2025) — lines merged of o3 — The GPT-5 system card maps the o-series into GPT-5 model-for-model, listing 'OpenAI o3 → gpt-5-thinking', 'OpenAI o4-mini → gpt-5-thinking-mini' and 'OpenAI o3 Pro → gpt-5-thinking-pro' alongside the GPT-line rows, and notes it 'focuses primarily on gpt-5-thinking and gpt-5-main'. o3 is the reasoning half of the unified system. Per cdn.openai.com/gpt-5-system-card.pdf.
- GPT-5.1 (Nov 12, 2025) — succession of GPT-5
- GPT-5.2 (Dec 11, 2025) — succession of GPT-5.1
- GPT-5.3-Codex (Feb 5, 2026) — succession of GPT-5.2
- GPT-5.3 Instant (Mar 3, 2026) — post-training of GPT-5.2 — OpenAI frames GPT-5.3 Instant as a behaviour update on the same model rather than a new one: 'Today, we're releasing an update to ChatGPT's most-used model that makes everyday conversations more consistently helpful and fluid,' and 'This update focuses on the parts of the ChatGPT experience people feel every day: tone, relevance, and conversational flow.' The improvements it lists are all post-training ones — fewer unnecessary refusals, less moralising preamble, better web synthesis, lower hallucination rates — measured throughout against GPT-5.2 Instant, which 'will remain available for three months for paid users in the model picker under the Legacy Models section, after which it will be retired on June 3, 2026.' Shipped as gpt-5.3-chat-latest, separate from the Codex-specialized GPT-5.3-Codex two days earlier. Per openai.com/index/gpt-5-3-instant (March 3, 2026).
- GPT-5.4 (Mar 5, 2026) — succession of GPT-5.3-Codex
- GPT-5.5 (Apr 23, 2026) — succession of GPT-5.4
- GPT-5.6 Sol (Jul 9, 2026) — succession of GPT-5.5 — GPT-5.6 (Sol / Terra / Luna) previewed June 26, 2026 and reached general availability July 9, 2026 across ChatGPT, Codex, and the OpenAI API, replacing GPT-5.5 as OpenAI's current frontier flagship. It introduces a new naming system: the number denotes the generation, while Sol (flagship), Terra (balanced), and Luna (fast, low-cost) are durable capability tiers. OpenAI's model catalog at developers.openai.com/api/docs/models lists gpt-5.6-sol as the 'Flagship model for complex professional work' (aliased gpt-5.6, 1.05M context, 128K max output, Feb 16, 2026 knowledge cutoff, $4 / $20 per M) alongside gpt-5.6-terra and gpt-5.6-luna, which share that context window and cutoff. OpenAI does not disclose GPT-5.6 as a fine-tune, post-training, or distillation of GPT-5.5 — per the openai.com/index/gpt-5-6/ GA announcement (preview: openai.com/index/previewing-gpt-5-6-sol/).
- GPT-6 Astra (Sep 3, 2026) — succession of GPT-5.6 Sol — GPT-6 Astra opens the GPT-6 generation on September 3, 2026, and across every OpenAI surface that documents it, GPT-5.6 Sol appears only as a benchmark baseline — never as a base model. The model card, the dated changelog entry, the safety overview and the full system card all compare the two without claiming continuity between them, so the edge is typed succession. The circumstantial case for a new base is strong and is recorded here so it is not re-derived: OpenAI describes Astra as 'the culmination of several long-running alignment workstreams (ranging from pre-training interventions to more careful and consistent grading during reinforcement learning)', describes pausing and then restarting 'the large frontier RL run' for Astra specifically, and moves the knowledge cutoff to April 30, 2026 — the first advance since GPT-5.6's February 16, 2026. What is missing is the plain new-base sentence this page requires before typing an edge that way. On availability: Astra is on the tree because it shipped, not because it is generally available. OpenAI's card says 'GPT-6 Astra is rolling out today for enterprises in our Trusted Access Program, with access through API and our Plus, Pro, Business and Enterprise plans coming in the coming days,' the catalog still calls gpt-5.6-sol the 'Flagship model for complex professional work' while its own 'Choosing a model' guidance has already moved to 'use GPT-6 Astra, our flagship model for complex reasoning and coding,' and Astra is the first model OpenAI has designated Critical for cybersecurity under its Preparedness Framework. 1.05M context, 128K max output, $10 / $50 per M. Per developers.openai.com/api/docs/models/gpt-6-astra, the September 3, 2026 changelog entry, openai.com/index/safety-overview-gpt-6-astra and openai.com/index/path-to-astra.
Google · Gemini
Versions page →Lineage as text (19 edges) ↓
Every edge in the Gemini tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- PaLM 2 (May 10, 2023) — new base of Bard (LaMDA) — The PaLM 2 Technical Report documents a fresh pre-training run: 'We introduce PaLM 2, a new state-of-the-art language model that has better multilingual and reasoning capabilities and is more compute-efficient than its predecessor PaLM. PaLM 2 is a Transformer-based model trained using a mixture of objectives.' Per arxiv.org/abs/2305.10403. Google's May 10, 2023 I/O post is what ties it to Bard — 'PaLM 2's improved multilingual capabilities are allowing us to expand Bard to new languages' — per blog.google/technology/ai/google-palm-2-ai-large-language-model. Note the report names PaLM, not the LaMDA model behind the first Bard, as the predecessor; this edge is Google's consumer-assistant line changing base, not a claim that PaLM 2 descends from LaMDA.
- Gemini 1.0 Pro (Dec 6, 2023) — new base of PaLM 2 — Google describes Gemini 1.0 as pre-trained from scratch rather than adapted from PaLM 2: 'We designed Gemini to be natively multimodal, pre-trained from the start on different modalities. Then we fine-tuned it with additional multimodal data to further refine its effectiveness,' and separately that it 'was built from the ground up to be multimodal.' Per blog.google/technology/ai/google-gemini-ai (December 6, 2023). The technical report opens the same way — 'This report introduces a new family of multimodal models, Gemini' — at arxiv.org/abs/2312.11805.
- Gemini 1.0 Ultra (Feb 8, 2024) — succession of Gemini 1.0 Pro
- Gemini 1.5 Pro (Feb 15, 2024) — new base of Gemini 1.0 Ultra — Google names the architecture change outright: Gemini 1.5 is 'more efficient to train and serve, with a new Mixture-of-Experts (MoE) architecture,' shipped alongside a context window 'running up to 1 million tokens consistently, achieving the longest context window of any large-scale foundation model yet.' Per blog.google/technology/ai/google-gemini-next-generation-model-february-2024 (February 15, 2024). The technical report calls the family 'the next generation of highly compute-efficient multimodal models' — arxiv.org/abs/2403.05530. A new architecture and a new pre-training run, not a pass over Gemini 1.0.
- Gemini 1.5 Flash (May 14, 2024) — distilled from of Gemini 1.5 Pro — Google states the teacher-student relationship outright: 1.5 Flash is fast and efficient 'because it's been trained by 1.5 Pro through a process called “distillation,” where the most essential knowledge and skills from a larger model are transferred to a smaller, more efficient model' — per the May 14, 2024 I/O announcement at blog.google/technology/ai/google-gemini-update-flash-ai-assistant-io-2024. This is the only disclosure of its kind anywhere in the Gemini line; every other small-sibling step in the family is a succession because Google never repeats the claim.
- Gemini 2.0 Flash (Dec 11, 2024) — succession of Gemini 1.5 Flash
- Gemini 2.0 Pro Exp (Feb 5, 2025) — succession of Gemini 1.5 Pro
- Gemini 2.5 Pro (Mar 25, 2025) — succession of Gemini 2.0 Pro Exp
- Gemini 2.5 Flash (Jun 17, 2025) — succession of Gemini 2.0 Flash
- 2.5 Flash-Lite (Jul 22, 2025) — succession of Gemini 2.5 Flash — Google uses no distillation or teacher-student language anywhere in the 2.5 Flash-Lite launch. The preview announcement (blog.google/products-and-platforms/products/gemini/gemini-2-5-model-family-expands, June 17, 2025) calls it 'our most cost-efficient and fastest 2.5 model yet,' and the GA post (developers.googleblog.com/en/gemini-25-flash-lite-is-now-stable-and-generally-available, July 22, 2025) says it was built 'to push the frontier of intelligence per dollar.' Both are positioning statements, so the edge is typed succession. The Gemini 2.5 model cards have since been retired, leaving these two posts as the only surviving sources for this step.
- Gemini 3 Pro (Nov 18, 2025) — new base of Gemini 2.5 Pro — Google states the new-base relationship as a negative, which is the strongest form of it: the Gemini 3 Pro model card's Model dependencies row reads 'Gemini 3 Pro is not a modification or a fine-tune of a prior model. Each subsequent model in the Gemini 3 Pro family is based on Gemini 3 Pro.' Per deepmind.google/models/model-cards/gemini-3-pro/, which is served as a PDF rather than a web page.
- Gemini 3 Flash (Dec 17, 2025) — post-training of Gemini 3 Pro — The Flash tier branches off Pro here rather than continuing from Gemini 2.5 Flash, because Google says so. The Gemini 3 Flash model card (deepmind.google/models/model-cards/gemini-3-flash/) states 'Gemini 3 Flash is based on Gemini 3 Pro' under Model dependencies, Architecture and Training Dataset alike, and its Description reads 'Gemini 3 Flash is built off of the Gemini 3 Pro reasoning foundation with thinking levels to control the mix of quality, cost and latency.' The Gemini 3 Pro card corroborates from the other side, naming Gemini 3 Flash in the list of models that 'is based on Gemini 3 Pro'. Same base, no new pre-training run. Note this is a shared-base disclosure, not a teacher-student one: Google uses no distillation language anywhere in the Gemini 3 Flash track.
- Gemini 3.1 Pro (Feb 19, 2026) — post-training of Gemini 3 Pro — The Gemini 3.1 Pro model card (deepmind.google/models/model-cards/gemini-3-1-pro/) states under Model dependencies: 'Gemini 3.1 Pro is based on Gemini 3 Pro.' Same base, no new pre-training run — so this is a post-training step, not a new generation. The launch blog never says this; the card is the disclosure.
- 3.1 Flash-Lite (Mar 3, 2026) — post-training of Gemini 3 Pro — The Gemini 3.1 Flash-Lite card (deepmind.google/models/model-cards/gemini-3-1-flash-lite/) states under Model dependencies: 'Gemini 3.1 Flash-Lite is based on Gemini 3 Pro' — the Pro model, not the Flash model it sits under in the tier stack, which is why this edge is drawn across lanes. The launch blog never says this; only the card does. Note this is a shared-base disclosure, not a teacher-student one: Google uses distillation language nowhere in the 3.x Flash-Lite track.
- Gemini 3.5 Flash (May 19, 2026) — post-training of Gemini 3 Flash — The Gemini 3.5 Flash card (deepmind.google/models/model-cards/gemini-3-5-flash/) repeats 'Gemini 3.5 Flash is based on Gemini 3 Flash' under Model dependencies, Architecture, Training Dataset and Hardware, and its Description reads 'Gemini 3.5 Flash is the next iteration in the Gemini 3 series … based on the Gemini 3 Flash reasoning foundation with thinking levels to control the mix of quality, cost and latency.' Same base, no new pre-training run, so this is a post-training step. Google's I/O framing of 3.5 Flash — the first Flash to outperform the prior generation's Pro flagship — is benchmark positioning and carries no lineage claim either way; the card is what settles it.
- 3.5 Flash-Lite (Jul 21, 2026) — post-training of 3.1 Flash-Lite — The Gemini 3.5 Flash-Lite card (deepmind.google/models/model-cards/gemini-3-5-flash-lite/) states under Model dependencies: 'Gemini 3.5 Flash-Lite is based on Gemini 3.1 Flash-Lite' — the Lite track continues from itself, not from the Flash model above it in the tier stack. The July 21, 2026 launch blog frames it only as 'our fastest, most cost-effective 3.5-class model' and discloses no distillation from 3.5 Flash.
- Gemini 3.6 Flash (Jul 21, 2026) — post-training of Gemini 3.5 Flash — The Gemini 3.6 Flash card (deepmind.google/models/model-cards/gemini-3-6-flash/) states under Model dependencies: 'Gemini 3.6 Flash is based on Gemini 3.5 Flash' — same base, no new pre-training run. The July 21, 2026 launch blog offers only feedback-and-benchmarks framing ('builds directly on developer and customer feedback from 3.5 Flash'), which would not by itself support anything stronger than succession; the card is the disclosure the blog never makes.
- Gemini 3.7 Flash (Aug 13, 2026) — post-training of Gemini 3.6 Flash — Gemini 3.7 Flash reached Stable GA August 13, 2026 as Google's workhorse Flash, three weeks after 3.6 Flash. Unlike every prior Flash-to-Flash step, Google discloses the base-model relationship outright: the 3.7 Flash model card (deepmind.google/models/model-cards/gemini-3-7-flash/, published 13 August 2026) states 'Gemini 3.7 Flash is based on Gemini 3.6 Flash' under Model dependencies, Architecture, Training Dataset, Hardware and Software alike, and describes the release as 'the next iteration in the Gemini 3 model family, featuring algorithmic improvements to its core reasoning foundation.' The launch post calls it 'a direct result of developer feedback and algorithmic innovations' shipped 'just three weeks after Gemini 3.6 Flash' — per blog.google 'Introducing Gemini 3.7 Flash' (August 13, 2026). Same base, no new pre-training run, so this edge is a post-training update rather than the succession the rest of the Flash track carries.
- Gemini 3.8 Flash (Sep 2, 2026) — post-training of Gemini 3.7 Flash — Gemini 3.8 Flash reached Stable GA September 2, 2026 — three weeks after 3.7 Flash, and Google's third Flash release in six weeks. The base-model disclosure carries forward verbatim: the 3.8 Flash model card (deepmind.google/models/model-cards/gemini-3-8-flash/, published 2 September 2026) states 'Gemini 3.8 Flash is based on Gemini 3.7 Flash' under Model dependencies, Architecture, Training Dataset, Hardware and Software alike, and its description reads 'the next iteration in the Gemini 3 model family, building on Gemini 3.7 Flash, delivering performance advancements across software engineering and agentic knowledge workflows.' Same base, no new pre-training run, so this is a post-training update — the second consecutive Flash-to-Flash step Google has documented that way. A gated sibling, Gemini 3.8 Flash Cyber, shipped the same day through the Fairwind Program with no public API id, so it is not drawn as a node.
xAI · Grok
Versions page →Lineage as text (14 edges) ↓
Every edge in the Grok tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- Grok 1.5 (Mar 28, 2024) — succession of Grok 1 — xAI states no base-model relationship between Grok-1 and Grok-1.5. The announcement (x.ai/news/grok-1.5, March 28, 2024) says 'we have improved reasoning and problem-solving capabilities in our latest model, Grok-1.5' and describes a custom JAX / Rust / Kubernetes training stack that lets the team 'train new architectures at scale' — capability and infrastructure prose that leaves the training relationship undisclosed, so the edge is typed succession.
- Grok 2 (Aug 13, 2024) — succession of Grok 1.5 — The Grok-2 Beta announcement (x.ai/news/grok-2, August 13, 2024) uses the words 'train', 'pretrain', 'architecture' and 'from scratch' nowhere in the post. It says only that Grok-2 is 'a significant step forward from our previous model Grok-1.5, featuring frontier capabilities in chat, coding, and reasoning' — positioning language rather than a new-base disclosure, so the edge is typed succession.
- Grok 3 (Feb 17, 2025) — new base of Grok 2 — xAI: Grok 3 blends 'strong reasoning with extensive pretraining knowledge' and was 'Trained on our Colossus supercluster with 10x the compute of previous state-of-the-art models' — an explicit new pretraining run, not a pass over Grok 2. Per x.ai/news/grok-3, which is bylined February 19, 2025 — two days after the February 17 launch the Versions row records.
- Grok 4 (Jul 9, 2025) — post-training of Grok 3 — The Grok 4 announcement (x.ai/news/grok-4, July 9, 2025) claims no new pretraining run — it credits the previous generation with the pretraining and describes Grok 4's own work as reinforcement learning on top of it. It says 'With Grok 3, we scaled next-token prediction pretraining to unprecedented levels' and then 'For Grok 4, we utilized Colossus, our 200,000 GPU cluster, to run reinforcement learning training that refines Grok's reasoning abilities at pretraining scale.' Grok 4 is the RL scale-up on the Grok 3 pretraining, plus RL training for native tool use.
- grok-code-fast-1 (Aug 28, 2025) — new base of Grok 4 — xAI documented grok-code-fast-1 (August 28, 2025) as a from-scratch architecture optimized for agentic coding rather than a fine-tune of Grok 4 — per the x.ai/news/grok-code-fast-1 announcement.
- Grok 4 Fast (Sep 19, 2025) — succession of Grok 4 — The announcement (x.ai/news/grok-4-fast, September 19, 2025) says Grok 4 Fast is 'Built on xAI's learnings from Grok 4' — learnings, not weights — and describes 'a unified architecture that blends reasoning and non-reasoning modes in one model' with a 2M-token context. No teacher-student or shared-base claim is made, so the edge is typed succession.
- Grok 4.1 (Nov 17, 2025) — post-training of Grok 4 — xAI describes Grok 4.1 as a post-training pass over the Grok 4 line: 'we used the same large scale reinforcement learning infrastructure that powered Grok 4 and applied it to optimize the style, personality, helpfulness, and alignment of the model,' and separately 'In Grok 4.1 post-training, we focus on reducing factual hallucinations' — per the x.ai/news/grok-4-1 announcement (November 17, 2025).
- Grok 4.1 Fast (Nov 19, 2025) — succession of Grok 4.1 — The launch post (x.ai/news/grok-4-1-fast, November 19, 2025) introduces Grok 4.1 Fast as 'our best tool-calling model with a 2M context window' alongside the Agent Tools API, and makes no smaller-sibling, distillation, or shared-base claim against Grok 4.1, so the edge is typed succession.
- Grok 4.20 (Mar 10, 2026) — succession of Grok 4.1
- Grok 4.3 (Apr 17, 2026) — succession of Grok 4.20
- Grok Build 0.1 (May 19, 2026) — succession of grok-code-fast-1 — Grok Build 0.1 (early-access launch May 19, 2026; the Grok Build CLI / TUI launched in beta five days earlier on May 14) takes over the agentic-coding slot grok-code-fast-1 held, but xAI never states a training relationship between the two. The API announcement (x.ai/news/grok-build-0-1, May 29, 2026) describes grok-build-0.1 only as 'a coding model specifically trained for agentic coding tasks, including web development, debugging, and MCP support' and 'the same model that powers Grok Build'; it does not mention grok-code-fast-1 at all. The edge is therefore chronological succession within xAI's coding track.
- Composer 2.5 (Jun 1, 2026) — succession of Grok Build 0.1 — xAI shipped Composer 2.5 (June 1, 2026) inside the Grok Build /model menu as a fast agentic-coding sibling to Grok Build 0.1; the original launch coverage identified Composer 2.5 as built on the open-source Kimi K2.5 checkpoint (Moonshot AI) and post-trained with roughly 25× more synthetic agentic tasks than Composer 2, so the line of descent is external — xAI does not disclose Composer 2.5 as a fine-tune, post-training, or distillation of any prior xAI model. The x.ai/news/composer-2-5 page has since been pared back to a two-paragraph availability note ('Composer 2.5 is a fast, state-of-the-art model that excels on long-running tasks and following complex instructions') that names neither Kimi K2.5 nor Composer 2's training; the Kimi K2.5 attribution is preserved on /ai/grok/versions/#grok-composer-2-5 with the original launch citation.
- Grok 4.5 (Jul 8, 2026) — succession of Grok 4.3 — Grok 4.5 (July 8, 2026) is xAI's newest chat and coding flagship, collapsing the Grok 4.x line's separate reasoning and non-reasoning SKUs into a single grok-4.5 model id with a low / medium / high reasoning-effort setting. xAI describes it as its 'strongest model ever,' trained across tens of thousands of NVIDIA GB300 GPUs on new coding / science / engineering / math data; it does not disclose Grok 4.5 as a fine-tune, post-training, or distillation of Grok 4.3 — per the x.ai/news/grok-4-5 announcement.
- Grok 4.6 (Aug 12, 2026) — post-training of Grok 4.5 — xAI documents Grok 4.6 (August 12, 2026) as continuous with Grok 4.5 rather than a new base: 'Grok 4.6 builds on Grok 4.5' and, under Training Grok 4.6, 'Grok 4.6 underwent a longer supplemental training run than Grok 4.5 … This produced a stronger foundation for the SFT and RL stages that followed. We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains.' So Grok 4.5 is both the starting point and the teacher for 4.6's supervised fine-tuning data. Per x.ai/news/grok-4-6 (August 12, 2026); the model page at docs.x.ai/developers/grok-4-6 gives a 500,000-token context window, a January 2026 knowledge cutoff, $2.00 / $6.00 per M tokens, and low / medium / high / xhigh reasoning effort.
Meta · Llama
Versions page →Lineage as text (16 edges) ↓
Every edge in the Llama tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- Llama 2 (Jul 18, 2023) — new base of LLaMA 1 — The Llama 2 paper documents its own pre-training run rather than an adaptation of LLaMA 1: 'In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama 2-Chat, are optimized for dialogue use cases.' Per arxiv.org/abs/2307.09288. The same release moved the weights from LLaMA 1's research-only licence to the Llama Community License.
- Code Llama (Aug 24, 2023) — fine-tune of Llama 2 — Meta's Code Llama paper opens: 'We release Code Llama, a family of large language models for code based on Llama 2 providing state-of-the-art performance among open models, infilling capabilities, support for large input contexts, and zero-shot instruction following ability for programming tasks.' Per arxiv.org/abs/2308.12950.
- Code Llama 70B (Jan 29, 2024) — succession of Code Llama — Meta documents no fine-tune of the earlier Code Llama releases into the 70B. Its Code Llama card (github.com/meta-llama/codellama) covers 7B / 13B / 34B / 70B together, says the released models 'have been trained and fine-tuned using the same data as Llama 2 with different weights', and dates the whole family's training 'between January 2023 and January 2024' — a shared origin in Llama 2 rather than a chain through the earlier Code Llama sizes. Chronological succession within the Code Llama line is all that is disclosed. Meta's own Code Llama 70B announcement on ai.meta.com is unreachable (HTTP 400), so the card is the surviving source.
- Llama 3 (Apr 18, 2024) — new base of Llama 2 — The Llama 3 model card documents a fresh pre-training corpus and a new tokenizer rather than continuity with Llama 2: 'Llama 3 is an auto-regressive language model that uses an optimized transformer architecture. Llama 3 uses a tokenizer with a vocabulary of 128K tokens, and was trained on sequences of 8,192 tokens,' over 'a new mix of publicly available online data' totalling 15T+ tokens against Llama 2's 2T. Per github.com/meta-llama/llama-models/blob/main/models/llama3/MODEL_CARD.md — the un-gated mirror, since the meta-llama HuggingFace repos require access approval.
- Llama 3.1 (Jul 23, 2024) — succession of Llama 3 — Meta discloses no base-model relationship between Llama 3 and Llama 3.1. The 3.1 card describes 'a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes' whose 'tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF)' — a description of 3.1's own pipeline, not a statement that it starts from Llama 3 weights — and it names Llama 3 8B only in the benchmark comparison table. Per github.com/meta-llama/llama-models/blob/main/models/llama3_1/MODEL_CARD.md, so the edge is typed succession.
- Llama Guard 3 (Jul 23, 2024) — fine-tune of Llama 3.1 — Meta's Llama Guard 3-8B card opens: 'Llama Guard 3-8B is a Llama-3.1-8B pretrained model, fine-tuned for content safety classification.' Per github.com/meta-llama/PurpleLlama/blob/main/Llama-Guard3/8B/MODEL_CARD.md — the un-gated mirror of the card, since the meta-llama HuggingFace repos require access approval.
- Llama 3.2 (Sep 25, 2024) — distilled from of Llama 3.1 — This is the rare teacher-student disclosure Meta actually makes, and it names the teacher checkpoints: 'For the 1B and 3B Llama 3.2 models, we incorporated logits from the Llama 3.1 8B and 70B models into the pretraining stage of the model development, where outputs (logits) from these larger models were used as token-level targets. Knowledge distillation was used after pruning to recover performance.' Per github.com/meta-llama/llama-models/blob/main/models/llama3_2/MODEL_CARD.md. Read the scope: the claim covers the 1B and 3B sizes only, and the same card records that Llama 3.2 overall 'was pretrained on up to 9 trillion tokens of data from publicly available sources' — which is why this edge is not typed post-training.
- Llama 3.3 70B (Dec 6, 2024) — post-training of Llama 3.2 — The Llama 3.3 70B model card reports the same pretraining corpus scale and knowledge cutoff as the rest of the Llama 3.x line (15T+ tokens, December 2023) and describes the release as 'a pretrained and instruction tuned generative model' whose 'tuned versions use supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF)'. Meta claims no new pretraining run for 3.3 — per github.com/meta-llama/llama-models/blob/main/models/llama3_3/MODEL_CARD.md.
- Llama 4 Scout (Apr 5, 2025) — new base of Llama 3.3 70B — The Llama 4 model card documents a new architecture rather than a pass over the Llama 3 line: 'The Llama 4 collection of models are natively multimodal AI models that enable text and multimodal experiences. These models leverage a mixture-of-experts architecture to offer industry-leading performance in text and image understanding,' and, under Model Architecture, 'The Llama 4 models are auto-regressive language models that use a mixture-of-experts (MoE) architecture and incorporate early fusion for native multimodality.' Per github.com/meta-llama/llama-models/blob/main/models/llama4/MODEL_CARD.md — the un-gated mirror of the card.
- Llama 4 Maverick (Apr 5, 2025) — new base of Llama 3.3 70B
- Llama Guard 4 (Apr 29, 2025) — fine-tune of Llama 4 Scout — Meta's Llama Guard 4 card names the base outright, and it is Scout rather than the Maverick sibling beside it: 'Llama Guard 4 is a natively multimodal safety classifier with 12 billion parameters trained jointly on text and multiple images. It is a dense architecture pruned from the Llama 4 Scout pre-trained model and fine-tuned for content safety classification.' No Meta artifact draws Guard 4 from Maverick. Per github.com/meta-llama/PurpleLlama/blob/main/Llama-Guard4/12B/MODEL_CARD.md — the un-gated mirror of the card, since the meta-llama HuggingFace repos require access approval.
- Muse Spark (Apr 8, 2026) — new base of Llama 4 Maverick — Meta names Llama 4 Maverick as the model Muse Spark replaces, and describes a rebuilt pre-training stack rather than a continuation of it: 'Over the last nine months, we rebuilt our pre-training stack with improvements to model architecture, optimization, and data curation … we can reach the same capabilities with over an order of magnitude less compute than our previous model, Llama 4 Maverick.' The launch post also frames Muse Spark as a fresh start — 'the first in the Muse family of models developed by Meta Superintelligence Labs' and 'the first product of a ground-up overhaul of our AI efforts' — and it is closed-weights: 'Muse Spark is available today at meta.ai and the Meta AI app. We're opening a private API preview to select users.' Per ai.meta.com/blog/introducing-muse-spark-msl (April 8, 2026).
- Muse Spark 1.1 (Jul 9, 2026) — succession of Muse Spark — Muse Spark 1.1 (July 9, 2026) is 'the latest model from Meta Superintelligence Labs and a significant upgrade from Muse Spark' — 'a multimodal reasoning model built for agentic tasks, with major gains in tool and computer use, coding, and multimodal understanding.' It launched in 'Thinking' mode in the Meta AI app alongside a public preview of the Meta Model API. Meta does not disclose it as a fine-tune, post-training, or distillation of Muse Spark, so the edge is typed succession. Per research.meta.ai/blog/introducing-muse-spark-meta-model-api, where Meta's own research blog index files the post; the 1,048,576-token context window is from Meta's model docs at dev.meta.ai/docs/models rather than the announcement.
- Muse Spark 1.2 (Aug 5, 2026) — succession of Muse Spark 1.1 — Muse Spark 1.2 (August 5, 2026) shipped alongside Muse Code, 'a terminal coding agent powered by Muse Spark 1.2, our newest model.' Under a 'Self-Improvement' heading Meta says 'We also used Muse Spark 1.1 to generate challenging coding environments and instruction-following templates. The model then graded candidate solutions on how well they satisfied those requirements' — the previous release supplied training data, which is not a statement that 1.2 shares its weights or base. Meta discloses no fine-tune, post-training, or distillation relationship, so the edge is typed succession. Per research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2.
- Muse Glimmer 30B (Aug 10, 2026) — distilled from of Muse Spark 1.2 — Meta Superintelligence Lab's model card opens: 'Muse Glimmer is a 30-billion-parameter causal language model with a dedicated perception encoder, distilled from Muse Spark and purpose-built for autonomous agentic tasks on consumer hardware.' Apache 2.0, 131K context, ~29.6B parameters including the vision encoder, knowledge cutoff January 4, 2026 — per huggingface.co/meta-models/Muse-Glimmer-30B. The card names the Muse Spark line generically rather than naming the teacher checkpoint — its only specific reference is a safety comparison, that Glimmer 'is broadly weaker than Muse Spark 1.0' — so the edge is drawn from Muse Spark 1.2 as the Spark release current when Glimmer shipped. Note the org: the weights are hosted under meta-models, not meta-llama.
- Muse Spark 1.3 (Sep 2, 2026) — succession of Muse Spark 1.2 — Muse Spark 1.3 (September 2, 2026) 'delivers improved performance across agentic and coding tasks', and Meta says 'We trained the model across a diverse set of harnesses to generalize to various agentic environments.' Meta's model docs describe the line as 'three versions, each sharing the same modalities and context window, and differing only by capability' — a statement about what the releases can do, not about shared weights or a shared base, so the edge is typed succession. Two caveats sit behind that. The announcement holds the top reasoning mode back — 'Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing' — although Meta's research-blog index now bills the release as shipping 'with max reasoning'. And the model docs, not the announcement, record that audio understanding 'in Muse Spark 1.3 is currently not fully supported, and response quality for requests including audio content may be degraded'; Meta routes audio work back to Muse Spark 1.2 or to Muse Voice Transcribe, so the 'same modalities' claim does not fully hold. Per research.meta.ai/blog/introducing-muse-spark-1-3 and dev.meta.ai/docs/models.
DeepSeek · DeepSeek
Versions page →Lineage as text (12 edges) ↓
Every edge in the DeepSeek tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- DeepSeek-Coder (Nov 2, 2023) — new base of DeepSeek-LLM — DeepSeek Coder is not a code fine-tune of DeepSeek-LLM, despite sitting next to it in the timeline. The official README opens: 'DeepSeek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese', and the feature list repeats 'Trained from scratch on 2T tokens'. Coder is an independently pretrained line — and it shipped first, November 2 against DeepSeek-LLM's November 29, 2023. Per github.com/deepseek-ai/DeepSeek-Coder.
- DeepSeek-V2 (May 6, 2024) — new base of DeepSeek-LLM — The DeepSeek-V2 paper presents a new architecture rather than an update of the dense DeepSeek-LLM line: 'We present DeepSeek-V2, a strong Mixture-of-Experts (MoE) language model characterized by economical training and efficient inference. It comprises 236B total parameters, of which 21B are activated for each token, and supports a context length of 128K tokens. DeepSeek-V2 adopts innovative architectures including Multi-head Latent Attention (MLA) and DeepSeekMoE.' Per arxiv.org/abs/2405.04434.
- DeepSeek-V3 (Dec 26, 2024) — new base of DeepSeek-V2 — The DeepSeek-V3 paper documents its own pre-training run at a new scale: 'We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2.' The architecture carries over from V2; the model does not — DeepSeek reports V3 pre-trained on 14.8T tokens of its own. Per arxiv.org/abs/2412.19437.
- DeepSeek-R1 (Jan 20, 2025) — post-training of DeepSeek-V3 — The DeepSeek-R1 paper names the base in its body: 'Specifically, we build upon DeepSeek-V3-Base and employ Group Relative Policy Optimization (GRPO) as our RL framework', and for R1 itself, 'we construct and collect a small amount of long CoT data to fine-tune the model as the initial RL actor.' Read the paper body, not the listing page: the arXiv abstract under this id has been replaced with the journal version, which describes the RL framework in general terms and never names DeepSeek-V3-Base. Per the full text at arxiv.org/html/2501.12948v2.
- DeepSeek-V3-0324 (Mar 24, 2025) — post-training of DeepSeek-V3 — DeepSeek presents the 0324 build as an update to the same model, not a new one: 'DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects,' and the card then lists benchmark deltas against V3 measured point-for-point (MMLU-Pro 75.9 to 81.2, GPQA 59.1 to 68.4, AIME 39.6 to 59.4, LiveCodeBench 39.2 to 49.2), plus front-end and Chinese-writing gains. The same release relicensed the weights to MIT. Per huggingface.co/deepseek-ai/DeepSeek-V3-0324.
- DeepSeek-R1-0528 (May 28, 2025) — post-training of DeepSeek-R1 — DeepSeek's card: 'The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training.' Per huggingface.co/deepseek-ai/DeepSeek-R1-0528. DeepSeek does not state that the two share an architecture, so the disclosure rests on that quoted post-training language rather than on any parameter-count continuity.
- DeepSeek-V3.1 (Aug 21, 2025) — lines merged of DeepSeek-R1-0528 — DeepSeek's release note folds the standalone reasoning line into the flagship: 'Introducing DeepSeek-V3.1: our first step toward the agent era!' with 'Hybrid inference: Think and Non-Think — one model, two modes' and 'Faster thinking: DeepSeek-V3.1-Think reaches answers in less time vs. DeepSeek-R1-0528.' The API change makes the merge concrete — 'deepseek-chat → non-thinking mode, deepseek-reasoner → thinking mode', both on one model. Per api-docs.deepseek.com/news/news250821 (August 21, 2025). No standalone R-series model has shipped since R1-0528.
- DeepSeek-V3.1 (Aug 21, 2025) — post-training of DeepSeek-V3-0324 — DeepSeek-V3.1 (August 21, 2025) merged the V-series and R-series into a hybrid Thinking / Non-Thinking architecture — per the api-docs.deepseek.com/news/news250821 release note (the generic api-docs.deepseek.com thinking-mode guide was subsequently rewritten around V4-pro).
- DeepSeek-V3.2-Exp (Sep 29, 2025) — new base of DeepSeek-V3.1 — DeepSeek-V3.2-Exp (September 29, 2025) is described by DeepSeek as 'an experimental version of our model' and 'an intermediate step toward our next-generation architecture' that 'builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention' — per the github.com/deepseek-ai/DeepSeek-V3.2-Exp README.
- DeepSeek-V3.2 (Dec 1, 2025) — post-training of DeepSeek-V3.2-Exp — The DeepSeek-V3.2 card declares the relationship in its own front matter — 'base_model: deepseek-ai/DeepSeek-V3.2-Exp-Base' with 'base_model_relation: finetune' — and its introduction credits DSA plus a scalable RL framework 'scaling post-training compute'. Per huggingface.co/deepseek-ai/DeepSeek-V3.2 (arXiv 2512.02556 is the technical report, but its abstract never names V3.2-Exp; the card's front matter is where the dependency is stated).
- DeepSeek-V4-Pro (Apr 24, 2026) — new base of DeepSeek-V3.2 — DeepSeek: 'We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens', built on new architecture (hybrid Compressed Sparse Attention + Heavily Compressed Attention, Manifold-Constrained Hyper-Connections, the Muon optimizer) and benchmarked against V3.2 on inference FLOPs and KV cache. A new base, not a pass over V3.2. Per huggingface.co/deepseek-ai/DeepSeek-V4-Pro. The tree dates this node to the April 24, 2026 preview; DeepSeek's own news index lists a separate 'DeepSeek-V4-Pro GA Release' on August 13, 2026, and the model card still opens on the preview wording.
- DeepSeek-V4-Flash (Apr 24, 2026) — succession of DeepSeek-V4-Pro — The V4-Flash model card claims no teacher-student relationship — it says the model 'outperforms DeepSeek-V4-Pro (Preview) on benchmarks listed below despite its far smaller activated parameter count', which is the opposite of a student trailing its teacher, so the smaller sibling here is a succession rather than a distillation. Its introduction reads 'DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities.' Per huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731.
Mistral · Mistral
Versions page →Lineage as text (18 edges) ↓
Every edge in the Mistral tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- Mixtral 8x7B (Dec 11, 2023) — succession of Mistral 7B — The Mixtral paper says the opposite of a new base — it says the architecture is the same one: 'We introduce Mixtral 8x7B, a Sparse Mixture of Experts (SMoE) language model. Mixtral has the same architecture as Mistral 7B, with the difference that each layer is composed of 8 feedforward blocks (i.e. experts).' It never claims Mixtral inherits Mistral 7B's weights either, and it never says the model was trained from scratch, so no base-model relationship is disclosed in either direction and the edge is typed succession. Per arxiv.org/abs/2401.04088.
- Mistral Large (Feb 26, 2024) — succession of Mixtral 8x7B — Mistral's 'Au Large' announcement makes no training claim at all: 'We are releasing Mistral Large, our latest and most advanced language model … Mistral Large is our new cutting-edge text generation model. It reaches top-tier reasoning capabilities.' New-flagship framing is not a new-base disclosure, and the post never mentions Mixtral, so the edge is typed succession. What the release did change is the licence — Mistral Large shipped closed, on la Plateforme and Azure, against Mixtral's Apache 2.0. Per mistral.ai/news/mistral-large (February 26, 2024).
- Mistral Large 2 (Jul 24, 2024) — new base of Mistral Large — Mistral's 'Large Enough' announcement (mistral.ai/news/mistral-large-2407, July 24, 2024) presents Mistral Large 2 as its own pretrained 123B base rather than an update of Mistral Large: it reports multilingual MMLU 'measured on the base pretrained model', says 'following our experience with Codestral 22B and Codestral Mamba, we trained Mistral Large 2 on a very large proportion of code', and benchmarks it against 'the previous Mistral Large'. Per mistral.ai/news/mistral-large-2407.
- Codestral 25.01 (Jan 13, 2025) — succession of Mistral Large 2 — Mistral frames Codestral 25.01 as an upgrade of the earlier Codestral, not of anything on the flagship line: 'today, Codestral is getting a big upgrade. Codestral 25.01 features a more efficient architecture and an improved tokenizer than the original, generating and completing code about 2 times faster.' The post never mentions Mistral Large 2, the release this edge follows chronologically, and the only model it benchmarks itself against inside the family is Codestral-2405 22B, which has no node on this tree. So nothing stronger than chronological succession is disclosed. Per mistral.ai/news/codestral-2501 (January 13, 2025).
- Mistral Small 3 (Jan 30, 2025) — new base of Mistral Large 2 — Mistral: 'Today we're introducing Mistral Small 3, a latency-optimized 24B-parameter model released under the Apache 2.0 license' — a new 24B base and the return to Apache 2.0, not a pass over the 123B Mistral Large 2. Per mistral.ai/news/mistral-small-3, bylined January 30, 2025 — a distinct, earlier event from the December 2025 Mistral 3 family relaunch that shipped Mistral Large 3 and Ministral 3.
- Mistral Medium 3 (May 7, 2025) — succession of Mistral Small 3 — Mistral Medium 3 (May 7, 2025) is the mid-tier proprietary flagship — per mistral.ai/news/mistral-medium-3.
- Devstral Small (May 21, 2025) — fine-tune of Mistral Small 3 — Devstral Small (May 21, 2025) is a 24B coding-agent fine-tune of Mistral Small 3.1. The model card declares base_model mistralai/Mistral-Small-3.1-24B-Instruct-2503, says outright 'It is finetuned from Mistral-Small-3.1, therefore it has a long context window of up to 128k tokens', and adds 'As a coding agent, Devstral is text-only and before fine-tuning from Mistral-Small-3.1 the vision encoder was removed.' Per the HuggingFace card at mistralai/Devstral-Small-2505 — the mistral.ai/news/devstral launch post names Mistral Small 3.1 only in a pricing comparison, so the card is where the relationship is stated.
- Magistral (Jun 10, 2025) — fine-tune of Mistral Small 3 — Magistral Small (June 10, 2025) declares base_model mistralai/Mistral-Small-3.1-24B-Instruct-2503 and is described as 'Building upon Mistral Small 3.1 (2503), with added reasoning capabilities, undergoing SFT from Magistral Medium traces and RL on top' — per the HuggingFace card at mistralai/Magistral-Small-2506. The launch post adds that 'Magistral is fine-tuned for multi-step logic' (mistral.ai/news/magistral; arXiv 2506.10910).
- Codestral 25.08 (Jul 30, 2025) — succession of Codestral 25.01 — Codestral 25.08 (July 30, 2025) is the next Codestral flagship, announced with +30% accepted completions and 50% fewer runaway generations over prior versions — per mistral.ai/news/codestral-25-08.
- Mistral Large 3 (Dec 2, 2025) — succession of Mistral Medium 3
- Ministral 3 (Dec 2, 2025) — succession of Mistral Small 3 — Ministral 3 (December 2, 2025) is the small-end open-weights line in the Mistral 3 family relaunch — per the mistral.ai/news/mistral-3 announcement.
- Devstral 2 (Dec 9, 2025) — succession of Devstral Small — Devstral 2 (December 9, 2025) is 'our next-generation coding model family available in two sizes: Devstral 2 (123B) and Devstral Small 2 (24B)' — per mistral.ai/news/devstral-2-vibe-cli.
- Mistral OCR 3 (Dec 18, 2025) — succession of Mistral Large 3 — Mistral OCR 3 (December 17, 2025) ships alongside the Mistral 3 family on la Plateforme as the upgraded structured-document model; Mistral's announcement positions OCR 3 as a major upgrade over Mistral OCR 2 (74% win rate) and does not document any lineage from Mistral Large 3 — per mistral.ai/news/mistral-ocr-3. The OCR-3 ↔ Mistral 3 family edge here is a co-shipped successor relationship, not a disclosed base-model claim.
- Mistral Small 4 (Mar 16, 2026) — succession of Mistral Large 3
- Leanstral (Mar 16, 2026) — succession of Mistral Small 4 — Leanstral (March 16, 2026) is 'the first open-source code agent designed for Lean 4' and was 'built as part of the Mistral Small 4 family' — a 119B-parameter MoE with 6.5B activated per token, 128 experts, a 256k context and multimodal input, under Apache 2.0. Family membership is not a fine-tune or distillation claim, so the edge is typed succession. The 'Mistral Small 4 family' line is on the HuggingFace card at mistralai/Leanstral-2603; the mistral.ai/news/leanstral launch post no longer carries it and describes the model only as 'highly efficient (with 6B active parameters)'.
- Mistral Med 3.5 (Apr 28, 2026) — succession of Mistral Small 4 — Mistral Medium 3.5 (April 28, 2026) is 'our first flagship merged model … a dense 128B model with a 256k context window, handling instruction-following, reasoning, and coding in a single set of weights,' and Mistral adds 'We trained the vision encoder from scratch to handle variable image sizes and aspect ratios.' What it merges is stated in terms of products it retires, not weights it inherits: it 'replaces its predecessor Mistral Medium 3.1 and Magistral in Le Chat. It also replaces Devstral 2 in our coding agent Vibe.' Mistral names no relationship at all to Mistral Small 4, the release this edge follows chronologically, so no more-specific type is claimed here. Per the HuggingFace card at mistralai/Mistral-Medium-3.5-128B; the spec sheet at docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04 dates the release and gives the Modified MIT licence, 256k context and $1.50 / $7.50 per M pricing, but carries no lineage statement.
- Mistral OCR 4 (Jun 23, 2026) — succession of Mistral OCR 3 — Mistral OCR 4 (June 23, 2026) is the next document-intelligence flagship, adding paragraph-level bounding boxes, typed block classification and inline confidence scores over the text-extraction focus of prior generations, and is benchmarked head-to-head against 'our own Mistral OCR 3.' Mistral discloses no base-model relationship between the two — per mistral.ai/news/ocr-4.
- Leanstral 1.5 (Jun 30, 2026) — post-training of Leanstral — Leanstral 1.5 (June 30, 2026) declares base_model mistralai/Leanstral-2603 in its front matter and notes 'This is an updated model of our previously released Leanstral model' — an explicit same-base statement, so the edge is a post-training step rather than a succession. Per the HuggingFace card at mistralai/Leanstral-1.5-119B-A6B; the launch post at mistral.ai/news/leanstral-1-5 carries neither the base_model key nor that sentence, but it is where the training recipe is described ('Trained through mid-training, supervised fine-tuning, and reinforcement learning with CISPO').
Alibaba · Qwen
Versions page →Lineage as text (22 edges) ↓
Every edge in the Qwen tree above, in chronological order. Each line shows: destination model, edge type, source model, and the provider documentation that establishes the relationship (where one was disclosed).
- Qwen-14B (Sep 25, 2023) — succession of Qwen-7B
- Qwen2 family (Jun 7, 2024) — new base of Qwen-14B — Alibaba documents Qwen2 as its own pre-training run with an architecture change: 'we are pleased to announce the evolution from Qwen1.5 to Qwen2. This time, we bring to you: Pretrained and instruction-tuned models of 5 sizes,' and 'previously in Qwen1.5, only Qwen1.5-32B and Qwen1.5-110B have adopted Group Query Attention (GQA). This time, for all model sizes, we apply GQA.' All base models were 'pretrained on data of the context length of 32K tokens'. The same release moved most variants to Apache 2.0. Per qwenlm.github.io/blog/qwen2, bylined June 7, 2024. Note Alibaba measures the step from Qwen 1.5, which this tree does not carry as its own node.
- Qwen2.5 family (Sep 19, 2024) — new base of Qwen2 family — The launch post is explicit that Qwen2.5 is a fresh pretraining run, not a post-training pass over Qwen2: 'In terms of Qwen2.5, the language models, all models are pretrained on our latest large-scale dataset, encompassing up to 18 trillion tokens. Compared to Qwen2, Qwen2.5 has acquired significantly more knowledge' — per qwenlm.github.io/blog/qwen2.5.
- QwQ-32B Preview (Nov 28, 2024) — fine-tune of Qwen2.5 family — QwQ-32B-Preview (November 28, 2024) is the first Qwen reasoning model — 'an experimental research model developed by the Qwen Team, focused on advancing AI reasoning capabilities,' a 32B dense Apache 2.0 release whose model card declares base_model Qwen/Qwen2.5-32B-Instruct. Per the HuggingFace card at Qwen/QwQ-32B-Preview and the launch post at qwenlm.github.io/blog/qwq-32b-preview.
- Qwen2.5-Max (Jan 29, 2025) — succession of Qwen2.5 family — Qwen2.5-Max is the first proprietary closed-weights Qwen-Max release. Alibaba describes it as 'a large-scale MoE model that has been pretrained on over 20 trillion tokens and further post-trained with curated Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) methodologies' — its own pre-training run, but no relationship to the open Qwen2.5 weights is claimed either way, so the edge is typed succession. Per qwenlm.github.io/blog/qwen2.5-max, bylined January 28, 2025, a day before the launch date the Versions row records.
- QwQ-32B (Mar 5, 2025) — fine-tune of Qwen2.5 family — The QwQ-32B model card declares 'base_model: Qwen/Qwen2.5-32B' in its front matter — the Qwen2.5 base model, not the QwQ-32B-Preview released three months earlier, which itself declares 'base_model: Qwen/Qwen2.5-32B-Instruct'. The two reasoning releases are therefore siblings off the Qwen2.5 base rather than a chain. Per huggingface.co/Qwen/QwQ-32B; the qwenlm.github.io/blog/qwq-32b launch post carries a redirect banner to qwen.ai but still serves its full body, and that body names no base model — it credits 'the effectiveness of RL when applied to robust foundation models pretrained on extensive world knowledge' without saying which one — so the card remains the only disclosure.
- Qwen3 family (Apr 28, 2025) — new base of Qwen2.5 family — Alibaba quantifies the new pre-training run: 'While Qwen2.5 was pre-trained on 18 trillion tokens, Qwen3 uses nearly twice that amount, with approximately 36 trillion tokens covering 119 languages and dialects,' across three staged passes ending in a long-context extension to 32K. Qwen3 also introduces the hybrid Thinking / Non-Thinking design that the separate reasoning track had carried. Per qwenlm.github.io/blog/qwen3 (April 28, 2025).
- Qwen3 family (Apr 28, 2025) — lines merged of QwQ-32B — Qwen folds reasoning into the flagship rather than shipping it as a separate model: 'In the third stage, we integrated non-thinking capabilities into the thinking model by fine-tuning it on a combination of long CoT data and commonly used instruction-tuning data,' and, in the conclusion, 'We have seamlessly integrated thinking and non-thinking modes, offering users the flexibility to control the thinking budget.' The post benchmarks Qwen3-30B-A3B as outcompeting QwQ-32B but never says in words that QwQ is being retired; what it documents is the unification itself. No standalone QwQ model has shipped since. Per qwenlm.github.io/blog/qwen3 (April 28, 2025).
- Qwen3-Coder (Jul 22, 2025) — succession of Qwen3 family — Qwen3-Coder (July 22, 2025) is Alibaba's open-weights agentic-coding flagship, led by Qwen3-Coder-480B-A35B-Instruct. The launch post documents its own pre-training run — 7.5T tokens at a 70% code ratio, native 256K context extended to 1M with YaRN, and scaled code RL in post-training — rather than a fine-tune of a Qwen3 chat checkpoint, so the edge is a succession within the Qwen3 generation. Per qwenlm.github.io/blog/qwen3-coder.
- Qwen3-Next 80B (Sep 11, 2025) — new base of Qwen3 family — Qwen3-Next-80B-A3B (September 11, 2025) introduced a novel ultra-sparse MoE architecture (hybrid Gated DeltaNet + Gated Attention, plus a 'High-Sparsity Mixture-of-Experts' layout that 'achieves an extreme low activation ratio') as a new base — per huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct.
- Qwen3-Max (Sep 15, 2025) — succession of Qwen3 family — Qwen3-Max is the largest Qwen3-generation proprietary flagship.
- Qwen3-VL (Sep 23, 2025) — fine-tune of Qwen3 family — The Qwen3-VL Technical Report states the base outright: 'Qwen3-VL is instantiated in three dense variants (Qwen3-VL-2B/4B/8B/32B) and two MoE variants (Qwen3-VL-30B-A3B, Qwen3-VL-235B-A22B), all built upon Qwen3 backbones', with a SigLIP-2 vision encoder on top. The claim is in the paper body only — neither the arXiv abstract nor the HuggingFace card carries it, and the card declares no base_model key. Per the full text at arxiv.org/html/2511.21631v1.
- Qwen3-Coder-Next (Feb 4, 2026) — succession of Qwen3-Coder — Alibaba discloses no base for Qwen3-Coder-Next. The model card at huggingface.co/Qwen/Qwen3-Coder-Next never names Qwen3-Coder or Qwen3-Next as a base: it lists 'Training Stage: Pretraining & Post-training' for itself, 80B total / 3B activated, and the Gated DeltaNet + Gated Attention hybrid layout, and declares no base_model key. The Qwen blog entry for this release (qwen.ai/blog?id=qwen3-coder-next) returns a client-rendered shell with no post body, so the card is the only readable first-party artifact and the edge is chronological only.
- Qwen3.5 + Plus (Feb 16, 2026) — succession of Qwen3-Next 80B — Qwen3.5 (February 16, 2026) carries forward the hybrid Gated DeltaNet + sparse MoE architecture established in Qwen3-Next — per the qwen3.5 release.
- Qwen 3.6-Max (Apr 2, 2026) — succession of Qwen3-Max — Qwen 3.6-Max-Preview (April 2, 2026) is the next-generation proprietary closed-weights flagship, replacing Qwen3-Max in the Max product slot.
- Qwen3.6-35B-A3B (Apr 16, 2026) — succession of Qwen3.5 + Plus — Qwen3.6-35B-A3B (April 16, 2026) is the first open-weights Qwen3.6 release, a 35B-total/3B-active MoE — per HuggingFace Qwen/Qwen3.6-35B-A3B.
- Qwen3.6-27B (Apr 22, 2026) — succession of Qwen3.6-35B-A3B — Alibaba's card opens 'Following the February release of the Qwen3.5 series, we're pleased to share the first open-weight variant of Qwen3.6', and describes a 27B dense model with the 16 × (3 × Gated DeltaNet → FFN, 1 × Gated Attention → FFN) hybrid layout and a new Thinking Preservation option — a dense architecture rather than the 35B-A3B MoE beside it. It does not say the two share a base, and the 'Training Stage: Pre-training & Post-training' line it carries is a field every Qwen card carries, not a claim specific to this release, so nothing stronger than succession is disclosed. Per huggingface.co/Qwen/Qwen3.6-27B; the Qwen blog entry for this release is a client-rendered shell with no readable post body.
- Qwen3.7-Max (May 20, 2026) — succession of Qwen 3.6-Max — Qwen3.7-Max (May 20, 2026) is the closed-weights proprietary reasoning-agent flagship succeeding the 3.6-Max preview — per the Alibaba Cloud Summit announcement.
- Qwen3.7-Plus (May 31, 2026) — succession of Qwen3.7-Max — Qwen3.7-Plus (May 31, 2026) is the multimodal vision + language sibling to the text-only Qwen3.7-Max, together forming the Qwen3.7 generation announced at the May 20, 2026 Alibaba Cloud Summit. Alibaba's announcement describes Plus as retaining Max's coding / tool-use / reasoning strengths while adding image and video understanding, but does not disclose it as a fine-tune, post-training, or distillation of Qwen3.7-Max. Per the Alibaba Cloud Bailian / Model Studio Recommended models page; the Qwen blog entry at qwen.ai/blog?id=qwen3.7-plus returns a client-rendered shell with no post body.
- Qwen3.8-Max (Aug 3, 2026) — succession of Qwen3.7-Max — Qwen3.8-Max was previewed at the World AI Conference in Shanghai on July 19, 2026 and reached general availability on August 3, 2026: a 2.4-trillion-parameter sparse-MoE model with hybrid attention, 95 billion active parameters, native vision-language input and a 1M-token context, billed on Qwen Cloud and Alibaba Cloud Model Studio as qwen3.8-max. Qwen Cloud's dated changelog describes it as 'Qwen's most capable flagship model to date, featuring 2.4 trillion parameters with a Mixture-of-Experts architecture … with significant improvements over the 3.7 series' — a capability comparison, not a base-model continuity claim — and Alibaba discloses no fine-tune, post-training, or distillation relationship to Qwen3.7-Max, so the edge is typed succession. An upgraded snapshot, qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), shipped September 2, 2026; a dated snapshot of an existing model is not a separate lineage step, so it is not drawn as a node. Per docs.qwencloud.com/changelog/models (entries for August 3 and September 2, 2026) — the alibabacloud.com launch post that previously carried the architecture line no longer serves a readable body.
- Qwen3.8-27B (Aug 14, 2026) — succession of Qwen3.6-27B — Qwen3.8-27B (August 14, 2026) takes the dense open-weights slot Qwen3.6-27B held, but Alibaba claims no continuity between them. Its card introduces 'Qwen3.8, the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5' — an architecture statement about the generation, not a base-model one about this release — and lists 'Training Stage: Pre-training and Post-training' for the 27B itself, a line every Qwen card carries. The card's only reference to Qwen3.6-27B is a benchmark column. Qwen Cloud's changelog goes one step further, saying the model 'builds upon the 3.6-27B version, with key improvements in coding and office productivity capabilities across both text and visual modalities' (docs.qwencloud.com/changelog/models, August 19, 2026) — and that phrasing is specific to this entry rather than boilerplate. It is still a one-line product blurb against a model card that declares no base, so the edge stays at succession rather than claiming a shared base. Apache 2.0, 27B dense, natively vision-language. Per huggingface.co/Qwen/Qwen3.8-27B.
- Qwen3.8-Flash-Next (Aug 26, 2026) — succession of Qwen3.8-27B — Qwen3.8-Flash-Next (August 26, 2026) is shipped as an architecture preview rather than a lineage step: 'This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models work.' 125B total with 6B activated plus a 51B n-gram embedding table, natively vision-language, 262,144-token context extensible to 1M, released under the qwen-community-1.0 licence rather than Apache 2.0. The card also notes the hosted relationship in the other direction — 'Qwen3.8-Flash is the official version based on Qwen3.8-Flash-Next with more production features, e.g., 1M context length by default, official built-in tools' — and discloses no base-model continuity from any prior Qwen release, so the edge is typed succession. Per huggingface.co/Qwen/Qwen3.8-Flash-Next.
About this page
Cross-family comparison page in the /ai/ section. Each per-provider tree was hand-laid-out from the per-family Claude, ChatGPT, Gemini, Grok, Llama, DeepSeek, Mistral, and Qwen Versions pages on this site, each row's lineage edge sourced to the provider's own model card, technical paper, or release announcement.
Conservative claims. Where a provider has not formally documented base-model continuity between two releases, the edge is labeled "succession" and no further claim is made. The page does not infer "X is a fine-tune of Y" from external speculation, leaderboard chatter, or press coverage. The "succession" edge is honest about what is and isn't disclosed; the per-edge tooltip and the "Lineage as text" block name the source for every more-specific edge type.
Per-provider focus, not encyclopedic. Each tree highlights the lineage-meaningful releases — the new bases, the major post-trainings, the visible fine-tunes, the line merges. Per-release minutiae (small variants, intermediate checkpoints, every refresh on every tier) live on the per-family Versions pages where they belong. The /ai/release-cadence/ and /ai/context-windows/ pages are the right place for per-release counts and metrics.
What is intentionally excluded. Open-source community fine-tunes (the Llama-derivatives ecosystem — Vicuna, WizardLM, Nous Hermes, etc.) are not on the page; they are downstream community work, not frontier-lab releases. Capability comparisons ("which line is best") are out of scope. Speculation about undisclosed base-model continuity is out of scope. Unreleased models (announced but never publicly shipped, like Llama 4 Behemoth) are not included.
Refreshed daily, aligned with the per-family Versions pages. Each refresh re-verifies every disclosed lineage source (model cards / papers / announcements move under the same URLs but their content changes), adds nodes for new flagship releases, and prunes any node that the provider has formally deprecated. See release cadence and context windows for the cross-family ship-cadence and context-budget pictures this page complements.
Last updated: September 4, 2026. 151 models · 8 providers · 146 typed edges.