2023 – 2026

DeepSeek Versions

DeepSeek's current frontier model is DeepSeek-V4-Pro, first released in preview on April 24, 2026 and promoted to general availability on August 13, 2026 as the DeepSeek-V4-Pro-0813 build — 1.6 trillion total / 49B active parameters in a Mixture-of-Experts architecture with a 1,000,000-token context window, dual Thinking / Non-Thinking modes, and an MIT license. The companion DeepSeek-V4-Flash shipped the same day in April at 284B-total / 13B-active and reached its own official release — DeepSeek-V4-Flash-0731 — on July 31, 2026. Both GA builds put open weights on HuggingFace the same day they hit the API. The newest release is DeepSeek-V4-Flash-Vision-Exp (August 21, 2026), an experimental multimodal build whose open weights followed ten days later, on August 31. The V4 GA wave also ended DeepSeek's price war: from August 16, 2026 the API bills at peak / off-peak rates that are higher than the old flat rates on every leg — though the weekend peak windows were dropped a week later, on August 23, 2026. I track every DeepSeek release here — from DeepSeek-LLM in November 2023 onward — with HuggingFace ids, ship dates, family (Flagship / Reasoning / Specialized), and license terms. Below the table: the High-Flyer hedge-fund parentage, the December 2024 / January 2025 V3 / R1 inflection that triggered the largest single-day market-cap loss in U.S. stock-market history, the U.S. chip export-control context, the DeepSeek License vs. MIT evolution, and the funding arc that ran from the April 2026 Tencent / Alibaba talks through a ~$7.4 billion state-anchored first external round in June 2026 to the ~$7.4 billion pre-IPO second round that was still unclosed when its reported end-of-August deadline passed, ahead of a Shanghai STAR Market listing targeted for 2027.

Family & status

Family

Flagship — the main DeepSeek chat lineage from DeepSeek-LLM through V4
Reasoning — DeepSeek-R1 and R1-0528; converged into the V-series at V3.1's hybrid mode
Specialized — DeepSeek-Coder, DeepSeek-Math, DeepSeek-VL / VL2, DeepSeek-OCR, and Janus-Pro

Status

Current — actively recommended; the latest in its family
Available — weights still served via HuggingFace and partner inference providers, but superseded
Legacy — deprecated, experimental and superseded, or no longer recommended

DeepSeek version table

Model
DeepSeek-V4-Flash-Vision-Exp
deepseek-v4-flash-vision-exp, deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Flagship
Current
Aug 21, 2026
Experimental multimodal V4-Flash build — image understanding added, pure-text capability at parity with V4-Flash. Open weights followed on August 31, 2026, ten days after the API. MIT-licensed. Billed at V4-Flash rates.
  • Released August 21, 2026; the announcement is at api-docs.deepseek.com/news/news260821. Called via model='deepseek-v4-flash-vision-exp' on the same OpenAI-compatible endpoint as the rest of the V4 line.
  • Open weights followed on August 31, 2026 at deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, under MIT — the repo's initial commit is stamped August 31 at 06:16 UTC, the weights upload at 06:57, and the init release commit at 07:57. That is a ten-day gap between the API and the weights, the longest of the V4 line: both GA builds (V4-Flash-0731, V4-Pro-0813) published weights the same day they hit the API. The card describes the model as DeepSeek's “first experimental multimodal model in the DeepSeek-V4 family,” built on the V4-Flash architecture with visual modules added and continued training on top. FP8 checkpoint with FP4 experts, same as V4-Flash.
  • Mixed text + image input via base64, external URL, or the new Files API (free, and it lets one uploaded image be re-referenced by file_id across requests). Images are tokenized for billing at up to 384 tokens each and charged at V4-Flash input rates. Feature parity with V4-Flash except FIM completion, which is unsupported here.
  • Text capability at parity with the official V4-Flash; the leap is on multimodal agent work, where DeepSeek says the model lands close to Opus 4.8. Published figures: 83.9 Terminal Bench 2.1, 64.3 Chartography, 57.7 NL2Repo, 59.3 DeepSWE, 36.5 ApexBench (Pass@1), 35.0 ZeroBench (Pass@5), 27.3 Agents' Last Exam. The numbers are DeepSeek-reported and not independently verified.
  • Same 1,000,000-token context window, 384K max output, and dual Thinking / Non-Thinking modes as V4-Flash. DeepSeek Harness 0.1.1 shipped the same day with out-of-the-box support.
Model
DeepSeek-V4-Pro
deepseek-ai/DeepSeek-V4-Pro, deepseek-ai/DeepSeek-V4-Pro-0813
Flagship
Current
Apr 24, 2026
Current frontier flagship. 1.6T total / 49B active MoE. 1M-token context. Dual Thinking / Non-Thinking modes. General availability — V4-Pro-0813, an agent-capability upgrade — August 13, 2026, with open weights the same day. MIT-licensed.
  • Released April 24, 2026 in public preview; the announcement is at api-docs.deepseek.com/news/news260424. HuggingFace card: deepseek-ai/DeepSeek-V4-Pro; technical report: “DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence” (arXiv 2606.19348).
  • General availability on August 13, 2026, four months after the preview. Per the GA release post, the build rolled out across the app, the web client, and the API simultaneously; the deepseek-v4-pro model string is unchanged, so callers picked it up without a code change. The pricing page now names DeepSeek-V4-Pro-0813 as the model version behind that string, retiring the April V4-Pro-Preview build. Open weights landed the same day at deepseek-ai/DeepSeek-V4-Pro-0813 — the repo's Release DeepSeek-V4-Pro-0813 commit is stamped August 13, 2026 at 12:25 UTC.
  • The GA update is an agent-capability upgrade. DeepSeek's published figures are 87.9 on Terminal Bench 2.1, 83.3 on Cybergym, 74.1 on Toolathlon (verified), 62.7 on DeepSWE, 61.5 on NL2Repo, 71.1 / 67.2 on its internal DSBench-FullStack and DSBench-Hard sets, and 42.7 / 60.0 on Humanity's Last Exam without and with tools. The build adds native Responses API support with a Codex-specific adaptation — the promise the pricing page had been carrying past due since early August — and a three-level reasoning_effort control (low / high / max) that now applies to V4-Flash too. The numbers are DeepSeek-reported and not independently verified.
  • 1.6 trillion total parameters / 49B active per token — the largest publicly-released DeepSeek model. Mixture-of-Experts architecture pretrained on more than 32 trillion tokens with the Muon optimizer and Manifold-Constrained Hyper-Connections (mHC).
  • 1,000,000-token context window with up to 384K tokens of output, built on a new Hybrid Attention Architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) — the successor to V3.2's DeepSeek Sparse Attention. Per DeepSeek's technical report, V4-Pro requires only ~27% of single-token inference FLOPs and ~10% of KV cache compared with V3.2 in the 1M-token setting.
  • Reasoning-effort modes in a single model ID — V4 collapses the V3.x chat-vs-reasoner split into one model with a reasoning_effort parameter. The GA build takes three levels — low for simple tasks, high for everyday agent work (the default thinking mode), and max (marketed as "V4-Pro-Max", the maximum-effort mode, not a separate model) — plus a disabled non-thinking mode. Continues the hybrid-reasoning architecture introduced in V3.1.
  • License: MIT. API pricing at launch was $1.74 / M input, $3.48 / M output, discounted 75% as a launch promo; on May 22, 2026 DeepSeek made the discount permanent, settling at a flat $0.435 / M input ($0.003625 / M cached) and $0.87 / M output. That direction reversed on August 16, 2026. The August 6 warning of a “significant increase” (TechNode) landed with the V4 GA release: effective 16:00 UTC on August 16, the pricing page bills peak and off-peak rates, with off-peak set at half of peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday; everything else is off-peak. That weekday qualifier is a revision to the launch terms — the schedule originally ran its peak windows every day, and DeepSeek dropped the weekend ones effective 00:00 Beijing time on Sunday, August 23, 2026 (16:00 UTC on the 22nd), so Saturdays and Sundays now bill at off-peak rates around the clock (Bloomberg; the notice was first reported by Science and Technology Innovation Board Daily). V4-Pro now costs $0.66 / M input and $1.98 / M output off-peak ($0.022 / M cached), and $1.32 / $3.96 at peak ($0.044 / M cached). Both legs are a raise: off-peak is ~52% above the old flat input rate and ~128% above the output rate, and peak roughly triples input and quadruples output. Concurrency is capped at 500 requests. Available at chat.deepseek.com via Expert Mode, and via the OpenAI-compatible API endpoint.
  • Coverage of the launch in CNBC, Simon Willison, and Euronews. The release re-opened the U.S. Department of Commerce inquiry into whether V4 was trained on smuggled Nvidia Blackwell GPUs in violation of export controls (covered in the prose history below).
Model
DeepSeek-V4-Flash
deepseek-ai/DeepSeek-V4-Flash, deepseek-ai/DeepSeek-V4-Flash-0731
Flagship
Current
Apr 24, 2026
Fast / cheap V4 companion. 284B total / 13B active MoE. Same 1M context and dual-mode hybrid-attention architecture as V4-Pro. Official API release — V4-Flash-0731, an agent-capability upgrade — July 31, 2026, with open weights on HuggingFace. MIT-licensed.
  • Released April 24, 2026 alongside V4-Pro; HuggingFace card: deepseek-ai/DeepSeek-V4-Flash.
  • 284B total parameters / 13B active per token; same MoE-with-Hybrid-Attention (CSA + HCA) architecture as V4-Pro at a smaller scale, positioned for fast and economical inference.
  • Same 1,000,000-token context window and dual Thinking / Non-Thinking modes as V4-Pro.
  • Official API release on July 31, 2026, in public beta, as DeepSeek-V4-Flash-0731. Per DeepSeek's change log, the 0731 build “keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained”; the deepseek-v4-flash model string is unchanged, so callers pick it up without a code change.
  • Open weights shipped the same day at deepseek-ai/DeepSeek-V4-Flash-0731 — the repo's Release DeepSeek-V4-Flash-0731 commit is stamped July 31, 2026 at 12:02 UTC, a few hours after the initial commit at 07:30 UTC. (The repo's August 1 timestamp is a later model-card edit adding an SGLang cookbook link, not the weights drop.) The card describes it as “the official release of DeepSeek-V4-Flash, superseding the preview version,” shipping with the DSpark speculative-decoding module attached (the same structure as DeepSeek-V4-Flash-DSpark), so target and draft weights come from one checkpoint. The checkpoint is FP8-quantized with FP4 experts; reasoning_effort now takes three levels — low, high, and max. There is no Jinja chat template: DeepSeek ships an encoding folder of Python scripts for building prompts and parsing output instead.
  • The 0731 update is an agent-capability upgrade. DeepSeek's published figures are 82.7 on Terminal Bench 2.1, 54.4 on DeepSWE, 76.7 on Cybergym, 70.3 on Toolathlon (verified), and 54.2 on NL2Repo — scores the lab says exceed its own V4-Pro-Preview. The official build natively supports the Responses API and is adapted for Codex. The numbers are DeepSeek-reported and not independently verified; coverage at TechNode.
  • License: MIT. API pricing held at a flat $0.14 / M input ($0.0028 / M cached) and $0.28 / M output through the July 31 update, then went up with the rest of the V4 line on August 16, 2026. Under the peak / off-peak schedule described in the V4-Pro row, V4-Flash bills $0.22 / M input and $0.66 / M output off-peak ($0.007 / M cached) and $0.44 / $1.32 at peak ($0.014 / M cached) — a raise on every leg, roughly +57% / +136% off-peak and +214% / +371% at peak. It stays about a third of V4-Pro's rate. Concurrency is capped at 2,500 requests. Contemporaneous coverage tied the increase to the demand surge that followed this model's cheap launch rate.
  • Available at chat.deepseek.com via Instant Mode, and via the OpenAI-compatible API.
Model
DeepSeek-OCR-2
deepseek-ai/DeepSeek-OCR-2
Specialized
Available
Jan 27, 2026
Second-generation document-OCR model. ~3B parameters, second-gen DeepEncoder (“Visual Causal Flow”). Apache 2.0-licensed, following DeepSeek-Math-V2 off the MIT default.
  • Released January 27, 2026; HuggingFace card: deepseek-ai/DeepSeek-OCR-2.
  • ~3B-parameter vision model for document understanding / OCR — the successor to DeepSeek-OCR, with a second-generation DeepEncoder (“Visual Causal Flow”) for parsing dense document, table, and chart layouts.
  • License: Apache 2.0 — the second DeepSeek model on Apache 2.0, after DeepSeek-Math-V2 (November 2025) first moved off the MIT default the lab has used for V-series, R-series, and Janus-Pro releases since R1.
Model
DeepSeek-V3.2 (+ Speciale)
deepseek-ai/DeepSeek-V3.2, deepseek-ai/DeepSeek-V3.2-Speciale
Flagship
Available
Dec 1, 2025
DeepSeek Sparse Attention productionized. 685B total MoE. 128K context. The high-compute Speciale variant claimed IMO and IOI gold medals at this scale.
  • Released December 1, 2025; the announcement is at api-docs.deepseek.com/news/news251201. Technical paper: “DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models” (arXiv 2512.02556).
  • 685B total parameters MoE with Multi-Latent Attention and the productionized DeepSeek Sparse Attention (DSA) from V3.2-Exp. 128K-token context window.
  • DeepSeek-V3.2-Speciale is the high-compute reasoning variant shipped alongside the standard V3.2; per DeepSeek's release post, Speciale achieved gold-medal performance at the 2025 International Mathematical Olympiad (IMO) and the International Olympiad in Informatics (IOI), and was claimed to be on par with Gemini 3 Pro on broad reasoning benchmarks.
  • License: MIT. HuggingFace cards: DeepSeek-V3.2, DeepSeek-V3.2-Speciale.
  • Superseded by V4-Pro / V4-Flash four months later; weights remain available for self-host and via hosted inference.
Model
DeepSeek-Math-V2
deepseek-ai/DeepSeek-Math-V2
Specialized
Available
Nov 27, 2025
Open-weights mathematical-reasoning model built on V3.2-Exp-Base. Self-verification training (a learned verifier rewards a proof generator). The first DeepSeek release under Apache 2.0 rather than MIT.
  • Released November 27, 2025; HuggingFace card: deepseek-ai/DeepSeek-Math-V2.
  • 685B-parameter MoE built on DeepSeek-V3.2-Exp-Base, centered on self-verification — the lab trains an LLM verifier and then a proof generator rewarded by it, so the model checks its own reasoning rather than only producing a final answer.
  • Competition-math results: gold-level scores on IMO 2025 and CMO 2024, and a near-perfect 118/120 on Putnam 2024 with scaled test-time compute.
  • License: Apache 2.0 — the first DeepSeek model released under Apache 2.0 instead of the MIT default the lab has used for V-series, R-series, and Janus-Pro releases since R1; it predates DeepSeek-OCR-2 in making that move.
Model
DeepSeek-OCR
deepseek-ai/DeepSeek-OCR
Specialized
Legacy
Oct 20, 2025
First DeepSeek OCR model. “Contexts Optical Compression” — a DeepEncoder paired with a 3B-MoE (~570M-active) decoder. MIT-licensed.
  • Released October 20, 2025; HuggingFace card: deepseek-ai/DeepSeek-OCR; GitHub: deepseek-ai/DeepSeek-OCR; paper arXiv 2510.18234.
  • “Contexts Optical Compression” — a DeepEncoder vision frontend paired with a DeepSeek-3B-MoE (~570M activated) decoder, framing long-context text as a vision / OCR compression problem.
  • License: MIT. Superseded by DeepSeek-OCR-2 (January 2026).
Model
DeepSeek-V3.2-Exp
deepseek-ai/DeepSeek-V3.2-Exp
Flagship
Legacy
Sep 29, 2025
Experimental release. Introduced DeepSeek Sparse Attention (DSA) for long-context efficiency. Superseded by V3.2 stable two months later.
  • Released September 29, 2025 as an explicitly experimental release branched off V3.1; HuggingFace card: deepseek-ai/DeepSeek-V3.2-Exp; GitHub: github.com/deepseek-ai/DeepSeek-V3.2-Exp.
  • Introduced DeepSeek Sparse Attention (DSA) — an efficient attention mechanism that substantially reduces computational complexity in long-context scenarios while preserving model performance. The recipe became the architectural backbone of V3.2 stable and V4.
  • License: MIT. Status is Legacy on the “experimental, superseded by the stable V3.2 release” reading.
Model
DeepSeek-V3.1 (+ Terminus)
deepseek-ai/DeepSeek-V3.1, deepseek-ai/DeepSeek-V3.1-Terminus
Flagship
Legacy
Aug 21, 2025
Hybrid Thinking / Non-Thinking modes in a single model. 671B / 37B active. 128K context. The architectural convergence of V-series and R-series.
  • DeepSeek-V3.1 released August 21, 2025; the V3.1-Terminus stability / instruction-following update followed on September 22, 2025. Coverage in InfoQ; the API thinking-mode docs are at api-docs.deepseek.com/guides/thinking_mode.
  • Hybrid reasoning architecture — a single model that supports both fast non-thinking mode and a chain-of-thought thinking mode (called DeepSeek-V3.1-Think), governed by tokenizer parameters rather than separate model architectures.
  • 671B total parameters / ~37B active per token; 128K-token context window. Per DeepSeek, the thinking mode delivers quality comparable to DeepSeek-R1-0528 while reducing output tokens by 20–50% on the same tasks.
  • License: MIT. The convergence of the V-series and R-series in a single hybrid model is the architectural reason no further Reasoning-family rows have shipped after R1-0528.
  • HuggingFace cards: DeepSeek-V3.1, DeepSeek-V3.1-Terminus.
Model
DeepSeek-R1-0528
deepseek-ai/DeepSeek-R1-0528
Reasoning
Available
May 28, 2025
R1 update shipped in lieu of the rumored R2. Improved reasoning depth, hallucination rate, and tool use. The last standalone R-series release before V3.1's hybrid convergence.
  • Released May 28, 2025; HuggingFace card: deepseek-ai/DeepSeek-R1-0528.
  • An R1 update, not a DeepSeek-R2. The April–May 2025 rumor cycle had widely predicted an R2 release; Reuters later reported R2 was delayed by data-labelling and chip-availability constraints. R1-0528 shipped instead and was widely read as the substitute.
  • Same 671B / 37B-active MoE architecture as R1, with substantially-improved reasoning depth, lower hallucination rate, and improved tool / function calling per DeepSeek's release notes.
  • License: MIT. Last standalone R-series release; subsequent reasoning capability ships as the “Thinking” mode of the V-series starting at V3.1 (August 2025).
Model
DeepSeek-V3-0324
deepseek-ai/DeepSeek-V3-0324
Flagship
Legacy
Mar 24, 2025
First MIT-relicensed flagship. Improved reasoning, coding, and tool use over V3. Often cited as “DeepSeek-V3.1” informally before V3.1 proper shipped.
  • Released March 24, 2025; HuggingFace card: deepseek-ai/DeepSeek-V3-0324; coverage in SiliconANGLE.
  • First DeepSeek flagship released under the MIT License rather than the bespoke “DeepSeek License.” The shift signaled a structural pivot toward fully-open licensing for the V-series; subsequent V3.1 / V3.2 / V4 releases all ship under MIT.
  • Same 671B / 37B-active MoE architecture as V3, with improved reasoning, coding, and tool / function calling per the model card.
  • Often referenced informally as “DeepSeek-V3.1” in third-party coverage during March–August 2025; the proper DeepSeek-V3.1 (the row above) shipped on August 21, 2025 and is architecturally distinct.
Model
Janus-Pro (1B / 7B)
deepseek-ai/Janus-Pro-{1B, 7B}
Specialized
Available
Jan 27, 2025
Unified multimodal understanding-and-generation. SigLIP-L vision encoder. Outperformed DALL-E 3 and SD3 on GenEval at 7B. MIT-licensed.
  • Released January 27, 2025 — the same trading day as the Nvidia stock crash triggered by the broader DeepSeek-R1 narrative (covered in the prose history below). Paper: “Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling”.
  • Unified multimodal architecture — a single model for both image understanding and text-to-image generation, building on the earlier Janus / JanusFlow family. Two sizes shipped (1B and 7B activated parameters); SigLIP-L vision encoder.
  • Per DeepSeek's release post, Janus-Pro 7B outperformed OpenAI's DALL-E 3 and Stability AI's Stable Diffusion 3 medium on GenEval and DPG-Bench at launch.
  • License: MIT. HuggingFace card: deepseek-ai/Janus-Pro-7B.
Model
DeepSeek-R1 (+ R1-Zero)
deepseek-ai/DeepSeek-R1, deepseek-ai/DeepSeek-R1-Zero
Reasoning
Legacy
Jan 20, 2025
Open-weights reasoning model. RL-only training (R1-Zero) demonstrated emergent chain-of-thought. Triggered the largest single-day market-cap loss in U.S. stock-market history a week later.
  • Released January 20, 2025; the announcement is at api-docs.deepseek.com/news/news250120. Paper: “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning” (arXiv 2501.12948); Nature follow-up: DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning.
  • 671B total / ~37B active per token MoE built on the V3 base, with reasoning capability incentivized through reinforcement learning. Performance was characterized by DeepSeek as comparable to OpenAI's o1 across math, coding, and reasoning benchmarks.
  • DeepSeek-R1-Zero is the load-bearing scientific result — trained with RL only (no supervised fine-tuning at all), R1-Zero exhibited emergent chain-of-thought reasoning, self-verification, and reflection behaviors. R1 itself adds an SFT cold-start to fix the readability and language-mixing issues in R1-Zero outputs.
  • The earlier DeepSeek-R1-Lite preview (November 20, 2024) had been the public hint at the reasoning track; R1 itself was the production release.
  • Distilled smaller variants (1.5B / 7B / 8B / 14B / 32B / 70B based on Qwen and Llama base models) shipped alongside R1; status of the distilled set varies by base model and is tracked on the HuggingFace org.
  • License: MIT. The R1 release week culminated in the January 27, 2025 Nvidia crash described in the prose history below; the broader market-cap impact and the dispute over causation are factual events documented there.

The V3 / R1 inflection — late December 2024 / January 2025. Above this line: every DeepSeek model from V3 onward, all under MIT or transitioning to MIT, shipped after the global-attention moment that DeepSeek-V3 (December 26, 2024) and DeepSeek-R1 (January 20, 2025) created. Below: the pre-inflection lineage — DeepSeek-LLM, the V2 architectural foundation, Coder, Math, VL, VL2 — mostly under the bespoke “DeepSeek License,” quietly building the MoE / MLA recipe that V3 productionized at frontier scale.

Model
DeepSeek-V3
deepseek-ai/DeepSeek-V3, deepseek-ai/DeepSeek-V3-Base
Flagship
Legacy
Dec 26, 2024
671B total / 37B active MoE. 14.8T-token pretraining. Reported $5.576M training cost on 2.788M H800 GPU hours. The frontier-class disclosed-compute moment.
  • Released December 26, 2024 (V3-Base + V3 chat); technical report at arXiv 2412.19437; HuggingFace card: deepseek-ai/DeepSeek-V3.
  • 671 billion total parameters / ~37 billion activated per token in a MoE architecture, with Multi-Latent Attention (MLA) and the DeepSeekMoE recipe productionized at frontier scale for the first time. Pretrained on 14.8 trillion tokens.
  • Disclosed training cost of $5.576 million on 2.788 million H800 GPU hours — the figure that reset the public conversation about AI training cost. The number excludes prior-stage research and infrastructure capex (the Fire-Flyer cluster the run depended on); see the prose history below for the dispute.
  • Performance was characterized at launch as competitive with GPT-4o and Claude 3.5 Sonnet on broad benchmarks at a fraction of the disclosed compute. 128K context window.
  • Originally distributed under the bespoke DeepSeek License for the model weights (an OpenRAIL-derived license with use-based restrictions) and MIT for the code repository; relicensed to MIT for the V3-0324 update three months later.
Model
DeepSeek-VL2 (Tiny / Small / VL2)
deepseek-ai/deepseek-vl2-{tiny, small}, deepseek-ai/deepseek-vl2
Specialized
Available
Dec 13, 2024
First MoE vision-language line. Three sizes (1.0B / 2.8B / 4.5B activated). Dynamic-tiling vision encoder. OCR / chart / document understanding.
  • Released December 13, 2024; paper at arXiv 2412.10302; GitHub: deepseek-ai/DeepSeek-VL2.
  • Three sizes — VL2-Tiny (1.0B activated), VL2-Small (2.8B activated), VL2 (4.5B activated). MoE language tower with MLA, dynamic-tiling vision encoder for variable-aspect-ratio inputs.
  • Targets visual question answering, optical character recognition, document / table / chart understanding, and visual grounding. Strong OCR results on OCRBench at launch.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository. The Janus-Pro line (one row above) is the more recent multimodal flagship.
Model
DeepSeek-V2.5
deepseek-ai/DeepSeek-V2.5, deepseek-ai/DeepSeek-V2.5-1210
Flagship
Legacy
Sep 5, 2024
Merged V2 chat and Coder-V2 into a single general-purpose model. Revised in December 2024 (V2.5-1210). The bridge to V3.
  • Released September 5, 2024 — DeepSeek's change log dates the deepseek-chat / deepseek-coder upgrade to V2.5 that day, and the weights repo's initial commit and Upload folder commit are both stamped September 5. Revised December 10, 2024 as V2.5-1210. HuggingFace cards: DeepSeek-V2.5, DeepSeek-V2.5-1210.
  • Merged the V2 chat and Coder-V2 lineages into a single general-purpose model, simplifying the deployment story and serving as the bridge release between V2 and V3.
  • Same 236B / 21B-active MoE architecture as V2, with improved general / coding capability and better instruction following per the release notes.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository.
Model
DeepSeek-Coder-V2 (16B / 236B)
deepseek-ai/DeepSeek-Coder-V2-{Lite-, }Instruct, -Base
Specialized
Available
Jun 17, 2024
First MoE coding model. Two sizes (16B / 236B). 338-language coverage. Reported parity with GPT-4 Turbo / Claude 3 Opus on HumanEval at launch.
  • Released June 17, 2024; HuggingFace collection: DeepSeekCoder-V2.
  • Two sizes — Coder-V2-Lite (16B total / 2.4B active) and Coder-V2 (236B total / 21B active), the latter built on the V2 base. 338-language coverage; 128K context.
  • Reported parity with GPT-4 Turbo and Claude 3 Opus on HumanEval and MBPP at launch — the first open-weights coding model to claim that benchmark range.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository. Superseded as the recommended coding option by general-purpose V3 / V3.1+ instruction tuning by 2025.
Model
DeepSeek-V2
deepseek-ai/DeepSeek-V2, deepseek-ai/DeepSeek-V2-Lite
Flagship
Legacy
May 6, 2024
First DeepSeek MoE flagship. 236B total / 21B active. Multi-Head Latent Attention. 42.5% training-cost reduction and 93.3% smaller KV cache vs. V1.
  • Released May 6, 2024; paper: “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model” (arXiv 2405.04434).
  • 236 billion total parameters / 21 billion activated per token in a Mixture-of-Experts architecture; supported a 128K-token context window.
  • Multi-Head Latent Attention (MLA) debuted here — the architectural innovation that compresses KV cache into a low-rank latent vector. MLA + DeepSeekMoE delivered 42.5% training-cost reduction, 93.3% smaller KV cache, and 5.76× maximum throughput versus DeepSeek-LLM 67B per the paper. The architecture is the backbone every subsequent V-series release builds on.
  • A smaller V2-Lite (15.7B total / 2.4B active) companion shipped alongside.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository.

The MoE / MLA architecture turn — May 2024. Above this line: every DeepSeek release built on the Multi-Head Latent Attention + DeepSeekMoE recipe introduced in V2. Below: the dense-architecture pre-MoE era — DeepSeek-LLM 7B / 67B, the original DeepSeek-Coder, DeepSeek-Math, and DeepSeek-VL — the foundation lineage that established the lab and the Fire-Flyer infrastructure but had not yet found the architecture that would let DeepSeek punch at frontier scale.

Model
DeepSeek-VL (1.3B / 7B)
deepseek-ai/deepseek-vl-{1.3b, 7b}-{base, chat}
Specialized
Legacy
Mar 8, 2024
First DeepSeek vision-language model. Two sizes. SigLIP / SAM-B hybrid vision encoder. Superseded by VL2 nine months later.
  • Released March 8, 2024 as the first DeepSeek vision-language model; HuggingFace org: deepseek-ai.
  • Two sizes (1.3B and 7B). Hybrid vision encoder combining SigLIP and SAM-B for fine-grained image understanding.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository. Superseded as the recommended VL by VL2 (December 2024) and Janus-Pro (January 2025).
Model
DeepSeekMath 7B
deepseek-ai/deepseek-math-7b-{base, instruct, rl}
Specialized
Legacy
Feb 5, 2024
7B math-specialist. Introduced GRPO — Group Relative Policy Optimization — the RL method later used to train DeepSeek-R1.
  • Released February 5, 2024; paper: “DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models” (arXiv 2402.03300).
  • 7B math-specialist built on DeepSeek-Coder-Base; pretrained on a 120B-token math-corpus filtered from Common Crawl. Three flavors shipped: base, instruct, and the RL-tuned variant.
  • Introduced Group Relative Policy Optimization (GRPO) — the RL recipe that DeepSeek later used to train DeepSeek-R1's reasoning behavior. The DeepSeekMath paper is the load-bearing methodology citation behind R1.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository.
Model
DeepSeek-LLM (7B / 67B)
deepseek-ai/deepseek-llm-{7b, 67b}-{base, chat}
Flagship
Legacy
Nov 29, 2023
First DeepSeek general-purpose LLM. Dense architecture. Two sizes. 2T-token pretraining. Grouped-Query Attention. The start of the Flagship lineage, four weeks after DeepSeek-Coder.
  • Released November 29, 2023; HuggingFace cards: deepseek-llm-7b-base, deepseek-llm-67b-base, plus -chat variants. GitHub: deepseek-ai/DeepSeek-LLM. All four weight repos were created and uploaded on November 29, and the GitHub repo the same day.
  • Two dense sizes (7B and 67B), pretrained from scratch on 2 trillion tokens in English and Chinese. Grouped-Query Attention at 67B. The lab's first general-purpose LLM line, four months after DeepSeek's July 2023 founding.
  • The 67B chat variant was widely considered competitive with Llama 2 70B and ChatGLM-3 at launch on broad benchmarks.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository. Architecturally superseded by the V2 MoE family six months later.
Model
DeepSeek-Coder (1.3B / 6.7B / 33B)
deepseek-ai/deepseek-coder-{1.3b, 6.7b, 33b}-{base, instruct}
Specialized
Legacy
Nov 2, 2023
DeepSeek's first public model release. Three sizes. 87-language coverage, 16K context. Strong open-weights coding-model competitor to Code Llama and StarCoder at launch.
  • Released November 2, 2023 as the first DeepSeek-Coder — and, by four weeks, the lab's first public model release of any kind, ahead of DeepSeek-LLM. HuggingFace collection: DeepSeek-Coder.
  • Three sizes (1.3B, 6.7B, 33B). 87-language coverage. 16,384-token context window and project-level pretraining for repo-aware completion.
  • Reported strong results vs. Code Llama and StarCoder on HumanEval and MBPP at launch; established DeepSeek's reputation in the coding-model niche before the broader brand became globally cited.
  • Distributed under the bespoke DeepSeek License for the model weights; MIT for the code repository. Superseded by Coder-V2 (June 2024) and the general-purpose V3 / V3.1 lineage.

Click any row to expand. Each row has a stable id for sharing — e.g. /ai/deepseek/versions/#deepseek-v4-pro, #deepseek-r1, #deepseek-v3, #deepseek-v2. DeepSeek API docs: api-docs.deepseek.com; HuggingFace org: huggingface.co/deepseek-ai; GitHub org: github.com/deepseek-ai.

High-Flyer, Liang Wenfeng, and the July 2023 founding

DeepSeek was founded on July 17, 2023 in Hangzhou, China, by Liang Wenfeng. The lab's parent and sole funder is High-Flyer, the Chinese quantitative hedge fund Liang co-founded in February 2016. Liang serves as CEO of both companies.

Liang's background is unusual for a frontier-AI founder: he studied electrical engineering at Zhejiang University, began trading equities during the 2008 financial crisis as an undergraduate, and built High-Flyer as an AI-driven quant fund. By 2021, High-Flyer was reportedly using AI exclusively for trading decisions and had become one of the largest quantitative funds in China. The fund's profitability is what underwrote the AI-research investment that produced DeepSeek; reporting in Fortune and ChinaTalk covers the trajectory.

DeepSeek was wholly owned by High-Flyer from incorporation through mid-2026. The April 2026 reports of Tencent and Alibaba investment talks — reshaped over May–June 2026 into a far larger state-anchored round that, per June 17, 2026 reporting, raised more than 50 billion yuan (~$7.4 billion) at a $50 billion-plus valuation (covered in the funding section below) — are the company's first external funding round. The round is structured so Liang retains control through a limited partnership he manages; the High-Flyer relationship stays intact.

Fire-Flyer 2 and the GPU stockpile

Before DeepSeek existed as an independent entity, High-Flyer had been building the GPU infrastructure the lab would inherit. Liang began acquiring Nvidia GPUs at scale starting in 2021, reportedly building a stockpile of around 10,000 Nvidia A100 chips before the U.S. October 7, 2022 export controls first restricted top-tier-AI-GPU exports to China. The pre-controls procurement window is the single most-cited piece of context for why a Chinese lab can train at scale despite the sanctions: the chips were already on the floor before the restriction took effect.

High-Flyer built the Fire-Flyer 2 cluster beginning in 2021 with a reported budget of 1 billion yuan. Per the cluster's published statistics, Fire-Flyer 2 had reached 5,000 PCIe A100 GPUs in 625 nodes with ~96% utilization through 2022, totaling ~56.74 million GPU-hours of capacity used. The cluster was the load-bearing infrastructure for everything DeepSeek shipped from DeepSeek-LLM through V3, and the implicit denominator behind the “$5.576M training cost” figure for V3 — that figure is the marginal compute cost of one training run, not the cumulative R&D and infrastructure cost of building the cluster the run depended on. The dispute over whether the headline cost number is misleading hinges on this distinction.

The V3 / R1 inflection — December 2024 / January 2025

DeepSeek-V3 shipped on December 26, 2024 as a 671B-total / 37B-active MoE model with a disclosed training cost of $5.576 million on 2.788 million H800 GPU hours, pretrained on 14.8 trillion tokens. Performance on broad benchmarks at launch was characterized as competitive with GPT-4o and Claude 3.5 Sonnet. The combination — frontier-adjacent quality, fully open weights, and an order-of-magnitude-lower disclosed compute number — was the data point the AI-infrastructure capex thesis had not previously had to absorb. The technical report is at arXiv 2412.19437.

DeepSeek-R1 followed on January 20, 2025, built on the V3 base, with reasoning capability incentivized through reinforcement learning. The accompanying R1-Zero result — trained with RL only and no supervised fine-tuning — demonstrated emergent chain-of-thought reasoning, self-verification, and reflection behaviors, and was the load-bearing scientific claim of the release. Performance was characterized as comparable to OpenAI's o1 across math, coding, and reasoning benchmarks. The paper is at arXiv 2501.12948; a Nature follow-up published September 17, 2025 is at nature.com/articles/s41586-025-09422-z.

On January 27, 2025, the public-equity reaction to the R1 narrative produced what was at the time the largest single-day loss in U.S. stock-market history: Nvidia fell ~17% and shed approximately $589 billion in market capitalization in a single trading session. The Nasdaq fell ~3% on the day; AI-infrastructure names (Broadcom, Marvell, Vertiv, Constellation Energy) sold off in sympathy. CNBC, Yahoo Finance.

The causation is disputed. Tim Lee at Understanding AI argued that the move was already in motion before the R1 release week, that the disclosed-compute figure excluded prior R&D and the Fire-Flyer build, and that an alternative reading is “efficiency gains expand the inference-compute market faster than they shrink the training-compute market” (the Jevons-paradox response that Microsoft, Meta, and Google all subsequently adopted publicly). DeepSeek did not retract the disclosed-compute figure; the dispute is over the framing, not the number.

The U.S. chip export-control context

DeepSeek's training infrastructure has been scrutinized by U.S. policymakers since the R1 release week. The October 7, 2022 U.S. Department of Commerce export controls restricted top-tier AI-GPU exports to China; Nvidia subsequently produced the H800, a deliberately-degraded H100 variant designed to fall under the export-control thresholds, which it sold legally to Chinese customers including DeepSeek. The H800 was banned in turn in October 2023, but the year-long gap between the original control and the H800 ban was sufficient for DeepSeek to procure the chips it disclosed using to train V3 in 2024. Coverage at CSIS, RAND.

Following the R1 release, the U.S. Department of Commerce opened an inquiry into whether DeepSeek had used U.S. chips not legally exportable to China. House Select Committee chairs Krishnamoorthi and Moolenaar issued a public call to tighten the existing controls in February 2025. Through 2025, several U.S. state governments and federal agencies banned the DeepSeek consumer chatbot on government devices on data-handling grounds (the same regime that had been applied to TikTok); the bans do not apply to the open-weights releases on HuggingFace, which can be self-hosted on Western infrastructure.

The April 2026 V4 release re-opened the chip-controversy docket. Reporting in 2026 alleged that the V4 training run used clusters of Nvidia Blackwell B200 GPUs — a chip class that is comprehensively export-controlled to China — reportedly housed at a data center in Inner Mongolia. In February 2026 a Trump-administration official told Reuters that DeepSeek's latest model was trained on Nvidia's most advanced Blackwell chip in possible violation of the export controls, alleging the company would strip the technical indicators that reveal U.S.-hardware provenance; a senior State Department official separately alleged DeepSeek had supported Chinese military and intelligence operations and used Southeast Asian shell companies to access restricted chips. DeepSeek has not publicly confirmed the chips it used to train V4.

Despite those allegations, DeepSeek has not been added to the Commerce Department's Entity List. Reuters reported on June 17, 2026 that an interagency committee (Commerce, Defense, Energy, State) had approved DeepSeek — along with memory chipmaker CXMT and more than 100 other Chinese firms — for blacklisting last year, but the Trump administration held the listings unpublished to avoid disrupting the fragile U.S.–China trade truce. No entity has been added to the list since October 2025 — a gap CSIS's Philip Luck called the longest between postings in over a decade, and one that still stood as of early September 2026, since every Entity List rule the Bureau of Industry and Security has published since is a revision or a removal. The DeepSeek consumer-chatbot bans on U.S. state-government and federal-agency devices remain in place separately; they do not apply to the open-weights releases on HuggingFace.

On July 7, 2026, Reuters reported, citing three people familiar with the matter, that DeepSeek is developing its own inference chip — silicon for serving trained models rather than training new ones — after roughly a year of talks with chip-design, foundry, and memory partners and private hiring of chip-design engineers. The reported aim is to reduce dependence on both Nvidia, whose most advanced parts are export-controlled to China, and Huawei, whose Ascend silicon reportedly runs DeepSeek's cloud service today. DeepSeek has not publicly confirmed the effort, and no tape-out or foundry partner has been named.

In the meantime the Huawei dependence is deepening rather than shrinking. On September 4, 2026, Bloomberg reported that DeepSeek plans to install at least 160,000 of Huawei's next-generation Ascend 950DT accelerators at the roughly one-gigawatt data center it is building in Inner Mongolia — which would be among the largest known clusters of Huawei AI silicon anywhere, and a concrete step in China's effort to substitute domestic parts for Nvidia's. Per the report the chips are earmarked for running models rather than training them, even though Huawei designed and markets the 950DT for the more demanding training workload as well: DeepSeek has tried training on Huawei parts before and has so far kept Nvidia accelerators for that step. The installation timetable is reported to depend on Huawei's production capacity, and that constraint is sharper than the phrasing suggests — shortages of components such as top-end memory will cap Huawei's 950DT output at the low hundreds of thousands of units this year, filling DeepSeek's order could take more than a year, and DeepSeek has asked Beijing to help press Huawei for a larger and faster allocation. The 160,000 accelerators would cover only one chunk of the site's eventual gigawatt-scale capacity; Bloomberg writes that it is unclear what silicon fills the rest, and that Trump-administration officials have alleged DeepSeek procured Nvidia's top Blackwell parts and installed them in Inner Mongolia — a claim the wire says it has not independently verified. All sources are unnamed, DeepSeek did not respond to requests for comment, and nothing in the report resolves what hardware trained V4.

The DeepSeek License vs. MIT — the licensing turn

DeepSeek's licensing has evolved across two distinct conventions. From DeepSeek-LLM (November 2023) through DeepSeek-V3 (December 2024), the model weights shipped under the bespoke “DeepSeek License” — an OpenRAIL-derived custom license with use-based restrictions (military, surveillance, deceptive content, certain weapons applications) and a separate commercial-license track. The associated GitHub source code repos shipped under MIT separately. This is the same code-vs-weights split Meta uses for the Llama lineage, but DeepSeek's bespoke license is differently shaped and was not OSI-approved. Black Duck's model-license review from January 2025 walks the original terms.

The licensing turn is DeepSeek-R1 (January 20, 2025), which was the first DeepSeek flagship released under the MIT License. DeepSeek-V3-0324 (March 24, 2025) re-released the V3 weights under MIT, retroactively bringing the V-series flagship into MIT-compliance for the post-V3 era. Every subsequent V-series and R-series release — R1-0528, V3.1, V3.1-Terminus, V3.2-Exp, V3.2, V3.2-Speciale, V4-Pro, V4-Flash, the V4-Pro-0813 / V4-Flash-0731 GA builds, and the experimental V4-Flash-Vision-Exp — has shipped under MIT. Janus-Pro (January 27, 2025) also shipped under MIT, as did DeepSeek-OCR (October 2025). DeepSeek-Math-V2 (November 27, 2025) was the first DeepSeek model to ship under Apache 2.0 instead of MIT, followed by its OCR successor DeepSeek-OCR-2 (January 2026).

The pre-R1 specialized models (Coder, Coder-V2, Math, VL, VL2) remain on the original DeepSeek License for the model weights as of this page's publication date; whether DeepSeek will retroactively relicense the older specialized weights to MIT is open. The one release that briefly sat outside both buckets no longer does: DeepSeek-V4-Flash-Vision-Exp was API-only for its first ten days, and its August 31, 2026 weights drop landed under MIT like the rest of the V4 line. For new builds, the practical guidance is “everything from R1 forward is MIT — with DeepSeek-Math-V2 and DeepSeek-OCR-2 the two Apache-2.0 exceptions — while the older specialized models retain the use-restriction terms of the DeepSeek License.” Read the LICENSE-MODEL file in the relevant GitHub repo before shipping at scale.

The 2026 first-external-round funding talks

Through April 2026, DeepSeek had raised no external capital — it was funded entirely by High-Flyer's profits since the July 2023 incorporation. On April 22, 2026, Bloomberg and The Information reported that Tencent and Alibaba were in talks to invest a combined ~$1.8 billion at a $20 billion+ valuation. Tencent had proposed acquiring up to a 20% stake but DeepSeek was reluctant to cede that share of control; Alibaba's role was reportedly smaller. The talks landed two days before the V4 release.

By mid-May 2026 the round had reshaped substantially — from the original $20 billion-plus Tencent / Alibaba framing into a far larger state-anchored round. Early-June 2026 reporting (Reuters, Bloomberg) described a round of roughly $7.4 billion (~50 billion yuan) at a 350–400 billion yuan ($52–59 billion) valuation, with Tencent (~10 billion yuan) and battery maker CATL (~5 billion yuan) named participants and founder Liang Wenfeng contributing roughly 20 billion yuan of his own capital.

By June 17, 2026, the Wall Street Journal and The Information reported that DeepSeek had raised more than 50 billion yuan (~$7.4 billion) in its first external round, valuing the company at more than $50 billion — among the largest AI funding rounds in Chinese history. The deal's defining feature is its control structure: outside investors place their capital into a limited partnership Liang manages rather than buying DeepSeek equity directly, accept a five-year lock-up, and receive no voting rights, so Liang retains control. China's state-backed National Artificial Intelligence Industry Investment Fund is the sole exception — it invested roughly 1 billion yuan directly into the company with voting rights and no lock-up. Reported participants include Tencent, CATL, JD.com, NetEase, Hillhouse, IDG Capital, and Monolith Capital; proceeds are earmarked for compute infrastructure and employee compensation. DeepSeek did not comment publicly. The state-vehicle lead, the founder's large self-contribution, and the voting-rights asymmetry are the load-bearing changes from the original Tencent / Alibaba framing.

Less than a month after that round closed, the cadence accelerated again. On July 14, 2026, the Financial Times reported — picked up by TechCrunch and Bloomberg — that DeepSeek is in talks to raise a further ~$1.5 billion at roughly a $71 billion valuation, and has begun preparing an initial public offering on a mainland Chinese exchange, with a filing targeted as early as the end of 2026 and a debut in 2027. The reported driver is the capital needed to build DeepSeek's own data-center capacity and secure more AI chips. Both the round and the IPO are reported-not-confirmed: DeepSeek has not commented publicly, and no prospectus has been filed.

Reuters put firmer numbers on both items the following week. In wire copy dated July 20, 2026 (carried by The Manila Times), Reuters reported that DeepSeek is planning a fresh round of as much as 50 billion yuan (~$7.4 billion) at a valuation of about 500 billion yuan ($74 billion) — a far larger raise than the ~$1.5 billion the FT had described — and that the June round had closed at a post-money valuation of roughly 450 billion yuan, though filings by two Chinese investors later implied a 350.88 billion yuan (~$52 billion) mark. On the listing, Reuters named Shanghai's Nasdaq-style STAR Market as the venue under early deliberation and reported an internal target to complete an IPO filing before the end of 2026. All sources were unnamed, DeepSeek did not respond to a request for comment, and the wire cautioned that both the terms and the timetable may change.

Five days later the second round stopped. On July 25, 2026, Bloomberg reported — carried by Fortune and Yahoo Finance — that DeepSeek had verbally told prospective investors it was suspending the deal for now, and that they would not be signing the investment agreements they had expected in the days ahead. The reported trigger is not the terms but a leak: a transcript of a meeting Liang held with unidentified parties circulated online, and Chinese outlets including Yicai reported that he had discussed DeepSeek's reliance on Nvidia chips and China's persistent lag behind the U.S. in AI sophistication. Bloomberg did not verify the transcript's authenticity, and DeepSeek did not respond to a request for comment on either the transcript or the fundraising. Bloomberg put the paused round at at least 10 billion yuan (~$1.5 billion) at a pre-money valuation of at least 480 billion yuan (~$71 billion) — the Financial Times' figures rather than the larger ones Reuters described — and reported that negotiations remain fluid and the company may resume the process later. The IPO preparation was described as continuing, with a filing possible before the end of the year.

It resumed twelve days after that. On August 6, 2026, Bloomberg reported that DeepSeek had restarted the round and is now seeking close to $8 billion at a valuation near 500 billion yuan (~$74 billion) — the larger figures Reuters had described in July rather than the ~$1.5 billion the FT reported, and roughly a second raise the size of the first. Monolith Management, an early backer of Moonshot AI, is in talks to join. Reported use of proceeds is DeepSeek's own data-center capacity, led by a large facility in Inner Mongolia. Bloomberg cautioned that the size, timing, and investor list can all still change, and DeepSeek has not commented publicly. Two other moves landed the same day and cut against the lab's cheap-compute reputation: the API price-increase warning noted in the V4 rows above, and a 140.8 million yuan ($20.8 million) investment in humanoid-robot maker Unitree Robotics — 933,399 shares, or 2.31% of the strategic placement in Unitree's Shanghai IPO, under a 36-month lock-up — paired with an agreement to jointly develop AI models for humanoid machines, with each company favoring the other for model-training services and robot purchases respectively (Reuters, from a stock-exchange filing). The partnership is DeepSeek's first substantial move toward embodied AI, a direction its language-first model line has not covered.

Three weeks later the round was reported to be closing. On August 26, 2026, the South China Morning Post reported that DeepSeek was nearing completion of a raise of about 50 billion yuan (~$7.4 billion) at a pre-money valuation near 500 billion yuan (~$74 billion), expected to close before the end of August — which would imply a post-money mark around $81 billion. The Wall Street Journal reported the same $74 billion figure and framed the raise as building a war chest ahead of a listing. Named participants are returning backers Monolith, Shixiang Capital, and CATL; new investors reported to be in talks include CPE, Legend Capital, and Stony Creek Capital, a semiconductor-focused private-equity firm, alongside funds backed by chipmaker GigaDevice and state investment vehicles from Hefei. Reported use of proceeds is roughly a gigawatt of added compute capacity. On the listing, the reporting firms up the venue and the clock: an IPO filing on Shanghai's STAR Market as early as the end of 2026, with a market debut targeted for 2027. CNBC covered the same week how High-Flyer affiliates have taken pre-IPO allocations in CXMT and Unitree, tying Liang's widening balance-sheet activity to sectors Beijing treats as strategic. All sources are unnamed and DeepSeek has not commented. That end-of-August window has now passed without an announced closing — as of September 8, 2026 the most recent reporting still describes the round as nearing completion rather than done, and the August 26 SCMP / WSJ stories remain the newest on it.

The same week put the first hard revenue figures on the table. On August 26, 2026, Reuters — carried by The Standard — relayed a report from The Information that DeepSeek booked roughly 475 million yuan (~$70.7 million) in revenue over the first seven months of 2026, about ten times its full-year 2025 revenue of roughly 47.5 million yuan. The same report puts the seven-month net loss at about 715 million yuan — more than the revenue it sits against — versus a 935 million-yuan net loss for all of 2025, and splits the margins: 82.9% gross margin on the API business against 44.6% company-wide. For scale, the report set those against OpenAI's 39% first-quarter gross margin and Anthropic's, which the outlet said is expected to reach 63% for the year. An earlier report the same summer had put DeepSeek's annualized run-rate at $400–500 million. The Information also reported that DeepSeek hired investment banks to prepare a Shanghai listing for next year immediately after closing the June round — the firmest sourcing yet on the IPO track. DeepSeek did not respond to the outlet's request for comment, and none of these figures are audited or company-published.

Where to run DeepSeek

DeepSeek is widely deployed because the weights are open and the API is OpenAI-compatible. Inference paths through 2025–2026 break into four categories.

DeepSeek's own API. The first-party endpoint at api-docs.deepseek.com is OpenAI-API-compatible, so any OpenAI SDK can be pointed at it with only a base-URL change; it also speaks the Anthropic message format and, since the V4 GA wave, the OpenAI Responses API. Pricing has historically been an order of magnitude cheaper than Western frontier-model APIs (V4-Flash at $0.14 / M input tokens at launch), though the August 16, 2026 move to peak / off-peak billing raised every rate — see the V4 rows above for the current numbers.

Self-host from HuggingFace. Download from the deepseek-ai org and run with vLLM, SGLang, llama.cpp, or Ollama. The full V3 / V3.1 / V3.2 / V4-Pro models require multi-node H100 / H200 / B200 deployments at full precision; quantized variants ship from the open-source community shortly after each release. On June 27, 2026 DeepSeek released DSpark — a speculative-decoding drafter that, per DeepSeek's own benchmarks, speeds up V4 per-user generation ~60–85% on V4-Flash and ~57–78% on V4-Pro with no change to the model — alongside DeepSpec, an MIT-licensed codebase for training and evaluating speculative-decoding draft models on open targets (Qwen3, Gemma). These are serving optimizations rather than new DeepSeek models, so neither is a table row; the numbers are DeepSeek-reported and not yet independently verified. As of the July 31, 2026 open-weights drop, DSpark is no longer a separate download for the current flagship — the official DeepSeek-V4-Flash-0731 checkpoint ships with the drafter built in, enabled in vLLM and SGLang with a single speculative-decoding flag and no separate draft-model path. The V4-Pro GA weights followed on August 13, 2026 under the same MIT terms, with a vLLM recipe linked from the model card. DeepSeek also open-sourced its own agent scaffold, DeepSeek Harness, the same day — the plugin-based framework its published agent benchmarks are run in. The multimodal V4-Flash-Vision-Exp weights landed on August 31, 2026, also MIT, with SGLang and vLLM recipes added to the card over the following day; that release broke the same-day pattern the two GA builds had set, trailing its API debut by ten days.

Hosted-inference providers. Together AI, Fireworks AI, OpenRouter, SiliconFlow, Groq, Perplexity's public-API tier. Most providers serve the post-MIT weights (R1 forward) and clearly label which version is hosted. This tier tracks DeepSeek closely: OpenRouter's catalog carries the entire V4 line — the April preview builds, the V4-Flash-0731 and V4-Pro-0813 GA builds, and V4-Flash-Vision-Exp — each listed within a day of its API release.

Hyperscalers. AWS Bedrock, Microsoft Azure AI Foundry, NVIDIA NIM, IBM watsonx, and Oracle OCI have all added DeepSeek SKUs across 2025–2026, and Google Cloud's Vertex AI Model Garden has tracked the line closely rather than lagging it — R1 and V3 in preview in February 2025, then V3-0324, V3.1, V3.1-Terminus, V3.2-Exp, DeepSeek-OCR, and V3.2 as each shipped, with V3.2 a fully-managed API since December 10, 2025. What no hyperscaler catalog carries yet is V4: checked on September 8, 2026, Bedrock's DeepSeek model cards top out at V3.2, Azure AI Foundry's featured-model list at V3-0324, and Vertex's release notes at V3.2 / V3.2-Speciale. On this family the hyperscalers run a release generation behind the hosted-inference providers; check the catalogs for the current state.

People who shaped DeepSeek

Liang Wenfeng — founder and CEO of DeepSeek, co-founder and CEO of High-Flyer. The 2021 GPU-stockpile decision, the July 2023 DeepSeek incorporation, the V2 MoE / MLA bet, and the 2025 MIT-licensing turn all trace through Liang's office. Profiled in Fortune; on Wikipedia at Liang Wenfeng.

High-Flyer (Hangzhou Huanfang Technology Co., Ltd.) — the parent quantitative hedge fund. Co-founded by Liang in February 2016; reported to be using AI exclusively for trading by 2021. The funder of the Fire-Flyer 2 cluster and DeepSeek's only investor until the first external round closed in mid-2026.

DeepSeek's research staff — the lab is known for an unusually flat structure, a young research team (many recent PhD graduates from Tsinghua, Peking University, and Zhejiang University), and a publication culture that ships technical reports alongside model releases. Named-author rosters appear on the V2 / V3 / R1 / V3.2 papers on arXiv. Several core researchers have reportedly been recruited away to ByteDance, Tencent, Xiaomi, and the autonomous-driving company Yuanrong Qihang during 2025; named-departure tracking is sparse compared to U.S. labs.

No publicly-named CTO, CEO-second, or board. Unlike OpenAI, Anthropic, xAI, and Google DeepMind, DeepSeek does not maintain a leadership page; corporate governance is held inside the High-Flyer / DeepSeek Liang-led structure. The 2026 first-external round was deliberately structured to preserve that concentration — investors hold limited-partnership interests with no voting rights, with only the state-backed National AI Industry Investment Fund taking a direct, voting stake — so it is unlikely to produce a conventional outside-investor board.

The competitive landscape

DeepSeek is, alongside Alibaba's Qwen line, one of the two dominant Chinese open-weights AI families through 2025–2026. The closest direct comparators on the open-weights axis are Alibaba's Qwen (also Apache-2.0-or-permissive across most releases, with strong HuggingFace-leaderboard presence — see Qwen Versions), Mistral (French; mixed Apache 2.0 / Mistral Research License / proprietary tiers across the line, with the December 2025 “Mistral 3” family relaunch re-committing the open releases to Apache 2.0 — see Mistral Versions), Meta's Llama (custom Llama Community License, see Llama Versions), and Moonshot AI's Kimi line. The closed-weights frontier competitors — ChatGPT, Claude, Gemini, Grok — are the practical benchmark for “is DeepSeek competitive at frontier scale,” which is the question the V3 / R1 / V3.2 / V4 release cycle has been answering in the affirmative since December 2024. This page does not attempt a benchmark roundup or a ranking.

Use DeepSeek

The browser cannot detect which DeepSeek model you've used — there's no fingerprint or header that exposes it. The block below carries the practical information instead: the current model identifiers, a copy-paste API call, the surfaces where DeepSeek is available, and the licensing summary.

Current model identifiers

DeepSeek API model strings on the left; HuggingFace ids on the deepseek-ai org on the right. Verify against api-docs.deepseek.com and huggingface.co/deepseek-ai for the freshest list.

# V4 — current frontier line (April 2026, GA August 2026)
deepseek-v4-pro     # frontier; serves DeepSeek-V4-Pro-0813, GA 2026/08/13
deepseek-v4-flash   # serves DeepSeek-V4-Flash-0731, public beta 2026/07/31
deepseek-v4-flash-vision-exp  # experimental multimodal, 2026/08/21
deepseek-ai/DeepSeek-V4-Pro-0813    # GA weights, 2026/08/13
deepseek-ai/DeepSeek-V4-Flash-0731  # GA weights, 2026/07/31 (DSpark attached)
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp  # multimodal weights, 2026/08/31
deepseek-ai/DeepSeek-V4-Pro         # the April preview weights
deepseek-ai/DeepSeek-V4-Flash       # the April preview weights

# Legacy aliases — retired 2026/07/24 15:59 UTC per the published schedule
deepseek-chat       # was the non-thinking mode of deepseek-v4-flash
deepseek-reasoner   # was the thinking mode of deepseek-v4-flash

# V3.2 — still widely served (December 2025)
deepseek-ai/DeepSeek-V3.2
deepseek-ai/DeepSeek-V3.2-Speciale

# V3.1 + Terminus — hybrid Thinking / Non-Thinking architecture (August/September 2025)
deepseek-ai/DeepSeek-V3.1
deepseek-ai/DeepSeek-V3.1-Terminus

# R1 line — reasoning
deepseek-ai/DeepSeek-R1
deepseek-ai/DeepSeek-R1-0528

# Specialized — Coder, Math, multimodal
deepseek-ai/DeepSeek-Coder-V2-Instruct
deepseek-ai/Janus-Pro-7B

Quick API call (OpenAI-compatible)

DeepSeek's API endpoint is OpenAI-API-compatible — point any OpenAI SDK at the DeepSeek base URL with a DeepSeek API key. Replace the placeholder values before running.

$ curl https://api.deepseek.com/chat/completions \
    -H "Authorization: Bearer $DEEPSEEK_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "model":    "deepseek-v4-flash",
      "messages": [{ "role": "user", "content": "Hello, DeepSeek." }]
    }'

Where to run DeepSeek

Four categories — DeepSeek's own API, self-host from HuggingFace, hosted-inference providers, and hyperscalers. Pricing varies by orders of magnitude; the open weights are the same across all of them.

# DeepSeek first-party
https://chat.deepseek.com/                  # consumer chat (Expert / Instant modes)
https://api-docs.deepseek.com/              # OpenAI-compatible API

# Self-host from HuggingFace
https://huggingface.co/deepseek-ai          # every model card lives here
https://github.com/vllm-project/vllm        # production-grade throughput
https://github.com/sgl-project/sglang
https://github.com/ggerganov/llama.cpp      # CPU + GPU, edge-friendly
https://ollama.com/                         # single-binary, easiest entry

# Hosted-inference providers
https://www.together.ai/
https://fireworks.ai/
https://openrouter.ai/
https://groq.com/
https://www.siliconflow.com/

# Hyperscalers
AWS Bedrock, Azure AI Foundry, NVIDIA NIM, IBM watsonx, Oracle OCI

Licensing

DeepSeek transitioned from a bespoke “DeepSeek License” on the model weights to MIT starting with R1 (January 2025). Read the LICENSE-MODEL file in the relevant GitHub repo before shipping at scale.

# MIT-licensed (R1 forward, January 2025+)
DeepSeek-R1, R1-0528, R1 distilled
DeepSeek-V3-0324, V3.1, V3.1-Terminus, V3.2-Exp, V3.2, V3.2-Speciale
DeepSeek-V4-Pro, V4-Pro-0813, V4-Flash, V4-Flash-0731
DeepSeek-V4-Flash-Vision-Exp (weights 2026/08/31)
Janus-Pro 1B / 7B
DeepSeek-OCR (October 2025)

# Apache 2.0 (first non-MIT modern release: Math-V2)
DeepSeek-Math-V2 (November 2025)
DeepSeek-OCR-2 (January 2026)

# Bespoke "DeepSeek License" on model weights, MIT on code
DeepSeek-LLM 7B / 67B
DeepSeek-Coder, Coder-V2
DeepSeekMath
DeepSeek-VL, VL2
DeepSeek-V2, V2.5
DeepSeek-V3 (original December 2024 release; relicensed to MIT as V3-0324)

# LICENSE-MODEL files live in each GitHub repo
https://github.com/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL
https://github.com/deepseek-ai/DeepSeek-V2/blob/main/LICENSE-MODEL

Sources: DeepSeek API docs; github.com/deepseek-ai; huggingface.co/deepseek-ai; research papers on arXiv (V2, V3, R1, V3.2, Janus-Pro); contemporaneous reporting in NYT, WSJ, FT, Bloomberg, CNBC, The Information, Fortune, Reuters, TechCrunch, Yahoo Finance, and Nature. Last updated September 8, 2026.

Mungomash LLC · More AI pages