2023 – 2026
Qwen Versions
The latest generally available Qwen flagship is Qwen3.8-Max (August 3, 2026) — a sparse-MoE model with 2.4 trillion parameters, 95 billion active, native vision-language input, a 1,000,000-token context window, and hybrid thinking on by default. It sits first on Qwen Cloud's capability-ordered model list, ahead of the mid-tier Qwen3.7-Plus (May 31, 2026) and the cost tier Qwen3.8-Flash (August 26, 2026); the prior Max-tier flagship, Qwen3.7-Max (May 20, 2026), is still served but no longer leads. The promise Alibaba made at the August 3 launch has now landed: on August 12, 2026 it published Qwen/Qwen3.8-2.4T-A95B, the first open-weights release of a Qwen-Max-class model and the largest open-weights language model published to date — though under a bespoke Qwen3.8-Max License with revenue-share conditions rather than Apache 2.0, and text-only (the hosted qwen3.8-max keeps the vision input). Two days later, on August 14, 2026, the companion Qwen3.8-27B dense checkpoint shipped under plain Apache 2.0. The newest open-weights model in the flagship line is Qwen3.8-Flash-Next (August 26, 2026), a 125B model activating 6B parameters that Alibaba published as an explicit early preview of the architecture Qwen4 will be built on — and which carries a fifth license, the Qwen Community License 1.0. The newest open-weights release of any kind is Qwen-Drive-1.0-4B (August 28, 2026), the line's first autonomous-driving model, under plain Apache 2.0. I track every Qwen / Tongyi Qianwen release here — from Qwen-7B in August 2023 onward — with HuggingFace ids, ship dates, family (Flagship / Reasoning / Specialized), and license terms (Apache 2.0 / Tongyi-Qianwen / the new Qwen3.8-Max License / proprietary). Below the table: the April 2023 Tongyi launch as Alibaba's ChatGPT response, the licensing turn at Qwen2, the QwQ reasoning track, the Qwen3 hybrid-reasoning era, the U.S. chip export-control context, and the HuggingFace-leaderboard dominance through 2025–2026.
The April 2023 Tongyi Qianwen launch
Alibaba Cloud formally launched Tongyi Qianwen (通义千问, “truth from a thousand questions”) in April 2023 as the company's response to ChatGPT, two months after Baidu's Ernie Bot launched and roughly five months after OpenAI's November 2022 ChatGPT release. The model was first demonstrated by then–Alibaba CEO Daniel Zhang at the Alibaba Cloud Summit on April 11, 2023 and rolled out to enterprise customers through the Tongyi product family on Alibaba Cloud.
The first open-weights release — Qwen-7B — followed on August 3, 2023. The four-month gap between the Tongyi consumer-product launch and the first open-weights release was characteristic of Alibaba's strategy: ship the proprietary chatbot first to enterprise customers via Alibaba Cloud, then open-source the underlying model line to developer communities for ecosystem effects. The same hybrid pattern ran unbroken until August 2026: open-weights flagships on HuggingFace alongside a closed Qwen-Max line on Alibaba's hosted API. That ended on August 12, 2026, when the Max line shipped weights for the first time — on its own license, and without the vision input the hosted model keeps.
The Apache 2.0 turn — from Tongyi-Qianwen License to permissive open-source
Qwen's licensing has moved through three eras — a bespoke house license, a broad Apache 2.0 turn, and a 2026 return to bespoke terms at the frontier — and it has never been as tidy inside an era as the era labels suggest. The Qwen 1 lineage (Qwen-7B / 14B / 72B, August–November 2023) shipped under the bespoke Tongyi-Qianwen License — an Alibaba-authored license with permissive terms for academic use and commercial use with restrictions, but not OSI-compliant. The Qwen 1.5 family in February 2024 continued the same pattern at 7B, 14B, 72B, 110B, and MoE-A2.7B. What is easy to miss, because it inverts the intuition that smaller weights come with fewer strings, is that the small checkpoints in both waves are the restricted ones: Qwen-1.8B, and Qwen1.5 at 0.5B, 1.8B, and 4B, ship under a separate Tongyi Qianwen RESEARCH License that grants no commercial rights at all. There is no size rule to lean on here; the per-checkpoint license is the answer (verified across all of them September 4, 2026). And the card is not always that answer: Qwen1.5-32B's base checkpoint is tagged research-only on HuggingFace while the LICENSE file in the same repo is the ordinary commercial agreement, byte-identical to its -Chat sibling's. When the tag and the license text disagree, read the license text.
The licensing turn arrived with Qwen2 on June 6, 2024. Most of the Qwen2 sub-family (0.5B, 1.5B, 7B, 57B-A14B) shipped under Apache 2.0 — the first Qwen flagship release with broad permissive coverage. Only Qwen2-72B retained the Tongyi-Qianwen License. The pattern continued through Qwen2.5 (September 19, 2024), where the 0.5B / 1.5B / 7B / 14B / 32B variants shipped Apache 2.0 while the 72B retained the commercial-with-restrictions Qwen License and the 3B — alone in the middle of an otherwise permissive range — took the stricter Qwen Research terms, with no commercial grant. The same 3B carve-out repeats across Qwen2.5-VL, Qwen2.5-Coder, and Qwen2.5-Omni.
Qwen3 (April 28, 2025) was the structural commitment: the entire Qwen3 family — six dense sizes from 0.6B to 32B, plus the 30B-A3B and 235B-A22B MoE variants — shipped Apache 2.0. Every open-weights Qwen flagship release since (Qwen3-Coder, Qwen3.5, Qwen3.5-Plus's open-weights variants, Qwen3.6-27B, Qwen3.8-27B, Qwen-Image-2512) has shipped Apache 2.0. The proprietary Qwen-Max-line (Qwen2.5-Max and Qwen3-Max through Qwen3.7-Max / Qwen3.7-Plus, with their hosted Plus / Flash siblings) runs as a parallel commercial track on DashScope, but the open-weights story has been Apache-2.0-or-permissive across every release since Qwen2. In July 2026 the specialized track picked up a hosted-only wrinkle: Qwen-Audio-3.0-Realtime (July 15, 2026), Qwen-Audio-3.0-TTS (July 20, 2026), Qwen-Image-3.0 (July 21, 2026), and Qwen-Audio-3.0-ASR-Flash (July 30, 2026) all shipped as API- or chat-only services with no license text and no weights — the first specialized lines to debut closed while their open-weights predecessors (Qwen3-TTS, Qwen3-Omni, Qwen-Image-2.0, Qwen3-ASR) stayed Apache 2.0.
In August 2026 the flagship track moved the other way, and added two more license conventions doing it. Announcing Qwen3.8-Max on August 3, 2026, Alibaba wrote that the launch “marks the first time we will open-source the weights of a Qwen-Max-class model” (Qwen3.8-Max: A New Bar for Coding and Cowork). The weights arrived on August 12, 2026 as Qwen/Qwen3.8-2.4T-A95B, ending six generations of an entirely closed Max line running from Qwen2.5-Max in January 2025 — but not under Apache 2.0. The repo carries a bespoke Qwen3.8-Max License: broad rights to use, modify, host, fine-tune, and sell, conditioned on prominent model-name attribution for products above 100 million MAU or $20 million monthly revenue, and on obtaining a separate commercial license if the licensee runs a model-as-a-service or AI-work-assistant business above $50 million of trailing-twelve-month group revenue. The checkpoint is also text-only; vision input stays behind the hosted API.
Two days later, on August 14, 2026, Qwen3.8-27B shipped under plain Apache 2.0 — so the split now runs by size rather than by track: the deployable sizes stay permissive, the frontier-scale checkpoint carries revenue-share terms. HuggingFace's State of Open Models: Summer 2026 (August 14, 2026) reads it as a sector-wide turn rather than an Alibaba quirk: of 178 Chinese releases above 20B parameters in 2026, 59% are Apache 2.0 and 22% MIT with almost none carrying non-commercial restrictions, “however, in the last few weeks, we started to see a change on this trend for the really large models, with Kimi K3 and Qwen 3.8 2.4T starting to include some non-commercial restrictions and revenue share requirements to their licenses.”
Twelve days after that report, the size-band reading stopped holding. Qwen3.8-Flash-Next (August 26, 2026) is a 125-billion-parameter model activating 6 billion — a size squarely in the band HuggingFace had just described as staying permissive, and one people actually self-host — and it shipped under a fifth convention, the Qwen Community License 1.0. The text is a near-copy of the Qwen3.8-Max License, with one substantive change that runs the stricter way: where the Max license triggers its separate-commercial-license requirement only once a model-as-a-service or AI-work-assistant licensee clears $50 million of trailing-twelve-month group revenue, the Community license drops the revenue threshold entirely and applies to any such business at any size. The 100-million-MAU / $20-million-monthly-revenue attribution condition and the internal-use carve-out are identical. So the line running through Alibaba's August is not really about parameter count — it is that Qwen now reserves the resale-inference and coding-assistant markets on its newest architectures, at every scale it has released them, while the established generations (Qwen3.8-27B, Qwen3.6, Qwen3.5, Qwen3) stay plain Apache 2.0. Two days after that, Qwen-Drive-1.0-4B (August 28, 2026) shipped under plain Apache 2.0 — and it is built on Qwen3.5-4B, which fits the reading rather than breaking it: the bespoke terms so far attach to the newest architectures, not to derivative work on the settled ones.
The Qwen3 hybrid-reasoning era — April 28, 2025
Qwen3 launched on April 28, 2025 as Alibaba's first frontier-AI line with hybrid reasoning architecture. The release shipped eight models simultaneously — six dense (0.6B, 1.7B, 4B, 8B, 14B, 32B) and two MoE (30B-A3B and the flagship 235B-A22B) — all under Apache 2.0, all sharing a single architecture that supports both thinking-mode chain-of-thought reasoning and non-thinking-mode fast responses. The announcement is at qwenlm.github.io/blog/qwen3; coverage in TechCrunch and Alibaba Cloud Community.
The Qwen3 architecture absorbed the standalone QwQ reasoning track into the Flagship V-series. The QwQ line had run for five months — QwQ-32B-Preview in November 2024, the production QwQ-32B in March 2025 — as Alibaba's open-weights answer to OpenAI's o1 series and DeepSeek's R1. Qwen3's hybrid-mode architecture made the standalone Reasoning family redundant; no further QwQ releases have shipped in the year since, and I don't expect to add another Reasoning row here unless Alibaba revives the standalone track.
Qwen3 was trained on 36 trillion tokens — double Qwen2.5's pretraining corpus — with native multilingual support across 119 languages and dialects, the broadest language coverage of any frontier-AI line at the time. The Qwen3.5 family released ten months later (February 2026) extended that to 201 languages and added native multimodality across text + image + video; Qwen3.6-27B (April 2026) added the Gated DeltaNet hybrid architecture and Thinking Preservation. The Qwen3 architecture and its descendants remain the load-bearing recipe for everything Alibaba has shipped since.
The HuggingFace-leaderboard dominance
Across late 2024 and most of 2025, Qwen-derived models held a majority of the top-trending positions on the HuggingFace Open LLM Leaderboard and the various community-derived benchmark trackers. The pattern was driven by a combination of three factors: the Apache 2.0 license (which let community fine-tuners commercially redistribute their derivatives), the breadth of base sizes (0.5B through 235B in the Qwen3 family alone), and the strength of Qwen2.5-Coder / QwQ / Qwen3 on math + coding benchmarks specifically (which weighted heavily in the leaderboard rankings). At various points in 2025, more than half of the top 20 trending HuggingFace models were Qwen-derived. The dominance shifted slightly in 2026 as DeepSeek-V4 / Llama 5-equivalents / Mistral 3 absorbed share, but Qwen-derivatives have remained a large fraction of the top-trending HuggingFace models through August 2026.
HuggingFace put numbers on it in State of Open Models: Summer 2026 Observations (August 14, 2026): Qwen-based models account for 151,448 derivatives on the Hub — 2.6× Meta's total footprint and 4.7× the Llama repositories specifically, with Google second at 82,506 — growing at roughly 180–210 new repositories a day across the first seven months of 2026. HuggingFace credits three things: a steady release cadence, coverage across sizes, and Apache 2.0. The third-largest derivative source is Unsloth, a community account publishing quantized builds, most of them Qwen. Alibaba's own reading of the same report (August 17, 2026) claims 460-plus open-sourced models, 300,000-plus derivatives and 3 billion cumulative downloads — a broader count than HuggingFace's Hub-only figures, and one the vendor is doing itself.
The U.S. chip export-control context
Qwen training has been constrained by the same U.S. chip export-control regime that shapes DeepSeek's training environment. The October 7, 2022 Department of Commerce export controls restricted top-tier Nvidia AI-GPU exports to China; Alibaba Cloud was already on various U.S. Entity List adjacencies prior to that, with subsequent expansions through 2023–2025 tightening the procurement environment. Like DeepSeek, Alibaba Cloud built Qwen's training infrastructure on a mix of Nvidia H800 chips procured during the gap before the October 2023 H800 ban, and on Chinese domestic alternatives (Huawei Ascend, Cambricon).
The empirical record — trillion-parameter Qwen3-Max trained on 36T tokens, Qwen3.6-27B with state-of-the-art coding benchmarks — demonstrates that the Qwen team has continued training at frontier scale despite the controls, and like DeepSeek has been the subject of Department of Commerce inquiries about the chip procurement that supported specific runs. Coverage in CSIS and South China Morning Post covers the broader policy environment.
At the May 2026 Alibaba Cloud Summit — the same event that launched Qwen3.7-Max — Alibaba pushed further toward domestic alternatives, unveiling the Zhenwu M890, a custom AI accelerator from its T-Head semiconductor subsidiary (144 GB on-chip memory, 800 GB/s interchip bandwidth, ~3× the performance of the prior Zhenwu 810E) purpose-built for agent workloads, alongside the Panjiu AL128 server packing 128 accelerators per rack. Alibaba outlined a multi-year in-house silicon roadmap (a V900 successor in Q3 2027 and a J900 in Q3 2028) and said T-Head has shipped more than 560,000 Zhenwu units to 400+ customers to date. The company has pledged more than 380 billion yuan (~$53 billion) across cloud and AI infrastructure over three years. Coverage in South China Morning Post, CNBC, and Reuters / Business Standard (May 20, 2026).
On June 8, 2026 the U.S. Department of Defense published an updated Section 1260H list of "Chinese military companies," adding 65 entities — 17 new parent-level listings and 48 subsidiaries — and naming Alibaba for the first time, alongside Baidu, BYD, NIO, Unitree, and TP-Link. The DoD described Alibaba and Baidu as contributors to the Chinese defense industrial base. The 1260H list is distinct from the Commerce Department's Entity List and imposes no export licensing requirement of its own; its effect is that the Defense Department is barred from contracting directly with listed companies, and from procuring their products or services through third parties beginning in June 2027. Alibaba disputes the designation and has taken it to court: on June 23, 2026 it sued the Defense Department in the U.S. District Court for the Northern District of California (Alibaba Group Holding Ltd. v. U.S. Department of Defense, No. 5:26-cv-06227), saying it does not work with the Chinese military and asking to be removed from the list. A companion provision — barring the Pentagon from working with any company whose lobbyists also represent a 1260H entity — took effect at the end of June and led all of Alibaba's registered lobbyists to withdraw; on July 5, 2026 Judge Eumi K. Lee ordered the department not to apply that lobbying restriction to Alibaba until she rules on the company's motion or 60 days after a hearing on it, whichever comes first. That order is temporary and does not touch the listing itself, which stood at this writing. Coverage in CNBC (June 9, 2026), a client alert from WilmerHale (June 11, 2026), and Bloomberg via Fortune (July 5, 2026); the docket is tracked at the Civil Rights Litigation Clearinghouse. At WAIC 2026 (July 20, 2026) Alibaba's T-Head unit open-sourced its SAIL AI software stack, optimized for the Zhenwu accelerators, continuing the domestic-silicon push.
The export-control story gained a China-side counterpart in July 2026. Per Financial Times reporting on July 20, 2026 (picked up by Reuters; see also Tom's Hardware), China's Ministry of Commerce has consulted leading AI and chip firms — Alibaba, ByteDance, and Z.ai among them — on possible export controls covering their most advanced AI models, including whether foreign users should be allowed to download model weights and whether key training data may be transferred overseas, alongside limits on foreign manufacture of advanced chips designed by Huawei, Alibaba, and ByteDance. The measures under discussion would ride the next revision of China's catalogue of technologies restricted from export; nothing has been decided, and regulators were still gathering industry feedback at this writing. A restriction on foreign downloads of open-weights models would apply directly to the Apache 2.0 Qwen releases this page tracks — I'll record whatever rule actually lands.
Where to run Qwen
Qwen is among the most widely-deployed AI lines because the open-weights releases are Apache 2.0 across nearly every size and the proprietary releases are available through Alibaba Cloud's Model Studio with OpenAI- and Anthropic-compatible APIs. Inference paths through 2025–2026 break into four categories.
Alibaba Cloud first-party. Qwen Chat is the consumer chat surface. Model Studio (formerly DashScope) is the long-standing developer API endpoint, OpenAI-API-compatible and serving both the open-weights and the proprietary Max-line models. The proprietary Max-line is exclusive to this surface. Since May 26, 2026 it has a front end built for agents: Alibaba Cloud launched Qwen Cloud for international markets in Singapore, an AI-native platform aggregating 150-plus model APIs — Qwen plus third-party lines including DeepSeek, GLM, Kimi, Wan, and HappyHorse — behind one API key, with three entry points (agent-readable Skills, a CLI, and a website) and a per-model marketplace listing price, context, and rate limits (Alibaba Cloud Community, May 28, 2026). Its model-releases changelog is now the fastest-updating first-party record of what has shipped. The older Model Studio docs lagged it badly through the summer, and unevenly: the recommended-models page sat frozen at a July 15, 2026 stamp for four weeks after Qwen3.8-Max went generally available, and the billing page stayed frozen at that stamp for six. Both have since caught up, and by September 2, 2026 both carried a same-day stamp along with the qwen3.8-max and qwen3.8-flash strings — so all three first-party surfaces are current again. The episode is a reminder that a frozen doc page answers “no” to every question you ask it, so the stamp is worth reading before the content. The underlying HTTP endpoint is unchanged — dashscope-intl.aliyuncs.com/compatible-mode/v1. A third first-party surface arrived on August 3, 2026: QwenWork, a workplace agent platform that entered public beta in China offering Qwen3.8-Max as its Flagship model tier and slated for embedding in DingTalk; an international edition followed on August 26, 2026. One consumer surface sits outside Alibaba's own apps: on July 15, 2026 the Cyberspace Administration of China approved Apple Intelligence for launch in China with Qwen as its system-level language engine across iOS / iPadOS / macOS / visionOS (TechCrunch, July 16, 2026).
Self-host from HuggingFace. Download from the Qwen org and run with vLLM, SGLang, llama.cpp, or Ollama. The Apache 2.0 open-weights flagships (Qwen3.8-27B, Qwen3.6-27B, Qwen3.5 family, Qwen3 family, QwQ-32B) self-host without commercial restriction. Three smaller sets carry conditions: the Qwen License variants (Qwen2-72B, Qwen2.5-72B, Qwen2.5-3B) require attestation of the bespoke terms; Qwen3.8-2.4T-A95B — the open Max checkpoint, and at 2.4 trillion parameters not a self-host most people will attempt anyway — requires a separate commercial license above $50 million of trailing-twelve-month revenue if you are reselling inference; and Qwen3.8-Flash-Next, which at 125B total and 6B active is a realistic self-host, requires that separate license for any model-as-a-service or AI-work-assistant business with no revenue floor at all. Running it for your own internal use is explicitly carved out of both.
Hyperscalers. AWS Bedrock and Azure AI Foundry have added Qwen SKUs across 2025–2026; ModelScope (Alibaba's own model-hub) hosts the broadest set. NVIDIA NIM has Qwen variants for the most-served sizes.
Hosted-inference providers. Together AI, Fireworks, OpenRouter, SiliconFlow, Groq. Most providers serve the Apache 2.0 lineage with similar latency / cost characteristics; the Tongyi-Qianwen-License variants (mostly the 72B and 3B sizes) are typically not carried by Western inference providers due to license-attestation overhead.
People who shaped Qwen
The Qwen / Tongyi Lab team is structured inside Alibaba Cloud rather than as a standalone lab. Junyang Lin (Lin Junyang) ran the Qwen project from the Tongyi Lab's formation in late 2022 — through the April 2023 Tongyi Qianwen launch and every release up to the Qwen3.5 small-model wave — and then left abruptly, posting “me stepping down. bye my beloved qwen.” on X in the early hours of March 4, 2026, with March 7 as his last day. Several Qwen engineers went in the same window, including post-training lead Yu Bowen. Alibaba Cloud CTO Jingren Zhou (Zhou Jingren), who built the Tongyi Lab and set the open-source strategy, took over oversight of the team, and Zhou Hao — previously a senior staff research scientist at Google DeepMind — joined as head of post-training research. Reporting attributed the split to disagreements over lab structure (Lin favored keeping pretraining and post-training integrated, against a reorganization separating them) and over the Qwen team's control of its own AI infrastructure: South China Morning Post (March 4, 2026) and Caixin Global (March 10, 2026). The release cadence did not slow: Qwen3.6-35B-A3B, Qwen3.6-27B, the Qwen3.7 generation, Qwen-AgentWorld, and the July 2026 wave all shipped after the departure.
Eddie Wu — CEO of Alibaba Group since September 2023; has publicly framed AI as Alibaba's strategic priority above e-commerce / cloud / logistics, with multi-year capex commitments to Qwen training infrastructure. He took a more direct hand after Lin's exit: a foundation-model task force under Wu stood up on March 5, 2026, and on March 16 Alibaba folded its AI teams and products into a new Token Hub Business Group, also led by Wu, sitting alongside Alibaba Cloud and the e-commerce divisions. The reshuffling continued: on June 8, 2026 Alibaba merged the Tongyi Lab with the Future Life Lab (the Taobao-and-Tmall team behind the Happy Horse video model and the Happy Oyster world model) into a new Token Foundry unit, again reporting directly to Wu, while Jingren Zhou moved up to Alibaba group chief scientist — the highest academic title in Alibaba's technology system — and took charge of a newly created AI Future Research Institute focused on frontier research; Zheng Bo, who led the Future Life Lab, brought that team into the merged unit. Sources: South China Morning Post and 36Kr (June 8–9, 2026). No further reshuffle or named-lead change had been reported as of September 8, 2026. Joseph Tsai — Chairman of Alibaba Group; has been the public spokesperson for Alibaba's AI strategy in international forums (the Davos annual meetings, the Bloomberg Tech Summit).
No publicly-named Qwen CTO or founding team in the Western-lab sense. Unlike OpenAI / Anthropic / Mistral / xAI / DeepSeek, Qwen is not structured as a startup-style lab with named founders and a public-facing leadership roster. The team operates under Alibaba Cloud's organizational umbrella, and the per-paper author lists on the Qwen / Qwen2 / Qwen2.5 / Qwen3 technical reports are the closest available roster of named contributors.
The competitive landscape
Qwen is, alongside DeepSeek, one of the two dominant Chinese open-weights AI families through 2024–2026. The closest direct comparators on the open-weights axis are DeepSeek (Chinese, MIT-licensed for the V3 / R1 line and onward, the December 2024 / January 2025 inflection — see DeepSeek Versions), Mistral (French; Apache 2.0 for the Mistral 3 family with a parallel proprietary tier — see Mistral Versions), Meta's Llama (custom Llama Community License, see Llama Versions), and the other Chinese frontier labs (Baidu Ernie, Zhipu GLM, MiniMax, Moonshot Kimi). The closed-weights frontier competitors — ChatGPT, Claude, Gemini, Grok — are the practical benchmark for “is Qwen competitive at frontier scale,” which the Qwen3 / Qwen3-Max / Qwen3.5 / Qwen3.6 release cycle has been answering in the affirmative since April 2025. Qwen's distinguishing variable is the breadth of its specialized track (Coder / VL / Audio / Math / Omni / Image) and the consistency of its Apache 2.0 commitment on the open-weights flagships, both of which continue to underwrite the line's HuggingFace-leaderboard dominance. This page does not attempt a benchmark roundup or a ranking.