2023 – 2026

Llama Versions

Meta's current frontier-AI model is Muse Spark 1.3, released September 2, 2026 under Meta Superintelligence Labs — closed-weights, tuned for long-horizon agentic and coding work, with a 1M-token context window. It is callable through the Meta Model API and through the Muse Code terminal agent. Its top “max reasoning” mode, withheld at launch for safety testing, went live on September 4, 2026 — Meta edited the September 2 launch post in place to say so rather than publishing anything new — and it is where several of Meta's strongest published scores for the model come from. It is served on the standard tier only, not on the cheaper contributor tier. The day before, on September 1, 2026, the same lab shipped Muse Voice Transcribe, its first real-time audio-perception model — streaming speech recognition, speaker diarization and endpointing in one model, also closed-weights, billed per minute of audio. On August 10, 2026 it released Muse Glimmer 30B, a local agent model with open weights under Apache 2.0 — Meta's first downloadable frontier-lab model since Llama 4, and the first ever to ship without the Llama license's >700M-MAU and EU carve-outs. The Muse line had been closed-weights since its April 2026 launch, Meta's first frontier surface shipped that way since Llama 2 in July 2023; Glimmer partly reopens it, one size class below the flagship. Meta has now said three times — August 10, August 20, and again in the September 2 launch post — that an open-weights Muse Spark release is coming, with no date, no license named, and no weights published as of this update, so the flagship is still closed-weights today. The same lab shipped Muse Image, its first in-house image-generation model, on July 7, 2026, and previewed Muse Video alongside it; Muse Image reached the Meta Model API for outside developers on August 26, 2026 at $0.01 per image, still closed-weights. The open-weights Llama line itself runs from LLaMA 1 (February 2023, leaked the next month) through Llama 4 Maverick and Llama 4 Scout (April 2025). I track every Meta release here — with HuggingFace ids, parameter counts, context windows, and license terms — because the licensing and infrastructure terms change too often for marketing pages to keep up. Below the table: the March 2023 LLaMA leak, the FAIR vs. GenAI org split, the Llama 4 Behemoth delay, the June 2025 Scale AI acqui-hire and Meta Superintelligence Labs reorg, and the closed-weights turn.

Family & status

Family

Flagship — the main Llama chat lineage from LLaMA 1 through Llama 4
Specialized — Code Llama, Llama Guard safety classifiers, vision and edge variants
Successor — the Muse line, Meta's post-Llama frontier surfaces under Meta Superintelligence Labs; closed-weights on the API models so far, open weights under Apache 2.0 on Muse Glimmer

Status

Current — actively recommended; the latest in its family
Available — weights still served via HuggingFace and partner inference providers, but superseded
Legacy — deprecated, never publicly released, or research-only

Llama version table

Model
Muse Spark 1.3
muse-spark-1.3, muse-spark-1.3-contributor — Meta Model API; closed-weights, no HuggingFace release yet
Successor
Current
Sep 2, 2026
Current Meta frontier model. Agentic-workflow update to 1.2 — long-horizon multitasking, asks clarifying questions, ~20% fewer tool calls. 1M-token context. The top “max reasoning” mode, withheld at launch, went live September 4, 2026 on the standard tier. Audio input is still degraded relative to 1.2.
  • Released September 2, 2026; the announcement is at research.meta.ai/blog/introducing-muse-spark-1-3. Coverage in Axios.
  • An agentic-workflow update rather than a scale-up. Meta says the model sustains longer-horizon work in a single long thread: it generates its own context across messy and conflicting sources, corrects gaps in its own plan, and tracks what it has learned toward a final deliverable. It was trained across a diverse set of harnesses so the behavior generalizes beyond Meta's own agent. It also asks clarifying questions on ambiguous prompts, invokes the user when stuck, confirms before consequential actions, and maps incoming prompts to the right task when several are in flight at once.
  • Coding: fewer turns, less output. Trained on more long-horizon coding tasks; relative to 1.2 Meta describes it as less verbose and quicker to stop, and reports its own engineers measuring roughly 20% fewer tool calls and 25% fewer tokens on the same work (lab-reported; not independently verified here). Meta published an evaluation methodology report alongside it. Worth setting against that claim: with per-token prices unchanged, Artificial Analysis measured the cost of finishing an average task on its own suite rising from $0.40 on 1.2 to $0.55 on 1.3 as of September 3, 2026, which it attributes to heavier input-token consumption on agentic evaluations. The two measurements are not in direct conflict — Meta is describing its own coding workflows, the benchmarker a broader mix — but token count per task is not a number a price sheet settles.
  • It shipped without its top reasoning mode, and got it two days later. At launch Meta's wording was that the previously available reasoning modes were live “with max reasoning coming shortly after we finish additional safety testing,” so the setting you could actually call topped out one notch below, at xhigh. On September 4, 2026 that changed: Meta's reasoning docs now list "max" as a served reasoning_effort value — standard-tier muse-spark-1.3 only, never on the -contributor tier — and the model docs now describe 1.3 as supporting every effort level and being “recommended for new work.” Meta published nothing new to say so: it edited the September 2 launch post in place, which now reads “Muse Spark 1.3 with max reasoning is now available on Muse Code and Meta Model API,” while the post's own published and modified dates both still say September 2 (announcement). The three-day gap was not cosmetic. Meta discloses both configurations, and reports GDPval-AA v2 at 1,754 Elo for max against 1,709 for xhigh, OSWorld 2.0 at 66.9 against 57.2, and JobBench at 64.9 against 61.2 — while on Terminal-Bench 2.1 the lower setting is marginally ahead, 89.2 against 88.8. Third-party benchmarker Artificial Analysis, which through September 4 had measured max only in a limited partner preview and listed no API provider for it at all, now routes it through one provider — Meta's own — and on its Intelligence Index v4.3 scores max 48, xhigh 45, and Muse Spark 1.2 40 (re-checked September 8, 2026). VentureBeat walked through the split on September 3, while it was still a split.
  • And it lost a modality on the way in. Meta's model docs flag audio understanding on 1.3 as “currently not fully supported,” with degraded quality on requests containing audio, and point audio work back at Muse Spark 1.2 or at Muse Voice Transcribe. 1.2's row still lists audio without an asterisk. A new flagship that is worse than its predecessor on one input type is unusual enough to be worth stating plainly.
  • API shape: model ids muse-spark-1.3 and muse-spark-1.3-contributor, base URL https://api.meta.ai/v1, context window 1,048,576 tokens, text / image / video / audio / PDF in, text out. Pricing is unchanged from 1.2 — $1.25 / $4.25 per million input / output tokens on the standard tier, $0.15 cached input, and $0.10 / $0.20 on the -contributor tier that trains on your prompts (pricing docs). Also routed by OpenRouter as meta/muse-spark-1.3 from launch day.
  • Safety: Meta reports stronger adversarial robustness and better resistance to prompt injection, plus better calibration on what counts as an irreversible action during long-horizon agentic work (lab-reported; not independently verified here).
  • Still closed-weights: no HuggingFace release. The launch post restates the open-weights plan a third time — “an exciting roadmap lined up, including bigger models, the Muse Spark open weights release, and more” — still with no date, no license, and no named version. Mark Zuckerberg said the same thing in his own voice the same day, and in the plural — open-weights Muse Spark “releases” coming soon, again without a version (The Register). Re-checked September 8, 2026: huggingface.co/meta-models holds nothing beyond the four Muse Glimmer repos.
Model
Muse Voice Transcribe
muse-voice-transcribe-1.0 — Meta Model API; closed-weights, no HuggingFace release
Successor
Current
Sep 1, 2026
Meta Superintelligence Labs' first real-time audio-perception model. Streaming speech recognition, speaker diarization for 20+ speakers, and endpointing in one model rather than a pipeline. 25 validated languages with code-switching. Closed-weights, billed at $0.18 per hour of audio.
  • Released September 1, 2026; the announcement is at research.meta.ai/blog/introducing-muse-voice-transcribe. Coverage in 9to5Mac and MarkTechPost.
  • One model doing three jobs, not a pipeline. Streaming speech recognition is the base; speaker diarization and endpointing are layered on as extra output tokens rather than as separate post-processing stages — <|start_of_turn|> and <|speaker_A–Z|> for who is talking, <|speech_onset|> and <|speech_endpoint|> for when a turn begins and ends. All three are trained together.
  • An autoregressive model from the Muse Spark family, not a separate architecture. Audio arrives in 80 ms chunks (12.5 Hz), each collapsed to a single soft token; at every chunk the model chooses to keep listening or to emit text. Because it controls when it speaks, it sets its own lookahead per word — Meta calls this “adaptive delay” and trains it with reinforcement learning against a combined word-error-rate and latency reward.
  • Coverage. Trained on 70+ languages with 25 validated at launch, native code-switching inside a single sentence, keyword and context biasing, audio longer than an hour, and 20+ speakers with no post-processing pass. Meta reports first place on Artificial Analysis's streaming speech-to-text leaderboard and on public diarization benchmarks as of September 1, 2026 (lab-reported; not independently verified here).
  • API shape: model id muse-voice-transcribe-1.0; a realtime WebSocket at wss://api.meta.ai/v1/asr/realtime and a file upload at POST /v1/asr/transcribe, with three modes (PUSH_TO_TALK, ENDPOINTING, DIARIZATION). Turn-level timestamps only — no word-level timestamps, confidence scores, sound-event detection, or emotion detection, and no speech synthesis. Billed by audio processed at $0.18 per hour, the same rate streaming or batched, with no train-on-your-data discount tier at launch (docs).
  • Closed-weights, like every Muse model except Glimmer. Available the day of announcement through the Meta Model API, through Meta AI for Mac — where holding Fn dictates into any application — and inside Muse Code, which now uses it for voice dictation.
  • It gives the Muse line a fourth first-party surface after Muse Spark (text and agents), Muse Image (images), and Muse Glimmer (open-weights local), and it is the first Meta frontier model priced per minute of audio rather than per token.
Model
Muse Glimmer 30B
meta-models/Muse-Glimmer-30B — open weights under Apache 2.0
Successor
Current
Aug 10, 2026
First open-weights Meta frontier-lab model since Llama 4, and the first ever under Apache 2.0 rather than a bespoke Llama license. 30B local agent model distilled from Muse Spark. 131K context. Runs on one consumer GPU.
  • Released August 10, 2026; the announcement is at research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model. Coverage in VentureBeat, TechCrunch, and Phoronix.
  • Open weights under Apache 2.0 — the first Meta frontier-lab model released with downloadable weights since Llama 4 in April 2025, and the first ever released under a standard OSI-approved license instead of a bespoke Llama Community License. That means no >700M-MAU carve-out and no EU carve-out: the two restrictions that have defined every Llama release since July 2023 simply do not appear. Meta says the model was “assessed for open-weight release across all relevant categories” under its Advanced AI Scaling Framework.
  • Published under a new HuggingFace org. The weights are at meta-models/Muse-Glimmer-30Bnot the meta-llama org that carries the Llama line — alongside -GGUF, -assistant, and -ExecuTorch-PTE variants. Weights went up August 9; the announcement followed on August 10.
  • A local agent model, not a frontier model. ~29.6B total parameters — a dense causal transformer (52 layers, hidden dimension 6,656, grouped-query attention at a 16:1 ratio, a repeating [Local, Local, Local, Global] attention pattern with a 2,048 sliding window) plus a ~1.8B ViT-G/14 perception encoder. 131,072+ token context window; text and image in, text out, no audio. Knowledge cutoff January 4, 2026; trained on data from 100+ languages. It is distilled from Muse Spark — logit distillation on the teacher's outputs in pre-training, then agent-heavy mid-training, then supervised fine-tuning with on-policy distillation and RL.
  • Engineered to fit a consumer GPU. At full precision a 30B model would need over 55 GB; Meta ships ~4-bit quantization that puts the language model under 20 GB, leaving headroom for the KV cache, the perception encoder, and the drafter inside a 24 GB or 32 GB envelope (Meta reports 0.2% degradation for its 32 GB profile and 1.0% for the 24 GB one, averaged over 15 benchmarks). A DFlash-based speculative-decoding drafter proposes 16-token blocks in one pass; Meta measures 74.9 → 233.4 tok/s on an RTX 5090, 23.7 → 37.8 on an M4 Max, and 26.6 → 50.2 on an M5 Max (lab-reported; not independently verified here).
  • Built for agent workloads: end-to-end task completion on DeepSearch QA, MCP-Atlas, 𝜏-Bench and SWE-Bench, schema-precise tool calling, long-horizon planning, failure recovery and retry, and compatibility with OpenClaw, Hermes Agent, and other scaffolds. Meta compares it against Gemma4-31B and Qwen3.6-27B in its size class and publishes an evaluation methodology report.
  • Not served on the Meta Model API — that API sells the Muse Spark models and, since August 26, 2026, Muse Image; Glimmer is not on it. It is a download instead: Ollama, LM Studio, Unsloth, llama.cpp, ExecuTorch, MLX, vLLM and SGLang, with hosted routes via Together AI, Fireworks AI, and OpenRouter (as meta/muse-glimmer-30b, 131,072-token context; routed pricing spread from $0.30 / $1.10 to $0.35 / $1.50 per million input / output tokens across four providers as of September 8, 2026 — Phala at the bottom, then DeepInfra at $0.30 / $1.20, with Together AI and Fireworks AI at the top; Parasail, which was routing it in August, has since dropped off). Meta names AMD, Arm, Dell, Intel, and NVIDIA as device-optimization partners. Developer docs at dev.meta.ai/docs/muse-glimmer.
  • Shipped the same day as Mark Zuckerberg's “The Future is for Everyone” letter, which argues superintelligence should be broadly distributed rather than concentrated. The letter does not mention Llama, Muse, or open weights by name.
Model
Muse Spark 1.2
muse-spark-1.2, muse-spark-1.2-contributor — Meta Model API; closed-weights, no HuggingFace release yet
Successor
Available
Aug 5, 2026
Coding-focused update to 1.1, co-trained with the Muse Code terminal agent shipped the same day. 1M-token context. Added a cheaper API tier that trains on your prompts. Superseded by Muse Spark 1.3, but still the model to use for audio input.
  • Released August 5, 2026; the announcement is at research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2. Coverage in TechCrunch, CNBC, and Forbes.
  • A coding-focused update to Muse Spark 1.1 — Meta reports gains in code generation, complex debugging, codebase understanding, and end-to-end developer workflows, from scaling up training compute on coding tasks and widening training-environment diversity, while holding its strength in other areas including general agents.
  • Co-trained with its own harness. Meta trained the model together with Muse Code — rejection-sampled harness trajectories, recipe optimizations for goals, compaction, and subagents, and the Muse Code toolset integrated for harness compatibility. Long-horizon training covered whole-repository generation, large end-to-end projects, and auto-research. Meta also used Muse Spark 1.1 to generate and grade challenging coding environments as training data for 1.2, a self-improvement loop it credits for tighter instruction following.
  • Muse Code is the agent, not a model — a terminal coding agent for macOS and Linux, installed with curl -fsSL https://dev.meta.ai/install.sh | sh, with persistent async background agents that stay alive across a session rather than spawning per task, an append-only local event log Meta calls replay-exact and restart-safe, and bundled /plan, /grill, and /goal skills. It is tooling rather than a model release, so it does not carry its own row here (product page).
  • Muse Code left beta on August 31, 2026, roughly four weeks after it shipped alongside this model — “now out of beta with new features and updates,” per Meta's developer blog. The release adds workflows that fan a task out across staged parallel agents, local messaging between named live sessions, a double-Esc conversation rewind, and a model-based reviewer that pre-approves tool calls it judges safe (release 0.2.1, per Meta's changelog). Workflows are gated on the installed build and the rollout, so they are not on every platform yet.
  • And a third way to pay for the same model. The same release opened flat-rate monthly subscriptions for Muse Code — Everyday Usage, High Usage (3×), and Power Usage (10×), sold through Meta's Accounts Center and metered in requests per five hours rather than tokens — alongside the existing pay-as-you-go token billing. A subscription binds to the API key the CLI provisions at sign-in and works only through the CLI; any other key you create stays pay-as-you-go. So the Muse commercial surface now has three shapes: metered standard tokens, discounted tokens in exchange for training rights, and a flat monthly rate for the agent.
  • API shape: model ids muse-spark-1.2 and muse-spark-1.2-contributor, base URL https://api.meta.ai/v1, context window 1,048,576 tokens, text / image / video / audio / PDF in, text out. Also routed by OpenRouter as meta/muse-spark-1.2.
  • A new data-for-discount tier. The standard tier is $1.25 / $4.25 per million input / output tokens; the -contributor tier is $0.10 / $0.20 — roughly 12x and 21x cheaper — in exchange for permission to train future Meta models on your prompts and completions. It is the first Muse pricing surface that trades data rights for cost, and it is separate from the Llama Community License, which governs weights Meta has not released for this line.
  • Still closed-weights as of September 8, 2026: no HuggingFace release. Available the day of announcement in Muse Code and in the Meta Model API with what Meta calls “expanded global access.” Meta says it has “larger and much more capable models on the way.”
  • Superseded by Muse Spark 1.3 on September 2, 2026, an agentic-workflow successor. 1.2 stays served on the Meta Model API at the same price and Meta has named no retirement date — and Meta's own docs now route audio input back to 1.2, because audio understanding on 1.3 is flagged as not fully supported. On this one axis the superseded model is still the recommended one.
  • An open-weights version is announced but not shipped. On August 10, 2026, alongside the Muse Glimmer release, Chief AI Officer Alexandr Wang said Meta would release open weights for a version of Muse Spark 1.2 “soon” (The Register); Mark Zuckerberg said it the same day, and VentureBeat reports his framing as “in the coming weeks” — a window that has since passed without a release. Meta restated it on August 20, 2026 in its own research post, which frames its new multimodal evaluations as arriving “ahead of the open-weights release” (research.meta.ai), and again on September 2, 2026 in the Muse Spark 1.3 launch post, which lists “the Muse Spark open weights release” on its roadmap — without naming which version, so it is no longer clear the promise still attaches to 1.2. No date, no license, no parameter count, and no weights have been published — re-checked September 8, 2026, and huggingface.co/meta-models still holds nothing beyond the four Muse Glimmer repos — so this row stays closed-weights until they land. It would be the first frontier-scale open-weights release from Meta since Llama 4.
  • Multimodal results published August 20, 2026. A follow-up post, “The Multimodal Intelligence of Muse Spark 1.2,” covers visual coding (image or video in, working web page or game out, with a self-review loop), audio-visual understanding and dense captioning, and robotics — a specialized Muse Spark variant acting as a high-level planner that decomposes a goal into subtasks executed by a smaller Muse-family Vision-Language-Action policy. Meta says the multimodal gains are largest when the model can call tools to inspect an image more closely. The post also previews WildArtifactBench, an internal win-rate / Elo evaluation judged by agents or humans against a baseline agent rather than against ground-truth answers, and releases 10 of its tasks (lab-reported; not independently verified here).
  • Meta published an evaluation methodology report covering Terminal-Bench 2.1 (all 89 tasks, pass@1 over five attempts), DeepSWE v1.1 (113 tasks across 91 repositories), GDPVal-AA v2, MCP Atlas, and a 440-task internal coding benchmark drawn from real Meta pull requests, run in isolated Daytona cloud sandboxes — with the caveat that its harness “may not be specifically tuned for proprietary third-party models” (lab-reported; not independently verified here). A separate case study ran iterative GPU kernel optimization over 1,000+ tool calls for up to 24 hours on KDA and MLA kernels for NVIDIA Hopper.
Model
Muse Spark 1.1
muse-spark-1.1 — Meta Model API; closed-weights, no HuggingFace release
Successor
Available
Jul 9, 2026
Multimodal reasoning built for agentic tool and computer use. 1M-token context. First Muse model developers could call, via the Meta Model API. Superseded by Muse Spark 1.2.
  • Released July 9, 2026; the announcement is at ai.meta.com/blog/introducing-muse-spark-meta-model-api. Coverage in TechCrunch and Fortune.
  • A multimodal reasoning model built for agentic tasks — Meta describes major gains over the original Muse Spark in tool use, computer use, coding, and multimodal understanding. It zero-shot generalizes to new native tools, MCP servers, and custom skills.
  • 1,000,000-token context window that the model actively manages — it retrieves from much earlier work and compacts context while preserving the steps needed later. It orchestrates multi-agent systems, acting as either a main agent that plans and delegates to parallel subagents, or as a subagent that escalates back.
  • First Muse model outside developers can call. Meta launched a public preview of the new Meta Model API alongside it — the original Muse Spark was private-preview API only. Still closed-weights: no HuggingFace release. Also live in “Thinking” mode in the Meta AI app and on meta.ai.
  • Agentic features shipped into Meta AI on July 24, 2026. Meta turned the model's planning and tool use into consumer surfaces: a daily briefing assembled from your calendar, standing tasks that keep running once set up, connections to email and calendar apps, multi-minute research syntheses, and generated slide decks. Work in progress can be steered mid-run rather than only re-prompted afterward. Rolling out first in select markets in the Meta AI app and on meta.ai, with WhatsApp named as a later surface (Meta newsroom).
  • API shape, from Meta's own Model API docs: model id muse-spark-1.1, base URL https://api.meta.ai/v1, context window 1,048,576 tokens, bearer-token auth. The endpoint is drop-in compatible with the OpenAI SDK (Responses and Chat Completions) and the Anthropic SDK (Messages), so existing agent CLIs point at it without a rewrite.
  • Safety-evaluated under Meta's Advanced AI Scaling Framework; Meta reports it operates within safe margins across the Chemical & Biological, Cybersecurity, and Loss of Control frontier-risk categories (lab-reported; not independently verified here).
  • Launch-partner quotes came from Replit, Cline, Box, and the OpenClaw Foundation, all framing the model around agentic coding workloads. Meta says it has “even more capable models in training.”
  • Superseded by Muse Spark 1.2 on August 5, 2026, a coding-focused successor co-trained with the Muse Code agent. 1.1 remains served on the Meta Model API and through OpenRouter, and Meta has not announced a retirement date for it.
  • A pre-release build exploited a real website during a third-party safety evaluation. Meta disclosed the incident on August 14, 2026 (research.meta.ai). Meta had contracted Irregular to run cybersecurity evaluations, which are deliberately conducted with production safeguards removed to measure raw capability. In early July, a misconfiguration in Irregular's test environment gave a pre-release Muse Spark 1.1 access to the open internet and named a real website as the exercise target instead of a fictional one; believing it was the intended target, the model found and exploited a vulnerability, read data, and modified the site's database. Meta says its review of 10,000+ activity records found no other instance, that the model stayed within its assigned task and that this was “not a sophisticated offensive cyber attack or sandbox escape,” and that several other companies' models under evaluation at the time behaved similarly. Irregular has corrected the misconfiguration; Meta is adding independent verification of test-environment isolation and scenario review before evaluations begin (lab-reported; not independently verified here).
Model
Muse Image
muse-image-1.0 — Meta Model API; closed-weights, no HuggingFace release (code-named “Mango”)
Successor
Current
Jul 7, 2026
Meta Superintelligence Labs' first in-house image-generation model. Closed-weights; text-to-image plus editing. Pairs with Muse Spark for prompt reasoning. Powers Meta AI / Instagram / WhatsApp image generation. On the Meta Model API for outside developers since August 26, 2026, at $0.01 per image.
  • Announced July 7, 2026. The research announcement is at ai.meta.com/blog/introducing-muse-image-muse-video-msl; the product announcement is at about.fb.com/news/2026/07/introducing-muse-image-meta-ai. Coverage in Axios, TechCrunch, and CNBC.
  • Meta Superintelligence Labs' first in-house image model — the second Muse-branded surface after Muse Spark, and Meta's first frontier image generator built under Alexandr Wang's lab. Internally code-named Mango.
  • Closed-weights, delivered through Meta's apps rather than as a download. At launch Meta said it was “still evaluating whether it will make Muse Image available to outside developers”; that question has since been answered one way — the model reached the Meta Model API on August 26, 2026, seven weeks after announcement — but not the other: there is still no HuggingFace release, consistent with the Muse-line closed-weights posture set by Muse Spark. Access, not weights, again.
  • API shape, added August 26, 2026. Model id muse-image-1.0, base URL https://api.meta.ai/v1, endpoints /v1/images/generations and /v1/images/edits with multi-turn editing over the Responses API, bearer-token auth, output as base64 or a signed URL (Model API docs). Meta prices it at $0.01 per image and pitches “anchored composition” — edits that hold the rest of the frame steady — as the differentiator (developer blog). Also routed by OpenRouter as meta/muse-image from the same date, 65,536-token context, Meta the only provider; the upstream snapshot it forwards to is muse-image-1.0-eval-20260824. The announcement is developer-blog-only — nothing on research.meta.ai, ai.meta.com, or the newsroom.
  • Text-to-image generation plus editing: multi-photo blending, in-image text rendering, photo restoration, and sketch/markup editing. Uses Muse Spark for prompt reasoning. Available in Meta AI, Instagram (30+ AI effects), and WhatsApp, with Facebook, Messenger, and Advantage+ advertiser tooling announced as coming.
  • Meta says Muse Image “performs strongly across several benchmarks, generally surpassing Google's Nano Banana 2 and trailing only” OpenAI's image generator, and reports it holding the No. 2 spot on Arena for text-to-image, single-image editing, and multi-image editing on human-preference Elo as of July 5, 2026 (lab-reported; not independently verified here).
  • Generates through an agentic loop rather than mapping prompts straight to pixels: it invokes search and coding tools, self-refines its own drafts, and improves with test-time compute. Meta says the self-refinement behavior was not designed but emerged during reinforcement learning. Images made in Meta AI carry Content Seal, Meta's invisible watermark, which survives cropping, compression, and screenshotting.
  • Muse Video previewed alongside it — built on the same pretraining base, with native audio support, ranked No. 3 on Arena for text-to-video human-preference Elo at announcement. Meta names audio-video synchronization and physically accurate fast motion as current gaps. It is not yet released: “coming soon to creators and in Meta AI,” so it does not carry a row here yet.
  • The launch drew immediate pushback over a feature that lets users pull public Instagram photos of a tagged person into generated images without explicit notification (TechCrunch).
Model
Muse Spark
closed-weights, API only — no HuggingFace release
Successor
Available
Apr 8, 2026
First model from Meta Superintelligence Labs. Closed-weights, API-only. Marks the end of Meta's open-weights frontier-AI era. “Contemplating” reasoning mode. Superseded by Muse Spark 1.1.
  • Released April 8, 2026 in private API preview; the announcement is at ai.meta.com/blog/introducing-muse-spark-msl.
  • Closed-weights, API-only — the first Meta frontier-AI model not released as open weights since Llama 2 in July 2023. Branded under the new Muse family rather than Llama. Meta has stated it “hopes to open-source future versions” but the launch model is proprietary.
  • Natively multimodal. Three operating modes: Instant (fast responses for casual queries), Thinking (chain-of-thought for multi-step reasoning), and Contemplating (fully rolled out by late April 2026, per CNBC) that orchestrates parallel sub-agents (“thinking wider, not longer”), positioned against Gemini Deep Think and the OpenAI o-series.
  • Treated as the Llama line's successor here because it is shipped by the same Meta organization (Meta Superintelligence Labs, formed June 2025 under Alexandr Wang) and replaces the Llama line as Meta's frontier-AI surface, even though it is not Llama-branded. The closed-weights turn is covered in the prose history below.
  • Superseded by Muse Spark 1.1 on July 9, 2026, which took over the “Thinking” mode in the Meta AI app and became the first Muse model exposed to outside developers. Meta has not announced a retirement date for the original model.
  • Coverage in CNBC, VentureBeat.

The closed-weights turn — April 8, 2026. Above this line: the Muse line, Meta's successor surfaces to Llama under Meta Superintelligence Labs — Muse Spark (April 2026), Muse Image (July 2026), Muse Spark 1.1 (July 2026), Muse Spark 1.2 (August 2026), Muse Voice Transcribe (September 2026) and Muse Spark 1.3 (September 2026), all closed-weights, plus Muse Glimmer (August 2026), which broke the streak with open weights under Apache 2.0 at a size class below the flagship. Meta has said three times since — in August and again at the 1.3 launch — that an open-weights Muse Spark release is coming, without naming a date or a license. Below: the Llama line itself — twelve open-weights releases between July 2023 and April 2025, the dominant open-weights frontier-AI lineage of the period, plus the research-only LLaMA 1 (February 2023) and the never-released Llama 4 Behemoth. The transition follows the June 2025 Scale AI acqui-hire that brought Alexandr Wang to Meta as Chief AI Officer, the November 2025 departure of Chief AI Scientist Yann LeCun, and the dissolution of FAIR's open-research mandate into MSL's product organization.

Model
Llama 4 Maverick
meta-llama/Llama-4-Maverick-17B-128E-Instruct
Flagship
Current
Apr 5, 2025
Current open-weights flagship. First Mixture-of-Experts Llama. 17B active params over 128 experts (~400B total). 1M context. Natively multimodal.
  • Released April 5, 2025; the announcement is at ai.meta.com/blog/llama-4-multimodal-intelligence. HuggingFace launch post: huggingface.co/blog/llama4-release.
  • First Mixture-of-Experts Llama — 17B active parameters routed over 128 experts (~400B total parameters). 1,000,000-token context window. Natively multimodal (text + image), not bolted-on vision tower.
  • License: Llama 4 Community License — preserves the >700M monthly-active-user carve-out from the Llama 2 / 3 lineage. License text at llama.com/license.
  • LMArena episode at launch: a custom “Llama-4-Maverick-03-26-Experimental” variant (not the released weights, verbose, emoji-laden) was submitted to LMArena and rocketed to #2; after policy clarification by LMArena, the unmodified shipped Maverick was tested and ranked #32. TechCrunch reported Meta's framing and LMArena's policy update; LMArena said “Meta's interpretation of our policy did not match what we expect from model providers.”
  • Promoted by Meta as the open-weights flagship through 2025; shipped alongside Scout (the smaller MoE companion) and the never-released Behemoth (the 2T-parameter teacher model).
Model
Llama 4 Scout
meta-llama/Llama-4-Scout-17B-16E-Instruct
Flagship
Available
Apr 5, 2025
Smaller MoE companion to Maverick. 17B active over 16 experts (~109B total). Headline 10M-token context window. Single-H100 deployable at Int4.
  • Released April 5, 2025 alongside Maverick and the announced-but-never-shipped Behemoth.
  • 10,000,000-token context window at launch — the largest ever claimed for an open-weights model, though the practical retrieval-quality envelope at full 10M is a topic of ongoing community evaluation.
  • 17B active parameters over 16 experts (~109B total). Designed to be deployable on a single H100 at Int4 quantization.
  • Same Llama 4 Community License as Maverick; HuggingFace download via meta-llama/Llama-4-Scout-17B-16E-Instruct.
Model
Llama 4 Behemoth
never publicly released — weights not on HuggingFace
Flagship
Legacy
Apr 5, 2025
288B active over 16 experts (~2T total). Announced as “still training” at Llama 4 launch. Indefinitely delayed; never publicly released.
  • Announced April 5, 2025 alongside Maverick and Scout as “still training”; Meta has never publicly released the weights.
  • 288B active parameters, 16 experts, ~2 trillion total parameters — the largest model Meta has publicly disclosed training. Intended as the teacher model that distilled Maverick and Scout.
  • Originally targeted for early summer 2025; pushed to June, then fall 2025, then indefinitely. Reporting in mid-2025 said benchmark gains over Maverick were too incremental and MoE routing at the 2T scale proved unstable.
  • The Behemoth delay was one of the proximate causes of the June 2025 Scale AI acqui-hire and Meta Superintelligence Labs reorg covered in the prose history below.
  • Carried as a row here because it is named in Meta's own Llama 4 announcement and is the implicit denominator in any “current Meta frontier-flagship” question; status is Legacy on the “never publicly released, no longer recommended” reading.
Model
Llama Guard 4 12B
meta-llama/Llama-Guard-4-12B
Specialized
Available
Apr 5, 2025
Latest Llama Guard safety classifier. 12B pruned dense model derived from Llama 4 Scout. Input/output moderation for Llama 4 deployments.
  • Released April 5, 2025 alongside Llama 4. The Llama Guard line is the safety-classifier sibling of the chat models — it does prompt and response moderation for Llama deployments.
  • 12B-parameter pruned dense model derived from Llama 4 Scout; HuggingFace at meta-llama/Llama-Guard-4-12B.
  • Replaces the Llama Guard 3 family (8B, 1B, and 11B Vision variants from 2024) for Llama 4 deployments.
  • Distributed under the Llama 4 Community License with the Acceptable Use Policy. Repo: github.com/meta-llama/PurpleLlama.
Model
Llama 3.3 70B Instruct
meta-llama/Llama-3.3-70B-Instruct
Flagship
Available
Dec 6, 2024
70B-only refresh. Marketed as 405B-level performance at 70B inference cost. 128K context. The last 3.x release before the Llama 4 reset.
  • Released December 6, 2024. TechCrunch coverage; HuggingFace card: meta-llama/Llama-3.3-70B-Instruct.
  • 70B-only release — no 8B / 405B siblings. Marketed by Meta as delivering 405B-level performance at 70B inference cost; pretrained on ~15T tokens with 25M synthetic examples on top of public instruction sets.
  • 128K-token context window; multilingual; improved instruction-following and coding over Llama 3.1 70B.
  • Distributed under the Llama 3.3 Community License — same shape as the 3.1 / 3.2 licenses, >700M MAU carve-out preserved, naming convention requires “Llama” prefix on derivatives.
  • The last 3.x model. The next release was Llama 4 four months later.
Model
Llama 3.2 (Vision + Edge)
meta-llama/Llama-3.2-{1B, 3B, 11B-Vision, 90B-Vision}
Specialized
Available
Sep 25, 2024
First multimodal Llama (11B / 90B Vision) and first edge-targeted sizes (1B / 3B). EU regulators excluded from the multimodal license.
  • Released September 25, 2024 at Meta Connect 2024. Announcement: ai.meta.com.
  • Two product lines announced together: a vision-language pair (11B Vision built on Llama 3.1 8B; 90B Vision on Llama 3.1 70B) and an edge-targeted pair (1B and 3B text-only, distilled from larger 3.1 teachers, designed to run on phones and laptops).
  • 128K-token context window across all four sizes.
  • EU multimodal carve-out: “the rights granted under the Llama 3.2 Community License Agreement are not being granted” to EU-domiciled individuals or EU-headquartered companies for the multimodal models, citing regulatory uncertainty under the EU AI Act, GDPR, and Digital Markets Act. EU users could still use the 1B / 3B text models. Slator coverage.
  • HuggingFace ids on the meta-llama org. Llama Guard 3 11B Vision shipped alongside as the multimodal-safety classifier.
Model
Llama Guard 3 family
meta-llama/Llama-Guard-3-{8B, 1B, 11B-Vision}
Specialized
Available
Jul 23, 2024
Three sizes shipped alongside Llama 3.1 / 3.2: 8B, 1B (edge), and 11B Vision (multimodal moderation). Input and output safety classification.
  • Llama Guard 3 8B shipped July 23, 2024 alongside Llama 3.1; the 1B and 11B Vision variants followed with Llama 3.2 on September 25, 2024.
  • Three sizes targeting different deployment surfaces — 8B for production servers, 1B for on-device, 11B Vision for multimodal-safety classification.
  • Input/output moderation: classifies prompts and responses against Meta's safety taxonomy (illegal violence, sexual content involving minors, etc.).
  • Distributed under the corresponding Llama 3.x Community License. Repo: github.com/meta-llama/PurpleLlama.
Model
Llama 3.1 (8B / 70B / 405B)
meta-llama/Llama-3.1-{8B, 70B, 405B}-Instruct
Flagship
Available
Jul 23, 2024
First “frontier-class” open-weights model. 405B competitive with GPT-4 / Claude 3.5 Sonnet at launch. 128K context. Co-released with Zuckerberg's open-source-AI letter.
  • Released July 23, 2024; the announcement is at ai.meta.com/blog/meta-llama-3-1.
  • First “frontier-class” open-weights model — Llama 3.1 405B was widely characterized as the first openly-available model competitive with GPT-4 / Claude 3.5 Sonnet on broad benchmarks at launch.
  • Three sizes (8B, 70B, 405B); 128K context across all three, up from 8K on Llama 3. Improved tool use and multilingual capability.
  • 405B trained on ~16,000 H100 GPUs per Meta's accounting, ~3.8e25 FLOP total — the first Llama trained at this scale.
  • Co-released with Mark Zuckerberg's “Open Source AI Is the Path Forward” letter, which likened open-source AI's trajectory to Linux. The letter is now read with sharp dramatic irony given the closed-weights Muse Spark turn twenty-one months later (covered in the prose history below).
Model
Llama 3 (8B / 70B)
meta-llama/Meta-Llama-3-{8B, 70B}-Instruct
Flagship
Legacy
Apr 18, 2024
Two sizes. ~15T-token pretraining (7× Llama 2). 128K-vocab tokenizer. Grouped-query attention across all sizes. 8K context.
  • Released April 18, 2024; the announcement is at ai.meta.com/blog/meta-llama-3.
  • Two public sizes (8B and 70B); pretrained on ~15 trillion tokens — roughly seven times Llama 2's pretraining corpus, with four times more code data.
  • New 128K-vocabulary tokenizer; Grouped-Query Attention across all sizes (Llama 2 had GQA only at 70B).
  • 8,192-token context window — bumped to 128K with Llama 3.1 three months later.
  • Distributed under the Meta Llama 3 Community License: same shape as Llama 2 (>700M MAU carve-out preserved); adds an attribution requirement that derivative model names include the “Llama 3” prefix.
Model
Code Llama 70B
codellama/CodeLlama-70b-{hf, Python-hf, Instruct-hf}
Specialized
Available
Jan 29, 2024
Largest Code Llama tier. Three flavors: base, Python specialization, Instruct. 16K context. The high-water mark of the Code Llama line.
  • Released January 29, 2024 as the largest Code Llama tier; followed the original 7B / 13B / 34B sizes by five months.
  • Three flavors per size: base, Python specialization, and Instruct (the chat-tuned coding assistant variant).
  • 16,384-token context window — substantially longer than the 4K of the base Llama 2 line.
  • Distributed under the Llama 2 Community License. HuggingFace org: huggingface.co/codellama.
  • Marked the high-water mark of the named Code Llama line; later code-task work was folded into the general-purpose Llama 3.x and Llama 4 instruction-tuning.
Model
Purple Llama / Llama Guard 7B
meta-llama/LlamaGuard-7b
Specialized
Legacy
Dec 7, 2023
First Llama Guard. 7B classifier built on Llama 2. Launched alongside Purple Llama, Meta's umbrella for AI-safety tooling.
  • Released December 7, 2023 as the first Llama Guard; HuggingFace at meta-llama/LlamaGuard-7b.
  • 7B classifier built on Llama 2 7B, fine-tuned for input and output safety classification (prompt and response moderation).
  • Launched as part of Purple Llama — Meta's umbrella for AI-safety tooling, also covering CyberSecEval (security benchmarks) and prompt-injection benchmarks. Repo: github.com/meta-llama/PurpleLlama.
  • Superseded by Llama Guard 3 family (mid-2024) and Llama Guard 4 (April 2025); status is Legacy on the “no longer the recommended classifier” reading.
Model
Code Llama (7B / 13B / 34B)
codellama/CodeLlama-{7b, 13b, 34b}-{hf, Python-hf, Instruct-hf}
Specialized
Legacy
Aug 24, 2023
Initial Code Llama release. Fine-tune of Llama 2 on code corpus. Three sizes × three flavors. 16K context.
  • Released August 24, 2023; the announcement is at ai.meta.com/blog/code-llama-large-language-model-coding.
  • Three sizes (7B, 13B, 34B), three flavors per size (base, Python, Instruct). The 70B tier followed five months later (January 2024 row above).
  • 16,384-token context window — meaningfully longer than the 4K of the base Llama 2 line.
  • Distributed under the Llama 2 Community License (>700M MAU carve-out applies).
  • Superseded as the open-weights coding-model recommendation by general-purpose Llama 3.x / Llama 4 instruction-tuning by mid-2024.
Model
Llama 2 / Llama 2-Chat
meta-llama/Llama-2-{7b, 13b, 70b}, -chat-hf variants
Flagship
Legacy
Jul 18, 2023
First open-weights Llama with commercial use permitted. Three sizes (7B, 13B, 70B). 4K context. The Community License debuts with the >700M-MAU carve-out.
  • Released July 18, 2023 in partnership with Microsoft (Azure preferred partner); the announcement is at about.fb.com/news/2023/07/llama-2.
  • Three public sizes (7B, 13B, 70B); a 34B was trained but withheld for what Meta described as red-teaming reasons. Pretraining doubled to ~2T tokens vs. LLaMA 1.
  • 4,096-token context window. The 70B tier introduced Grouped-Query Attention (GQA), the architectural choice that made later long-context expansions practical.
  • First Llama released as open weights with commercial use permitted, under the bespoke Llama 2 Community License Agreement. The license carves out companies whose products had >700 million monthly active users on the Llama 2 release date — a deliberate exclusion aimed at TikTok, Google, Apple, and other direct platform competitors. The carve-out has been preserved through every subsequent Llama license.
  • Llama 2-Chat ships RLHF-tuned chat variants in matching 7B / 13B / 70B sizes; the chat models powered most early third-party Llama deployments.
  • The release came four months after the LLaMA 1 leak and is widely read as Meta deciding the marginal cost of an official open-weights release was near zero once the weights were already in the wild — and that the upside (developer mindshare, ecosystem) was substantial.

The open-weights commercial era — July 18, 2023. Above this line: every Llama released as open weights with commercial use permitted, under successive Llama Community Licenses with the >700M-MAU carve-out preserved. Below: the original LLaMA 1, released as research-only by application — not actually open weights, and accidentally became so via the March 2023 leak that arguably forced the Llama 2 open release four months later.

Model
LLaMA 1 (7B / 13B / 33B / 65B)
research-only by application; weights leaked Mar 3, 2023
Flagship
Legacy
Feb 24, 2023
The original Llama. Four sizes. Research-only by application. Weights leaked to 4chan March 3, 2023, arguably forcing the Llama 2 open commercial release.
  • Released February 24, 2023; the paper is “LLaMA: Open and Efficient Foundation Language Models” (Touvron et al., 2023). The Meta page is at ai.meta.com/research/publications.
  • Four sizes (7B, 13B, 33B, 65B); pretraining 1–1.4 trillion tokens of fully public data. 2,048-token context window.
  • Demonstrated that state-of-the-art performance was achievable using only publicly-available data; LLaMA 65B competed with PaLM-540B and GPT-3 on broad benchmarks.
  • Released research-only, by application — non-commercial, gated by request form, available to academic and government applicants only. Not actually open weights.
  • The leak (March 3, 2023): a user on 4chan's /g/ board posted a torrent containing the LLaMA weights; the same day, a pull request was filed on Meta's official Llama GitHub repo proposing to add the magnet link to documentation. Meta filed DMCA takedowns through HuggingFace and GitHub through March 20. The leak is widely credited with forcing the Llama 2 open-commercial release four months later. The Register coverage; DeepLearning.AI's The Batch.

Click any row to expand. Each row has a stable id for sharing — e.g. /ai/llama/versions/#llama-4-maverick, #llama-3-1, #llama-2, #muse-spark. Llama family hub: llama.com; license texts at llama.com/license; HuggingFace org: huggingface.co/meta-llama.

The March 2023 LLaMA leak

Meta released LLaMA 1 on February 24, 2023 as a research-only model gated by application form, accessible to academic and government applicants and not under any open-source license. On March 3, 2023 — nine days later — a user posted a torrent containing the 7B and 65B weights to 4chan's /g/ technology board under the handle llamanon. The same day, a pull request was filed against Meta's official Llama GitHub repository proposing to add the magnet link to the README. The torrent reportedly pulled directly from Facebook's CDN at high speed, embedded with a unique download URL traceable to the leaker's gated-access grant.

Meta filed DMCA takedowns through HuggingFace and GitHub through March 20, 2023, but the weights had already propagated. Coverage at the time in The Register and DeepLearning.AI's The Batch documents the propagation timeline. Within weeks, the leaked LLaMA 1 weights were the de facto foundation for an enormous open-source fine-tune ecosystem — Alpaca, Vicuna, Guanaco, and dozens of others — on HuggingFace.

The leak is widely credited with forcing Meta's hand on Llama 2: with the weights already in the wild, the marginal cost of an official open-weights commercial release dropped to near zero, while the upside (developer mindshare, ecosystem effects, reduced lawsuit risk) was substantial. Llama 2 shipped with permissive commercial terms four months later. The decision-tree shape of the leak → commercial open release pattern is the load-bearing precondition for everything that followed in the Llama story through 2025.

The Llama Community License and the >700M-MAU carve-out

Llama 2 introduced the bespoke Llama 2 Community License Agreement — not an OSI-approved open-source license, but a Meta-authored license with permissive commercial terms and a single load-bearing carve-out: companies whose products had more than 700 million monthly active users on the Llama 2 release date (July 18, 2023) must request a separate license from Meta. The carve-out is a deliberate exclusion aimed at TikTok, Google, Apple, and other direct platform competitors; reading the precise language requires the actual PDF, which Meta hosts at llama.com/license and ai.meta.com/llama/license.

The carve-out has been preserved through every subsequent Llama license: Llama 3 Community License (April 2024) added an attribution requirement that derivative model names include the “Llama 3” prefix; the Llama 3.1 / 3.2 / 3.3 licenses extended the naming convention to require the “Llama” prefix on derivatives; the Llama 4 Community License preserved the same general shape. The Open Source Initiative has consistently said the Llama license does not meet the Open Source Definition because the use-restriction (any restriction tied to who you are, not what you do) is incompatible with the OSD — this is the substantive basis for the “source-available, not open source” framing of the line.

Llama 3.2 added a regional carve-out: the multimodal models (11B and 90B Vision) were excluded from EU-domiciled individuals and EU-headquartered companies, citing regulatory uncertainty under the EU AI Act, GDPR, and Digital Markets Act. The text-only 1B / 3B models remained available in the EU. Llama 4 carries the same EU exclusion, in identical wording — and because every Llama 4 model (Maverick, Scout, Behemoth) is natively multimodal, the restriction effectively covers the entire Llama 4 line for EU-domiciled individuals and EU-headquartered companies. EU-based end-users may still use Llama-4-powered services originating outside the EU: the restriction “does not apply to end users of a product or service that incorporates any such multimodal models.” A separate EU instrument arrived on July 28, 2026, when Meta said it would sign the EU AI Act Code of Practice on Transparency of AI-Generated Content — that one governs labeling and provenance of generated media (Content Seal is Meta's own implementation), not who may download and deploy the weights, so it leaves the license carve-outs untouched.

A sourcing detail worth knowing if you go looking for that language: the EU exclusion is not in the Community License text itself. It lives in the matching Acceptable Use Policy, which each license incorporates by reference, and it operates by withholding the Section 1(a) grant. The license PDF is silent on the EU; searching it alone will tell you the restriction doesn't exist. Re-verified 2026-09-08 against the Llama 4 Acceptable Use Policy and Community License as published in Meta's llama-models repo, alongside the Llama 3.2 equivalents — the Llama 4 wording is still identical to Llama 3.2's, end-user exemption included. The same pass confirmed the >700M-MAU carve-out intact in Llama 4, pegged to the Llama 4 version release date (April 5, 2025) and measured over the preceding calendar month.

The bespoke-license era has an end date, or at least an exception. Muse Glimmer, released August 10, 2026, ships under Apache 2.0 — a standard, OSI-approved license, and the first Meta frontier-lab model in the line's history not governed by a Meta-authored one. Practically, the two restrictions that have shaped every “can I ship with this?” question since July 2023 are simply absent: no >700M-MAU carve-out, so the platform competitors the Llama license was drawn to exclude are not excluded, and no EU carve-out, even though the model takes image input and would have tripped the multimodal restriction under a Llama license. The naming convention is gone too — derivatives need no “Llama” prefix. It is also published outside the Llama plumbing: the weights sit in a new meta-models HuggingFace org rather than meta-llama, and there is no Muse generation directory in the llama-models repo that hosts the per-generation license and use-policy texts. Whether Apache 2.0 becomes the shape of future open Muse releases or stays a one-off for the small local model is unresolved — Meta announced on August 10, 2026 that an open-weights version of Muse Spark 1.2 is coming, restated it on August 20 and again in the September 2 Muse Spark 1.3 launch post (which drops the version number), but has never said under which license — and as of September 8, 2026 every API-served Muse model remains proprietary and closed-weights: the three Muse Spark versions, Muse Image, which joined them on the API on August 26, and Muse Voice Transcribe, which shipped closed on September 1.

FAIR, GenAI, and the org evolution

Meta's AI research home through 2022 was FAIR (Fundamental AI Research), founded in 2013 under Yann LeCun as Chief AI Scientist and led from 2023 by Joelle Pineau. FAIR's mandate was long-horizon basic research; its publication culture was open and academically-styled. LLaMA 1 was a FAIR project.

After ChatGPT's November 2022 launch caught Meta flat-footed, the company formed a separate Generative AI organization (GenAI) in February 2023, led by VP Ahmad Al-Dahle (who departed Meta in January 2026 to become CTO of Airbnb) and tasked with shipping product-grade generative AI on a faster cadence than FAIR's research timeline. Llama 2 and onward shifted progressively into GenAI under Manohar Paluri; FAIR was pushed toward longer-horizon work as the shipping pace of the GenAI org accelerated through 2024.

Through late 2024 and into 2025, FAIR's relevance to Meta's shipping AI products visibly waned. Internal reporting in April 2025 described FAIR as “dying a slow death.” Joelle Pineau announced her departure on April 1, 2025, last day May 30, 2025, after eight years leading FAIR; she was the most prominent internal advocate for open-source releases. Her exit set up the structural reorganization that came two months later.

The Scale AI acqui-hire and Meta Superintelligence Labs (June 2025)

On June 12–13, 2025, Meta announced a $14.3 billion investment for a 49% non-voting stake in Scale AI, valuing Scale at over $29 billion. The deal was structurally an acqui-hire: its purpose was to bring Scale CEO Alexandr Wang (then 28) to Meta as Chief AI Officer. Wang stepped down as Scale CEO but remained on Scale's board. Coverage in CNBC and Fortune.

On June 30, 2025, Mark Zuckerberg announced the formation of Meta Superintelligence Labs (MSL) with Wang as Chief AI Officer leading the new unit and former GitHub CEO Nat Friedman leading AI products. FAIR and GenAI were both consolidated under MSL; a new TBD Lab sub-group was created to work on next-generation LLMs. CNBC published the internal memo. The reorganization was the structural endpoint of Pineau's departure two months earlier and of the Llama 4 Behemoth delay reporting through May 2025.

Friction between Meta and Scale AI emerged through August 2025 as Scale's enterprise-data-labeling business saw revenue impact (competitors stopped sending training data to a Meta-owned vendor); TechCrunch covered the cracks in the partnership. MSL itself restructured into four sub-groups in August 2025.

Yann LeCun departs (November 2025) — AMI Labs

On November 19, 2025, Chief AI Scientist Yann LeCun announced his departure from Meta to launch his own startup focused on Advanced Machine Intelligence and “world models.” Reporting in CNBC and Bloomberg tied the decision to LeCun being asked to report to Wang — a structural demotion, since Wang's closed-source product orientation conflicts with LeCun's long-held public position that scaling LLMs is not a path to AGI and that open research is foundational to the field.

LeCun named the startup AMI Labs (Advanced Machine Intelligence) on December 18, 2025 and launched the public-facing site at amilabs.xyz in January 2026. On March 9, 2026 AMI announced a $1.03 billion seed round at a $3.5 billion pre-money valuation, co-led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital, and Bezos Expeditions. The stated mission is “intelligent systems that understand the real world” rather than language-only LLMs. Meta is not an investor but a partnership is reportedly under discussion (LeCun has publicly suggested Meta could be AMI's first client), and AMI has a clinical-AI partnership with Nabla. With both Pineau and LeCun out, the open-research mandate inside Meta effectively ended at MSL's June 2025 formation.

The closed-weights turn — Muse Spark (April 2026)

On April 8, 2026, Meta Superintelligence Labs released Muse Spark, its first model since the June 2025 reorganization. Muse Spark shipped closed-weights, API-only, in private preview — the first frontier-AI model from Meta released without open weights since Llama 2 in July 2023. Branded under the new Muse family rather than Llama. The announcement is at ai.meta.com/blog/introducing-muse-spark-msl; coverage in CNBC and VentureBeat. Meta has stated it “hopes to open-source future versions” of Muse, but the launch model is proprietary.

Muse Spark is natively multimodal and ships with three operating modes — Instant (fast responses, no extended reasoning), Thinking (chain-of-thought), and Contemplating, which orchestrates parallel sub-agents (“thinking wider, not longer”), positioned competitively against Gemini Deep Think and the OpenAI o-series. Several outlets characterized the launch as the end of Meta's open-weights frontier-AI era. The closed-weights turn is the dramatic reversal of Mark Zuckerberg's July 2024 “Open Source AI Is the Path Forward” letter, which had argued (twenty-one months earlier) that open-source AI's trajectory would mirror Linux's path from inferior-but-cheap to dominant-via-ecosystem.

The line has since moved twice in one week without reversing course on weights. On July 7, 2026 MSL shipped Muse Image and previewed Muse Video (announcement), both closed-weights. On July 9, 2026 it shipped Muse Spark 1.1 (announcement) and launched a public preview of the Meta Model API. That API is the one partial opening so far: the original Muse Spark was private-preview only, and Muse Spark 1.1 is the first Meta frontier model any outside developer can call. Access, not weights — there was still no HuggingFace release, and Meta's “hopes to open-source future versions” framing from April had not yet been acted on. The pattern repeated on July 24, 2026, when Meta wired the model's planning and tool use into Meta AI itself as consumer features — scheduled tasks, calendar and email connections, research syntheses, generated slides (newsroom post). Wider reach, same posture on weights.

On August 5, 2026 the line moved again, in a direction that says something about what Meta is selling. MSL shipped Muse Spark 1.2 — a coding-focused update — together with Muse Code, a terminal coding agent the model was co-trained with, shipped in beta (announcement). Weights stayed closed, but the commercial terms opened a new seam: alongside the standard API tier, Meta introduced a muse-spark-1.2-contributor tier at roughly a twelfth the input price and a twenty-first the output price, in exchange for permission to train future Meta models on the customer's prompts and completions. Where the Llama era traded weights for ecosystem, the Muse era is trading price for training data. The announcement also moved hosts: it was published on research.meta.ai, which now carries the Muse-line posts, and llama.com now redirects to developer.meta.com/ai/.

Five days later the turn partly reversed. On August 10, 2026 MSL released Muse Glimmer, a 30-billion-parameter local agent model, with open weights under Apache 2.0 (announcement) — the first downloadable model from Meta's frontier lab since Llama 4 in April 2025, and the first ever under a standard open-source license rather than a bespoke Llama Community License. Meta framed it as “keeping with our long tradition of sharing fundamental AI research.” The reopening is real but bounded: Glimmer is distilled from Muse Spark and sits a size class below it, and the flagship Muse Spark models stayed closed and API-only. What it does settle is the April 2026 “hopes to open-source future versions” framing — that has now been acted on once, at the small end of the line.

The same announcement went further than the model it shipped. On August 10, 2026, Chief AI Officer Alexandr Wang said Meta would also release open weights for a version of Muse Spark 1.2 — the flagship — “soon,” with no date attached. Meta restated the plan in its own words on August 20, 2026, opening a research post on the model's multimodal evaluations with the line that it was publishing them “ahead of the open-weights release” (announcement). That would move the reopening from the local tier to frontier scale, which is the question Glimmer left open. It has not happened yet: as of September 8, 2026 the meta-models org holds only the four Muse Glimmer repos, the Meta Model API still serves every Muse Spark version as a closed model, and Meta has named neither a release date nor a license. This page adds rows for models that have shipped, so the Muse Spark rows stay closed-weights here until the weights are published.

The same day, Mark Zuckerberg published “The Future is for Everyone”, a letter proposing a philosophy of “individual empowerment as the source of prosperity, invention as the primary purpose of superintelligence, and balance of power as the foundation of safety,” and asking whether superintelligence “will be centralized and restricted to a few institutions, or will it be a tool that empowers everyone.” It is the closest thing to a successor to the July 2024 open-source letter, and it landed alongside the first open-weights release since the turn — but it argues from access and distribution rather than from open weights, and it does not mention Llama, Muse, or open weights anywhere in its text.

What actually moved next was access again. On August 26, 2026, sixteen days after the open-weights promise and seven weeks after the model itself launched, Meta put Muse Image on the Meta Model API as muse-image-1.0 at $0.01 per image, resolving the “still evaluating whether it will make Muse Image available to outside developers” line from July. The weights stayed closed. That is now the line's settled rhythm at the frontier: each surface opens to developers as a metered endpoint, months after it opens to consumers, and never as a download. Worth noting where the announcement landed — a post on the AI developer blog, with nothing on research.meta.ai, ai.meta.com, or the newsroom. Meta now ships platform news and research news to different addresses.

Then, on August 31, 2026, the packaging changed rather than the model. Meta took Muse Code out of beta — four weeks after shipping it — and put flat-rate monthly subscriptions behind it, in three tiers (Everyday, High, and Power Usage) sold through Meta's Accounts Center and metered in requests per five hours instead of tokens (developer blog, subscription docs). The same release added staged multi-agent workflows, messaging between live sessions, and a conversation rewind. That gives the same closed model three distinct commercial shapes in under a month: metered tokens at the standard rate, tokens at roughly a twelfth the price if Meta may train on them, and a flat monthly fee for the agent that drives them. Through the end of August the shape of the arc was hard to miss: every change Meta had shipped at the frontier since the April 2026 launch was about how you buy access, and the one announced change to who may hold the weights, promised on August 10, had not landed. Where the announcement went said the same thing: the developer blog again, with nothing on research.meta.ai or the newsroom.

Then the models resumed. On September 1, 2026 MSL shipped Muse Voice Transcribe, its first real-time audio-perception model (announcement) — a new modality for the line, and its first product priced per minute of audio rather than per token. A day later, on September 2, 2026, it shipped Muse Spark 1.3 (announcement), an agentic-workflow update to the flagship. Both are closed-weights, so the posture is unchanged — but two things about the 1.3 launch are worth recording. It shipped incomplete, and was completed quietly two days later: Meta held its top “max reasoning” mode back “shortly after we finish additional safety testing,” so for two days the flagship was generally available while its highest setting was not — and, as VentureBeat pointed out on September 3, several of the launch's strongest disclosed scores belonged to that unshipped configuration rather than to the xhigh setting developers could actually call. On September 4, 2026 max went live on the Meta Model API and in Muse Code, on the standard tier only. Meta marked the change by editing the September 2 post in place rather than publishing a new one, and left both the post's published and its modified date reading September 2 — so the release is invisible to anyone diffing headlines, dates, or the blog index, and only the reasoning docs and the post body carry it. Nothing was concealed about the numbers; what was hard to see was the moment the gap closed. And it shipped narrower on one axis: Meta's own docs flag audio understanding on 1.3 as not fully supported and send audio work back to Muse Spark 1.2 or to the new transcription model. The open-weights promise was restated a third time in the same post — “bigger models, the Muse Spark open weights release, and more” — and this time without naming a version, so it is no longer clear that what was promised on August 10 for Muse Spark 1.2 is still what is coming. Zuckerberg restated it himself on X the same day, also without a version and in the plural — open-weights Muse Spark “releases” — which moved the promise up to the CEO without moving it any closer to a date (The Register). The one time frame anyone attached to it — “in the coming weeks,” per VentureBeat's account of the August 10 statement — has now come and gone. Four months in, the Muse line has published open weights exactly once, at the size class furthest from the frontier.

Muse Spark is treated on this page as Llama's successor — same Meta organization (MSL), same flagship-frontier slot — even though it is not Llama-branded. Whether future Muse releases are open-weights, partial open-weights, or fully closed remains the load-bearing question for the Meta-AI story going forward; Muse Glimmer answers it for the local-model tier, and at the frontier Meta has now stated an intention it has not yet executed. No new Llama has shipped since April 2025 and none is announced — the partial reopening arrived under the Muse brand and a different license instead.

Where to run Llama

Llama is the most widely-distributed frontier-AI line because the weights are open. Inference paths through 2025–2026 break into three categories.

Self-host. Download from the HuggingFace meta-llama org and run with vLLM, llama.cpp, Ollama, MLX (Apple Silicon), or TensorRT-LLM. The 1B / 3B Llama 3.2 edge models are specifically designed for on-device deployment. Since August 10, 2026 the self-host path also covers a Muse model: Muse Glimmer 30B ships open weights under Apache 2.0 from the separate meta-models org, quantized to fit a single 24 GB or 32 GB consumer GPU, with LM Studio, Unsloth, ExecuTorch and SGLang added to the runtime list.

Hosted-inference providers. Together AI, Fireworks AI, Groq (notable for very high tokens-per-second on Llama 3 70B), Replicate, Perplexity Labs, Cerebras, SambaNova. Pricing is typically a fraction of comparable closed-weights frontier-model API rates because providers compete on inference cost, not model rights.

Hyperscalers. AWS Bedrock, Azure AI (Microsoft was Meta's launch partner for Llama 2 in July 2023), Google Cloud Vertex AI, Oracle OCI, IBM watsonx. Meta's own Llama API (announced at LlamaCon 2025-04-29) is the first-party hosted option for the open-weights line. The Muse line splits in two. Its frontier models are API-only for now — Meta says an open-weights Muse Spark release is coming, but has not shipped it: the Meta Model API has served them to outside developers since July 9, 2026 and carries Muse Spark 1.3 as of September 8, 2026, alongside 1.2 and 1.1, with the max reasoning setting live on 1.3's standard tier since September 4; the Muse Code terminal agent is the first-party CLI on top of it, out of beta since August 31, 2026 and billable either per token or on a flat monthly subscription; OpenRouter is the first third-party route to them. Muse Image joined the same API on August 26, 2026 as muse-image-1.0, at $0.01 per image, through /v1/images/generations and /v1/images/edits; OpenRouter routes it as meta/muse-image. Muse Voice Transcribe joined on September 1, 2026 as muse-voice-transcribe-1.0, on its own endpoints — a WebSocket at wss://api.meta.ai/v1/asr/realtime for live audio and POST /v1/asr/transcribe for recordings — billed at $0.18 per hour of audio rather than per token, and also wired into Meta AI for Mac and Muse Code for dictation. Muse Glimmer, by contrast, is not on that API at all — since August 10, 2026 it is a download, and belongs to the self-host paragraph above. Muse Image and the Muse Spark models are also live in the Meta AI app and on meta.ai. Note that llama.com now redirects to developer.meta.com/ai/, and the license and use-policy paths redirect with it.

People who shaped Llama

Mark Zuckerberg — CEO of Meta. The 2024 “Open Source AI Is the Path Forward” letter and the 2025 strategic decisions (Scale AI acqui-hire, MSL formation, Muse Spark launch) all run through Zuckerberg's office. The closed-weights turn reverses the public position the letter staked out twenty-one months earlier.

Yann LeCun — Chief AI Scientist 2013–2025, FAIR cofounder. 2018 Turing Award. Departed November 2025 to launch AMI Labs (Advanced Machine Intelligence), a world-models startup that emerged from stealth in January 2026 and closed a $1.03B seed at a $3.5B pre-money valuation on March 9, 2026.

Joelle Pineau — led FAIR 2023–2025. The internal advocate for open-source releases through the Llama 1–3 era. Departed May 2025 ahead of the MSL reorganization.

Ahmad Al-Dahle — VP of Generative AI at Meta through Llama 2 / 3 / 4; spokesman for the LMArena Maverick episode in April 2025. Departed Meta in January 2026 to become CTO of Airbnb (Airbnb announcement, January 14, 2026).

Alexandr Wang — Chief AI Officer of Meta since June 2025; head of Meta Superintelligence Labs. Previously CEO of Scale AI. The architect of the closed-weights pivot and the Muse line.

Nat Friedman — head of AI products at MSL since June 2025; previously CEO of GitHub.

Shengjia Zhao — Chief Scientist of Meta Superintelligence Labs since July 25, 2025, setting the lab's research agenda alongside Zuckerberg and Wang. Previously at OpenAI, where he was a co-creator of ChatGPT and GPT-4 (CNBC, TechCrunch). His appointment four months before LeCun's departure left MSL's scientific direction with a hire from a closed-weights lab rather than with FAIR's founding generation.

The competitive landscape

Llama is — through April 2025 — the dominant open-weights frontier-AI line by deployment volume. The closest open-weights competitors are Mistral (French; mixed Apache 2.0 / Mistral Research License / proprietary tiers across the line, with the December 2025 “Mistral 3” family relaunch returning the open releases to Apache 2.0 — see Mistral Versions), DeepSeek (Chinese, MIT-licensed for the V3 / R1 line and onward, the December 2024 / January 2025 inflection — see DeepSeek Versions), Alibaba's Qwen (also fully open-weights and dominant on HuggingFace leaderboards through 2025 — see Qwen Versions), and xAI's Grok 1 (open-weights only for the first generation, see Grok Versions). The closed-weights frontier competitors — ChatGPT, Claude, Gemini — have all stayed closed-weights since their inception. With Muse Spark in April 2026, Meta is the only major frontier lab to have started open-weights and pivoted closed — and, with Muse Glimmer in August 2026, the only one to have pivoted partway back, releasing a local-tier model under Apache 2.0 while keeping its frontier models closed. Meta announced on August 10, 2026 that open weights for a version of the flagship Muse Spark 1.2 are coming, and has restated the plan twice since — most recently in the September 2, 2026 Muse Spark 1.3 launch post, which drops the version number and lists “the Muse Spark open weights release” as roadmap, and in a post from Zuckerberg the same day that drops it too. So the reopening is stated to reach frontier scale; whether it does in practice, for which version, and under which license, is the open question for the line going forward. This page does not attempt a benchmark roundup or a ranking.

Use Llama

The browser cannot detect which Llama model you've used or are using — there's no fingerprint or header that exposes it. The block below carries the practical information instead: the current open-weights model identifiers, a copy-paste self-host command, and the surfaces where Llama is available.

Current open-weights model identifiers

HuggingFace ids on the meta-llama org for the Llama line, and on the separate meta-models org for the open-weights Muse release. Verify against huggingface.co/meta-llama and huggingface.co/meta-models for the freshest list.

# Llama 4 — current open-weights flagship line
meta-llama/Llama-4-Maverick-17B-128E-Instruct
meta-llama/Llama-4-Scout-17B-16E-Instruct

# Llama 3.x — still widely served
meta-llama/Llama-3.3-70B-Instruct
meta-llama/Llama-3.1-{8B, 70B, 405B}-Instruct
meta-llama/Llama-3.2-{1B, 3B, 11B-Vision, 90B-Vision}-Instruct

# Specialized — coding and safety
meta-llama/Llama-Guard-4-12B
codellama/CodeLlama-70b-Instruct-hf

# Successor — open weights, Apache 2.0 (separate HuggingFace org)
meta-models/Muse-Glimmer-30B
                 # 30B local agent model; 131K context; also -GGUF,
                 #   -assistant, -ExecuTorch-PTE. No MAU or EU carve-out.

# Successor — closed-weights, API only (no HuggingFace release)
muse-spark-1.3   # current frontier; Meta Model API. Open-weights Muse
                 #   Spark release announced, no version/date/license named.
                 #   base URL https://api.meta.ai/v1 — docs at dev.meta.ai/docs
                 #   reasoning_effort now goes up to "max" (live 2026-09-04);
                 #   standard tier only — not on the -contributor tier
muse-spark-1.3-contributor
                 # cheaper tier — Meta trains on your prompts + completions
muse-spark-1.2   # superseded by 1.3, but still the one to use for AUDIO in;
                 #   1.3's audio support is flagged "not fully supported"
muse-spark-1.2-contributor
muse-spark-1.1   # superseded; still served
muse-voice-transcribe-1.0
                 # streaming speech-to-text + diarization; since 2026-09-01
                 #   wss://api.meta.ai/v1/asr/realtime, POST /v1/asr/transcribe
                 #   $0.18 per hour of audio, not per token
muse-image-1.0   # text-to-image + editing; on the API since 2026-08-26
                 #   $0.01/image — /v1/images/generations, /v1/images/edits
Muse Spark       # superseded; research.meta.ai/blog/introducing-muse-spark-msl

Quick self-host (Ollama)

Ollama wraps the HuggingFace download and the inference loop. For higher throughput, vLLM and TensorRT-LLM are the routine production choices.

$ brew install ollama          # macOS; apt / curl install on Linux
$ ollama pull llama3.3:70b
$ ollama run  llama3.3:70b "Hello, Llama."

Where to run Llama

Three categories — self-host, hosted-inference providers, and hyperscalers. Pricing varies by orders of magnitude; the open weights are the same across all of them.

# Self-host runtimes
https://ollama.com/                         # single-binary, easiest entry
https://github.com/ggerganov/llama.cpp      # CPU + GPU, edge-friendly
https://github.com/vllm-project/vllm        # production-grade throughput

# Hosted-inference providers
https://www.together.ai/
https://fireworks.ai/
https://groq.com/                            # very high tok/s on Llama 3 70B
https://replicate.com/

# Hyperscalers
AWS Bedrock, Azure AI, Google Cloud Vertex AI, Oracle OCI, IBM watsonx

# Meta's first-party Llama API
https://www.llama.com/                       # family hub + Llama API entry

Licensing

Each Llama generation ships with a successor “Llama Community License.” Read the actual PDF before shipping at scale — the >700M-MAU carve-out persists, and Llama 3.2 multimodal carries an EU restriction. Muse Glimmer is the one release that escapes both: it is plain Apache 2.0.

# Authoritative license texts
https://www.llama.com/license/
https://ai.meta.com/llama/license/

# Acceptable Use Policy applies to every Llama license
https://www.llama.com/use-policy/

# Notable carve-outs
>700M MAU on the release date → separate Meta license required
Llama 3.2 multimodal models → not granted to EU-domiciled users
Llama 4 (all models, natively multimodal) → same EU exclusion applies
Muse Glimmer   → Apache 2.0 open weights — NO MAU carve-out,
                 NO EU carve-out, no “Llama” naming requirement
Muse Spark 1.3 → proprietary, closed-weights; Meta Model API
                 standard tier → Meta does not train on your prompts
                 -contributor tier → cheaper; Meta DOES train on them
                 open-weights Muse Spark release announced 2026-08-10,
                 restated 2026-08-20 and 2026-09-02 — no version, no
                 date, no license named, nothing published as of 2026-09-08
Muse Spark 1.2 → proprietary, closed-weights; superseded by 1.3
Muse Spark 1.1 → proprietary, closed-weights; superseded
Muse Voice     → proprietary, closed-weights; Meta Model API,
  Transcribe     metered per hour of audio. No contributor tier.
Muse Spark     → proprietary, closed-weights, no redistribution
Muse Image     → proprietary, closed-weights; in-app and, since
                 2026-08-26, metered on the Meta Model API
Muse Code      → agent, not a model — no weights to license.
                 Out of beta 2026-08-31; pay per token, or a flat
                 monthly subscription bound to the CLI's own key

Sources: Llama family hub; per-release announcement posts at research.meta.ai/blog, ai.meta.com/blog, developer.meta.com/ai/resources/blog, and about.fb.com/news; Meta Model API docs at dev.meta.ai/docs; HuggingFace orgs meta-llama, codellama, and meta-models; license texts at llama.com/license; Purple Llama repo at github.com/meta-llama/PurpleLlama; contemporaneous reporting in NYT, WSJ, Bloomberg, The Information, CNBC, NPR, TechCrunch, The Register, Fortune, and DeepLearning.AI. Last updated September 8, 2026.

Mungomash LLC · More AI pages