
Llama 4
Open-weight multimodal LLM family
Open-weight multimodal models you can run cheap at massive context
Reviewed by The Desk · Last verified July 2026
“A genuinely open-weight, natively multimodal MoE family with class-leading context length and rock-bottom inference costs, but it launched into a benchmark-credibility controversy and never delivered the frontier crown its hype promised. Treat it as a pragmatic, cheap building block rather than the best model you can run.”
Skip it ifSkip it if you need a built-in reasoning mode or true frontier-tier quality: Llama 4 shipped without a dedicated reasoning variant, its splashy launch benchmarks came from an experimental LMArena build rather than the open weights, and the flagship Behemoth remains unreleased. EU-based users are also barred outright by the license.

What it is
Llama 4 is Meta's April 2025 family of natively multimodal, open-weight language models and the company's first built on a mixture-of-experts (MoE) architecture. Two variants shipped publicly: Scout (17B active parameters across 16 experts, 109B total, with a headline 10M-token context window) and Maverick (17B active across 128 experts, 400B total, up to 1M tokens). A far larger teacher model, Behemoth, was announced as still in training and, as of 2026, has not been released. All variants accept both text and image inputs.
The appeal is concrete: genuinely downloadable weights, an efficient MoE design that keeps only about 17B parameters active per token, unusually long context windows, and very low per-token pricing across a wide field of inference hosts such as Groq, Together AI, Fireworks, Bedrock, and Vertex AI. For teams that want a private, self-hostable model or cheap high-volume inference, it's a credible building block.
The caveats are equally concrete. Llama 4's launch drew criticism after it emerged that the competitive benchmark numbers came from an experimental LMArena build rather than the released weights, raising questions about real-world quality. The family also ships without a dedicated reasoning model at a time when rivals lean into that, Behemoth's delay left the lineup without a true flagship, and the community license bars EU users and 700M+ MAU companies, so it falls short of OSI open source. It's a pragmatic, low-cost option, not the frontier leader its rollout suggested.
Key features
- Natively multimodal: accepts text and image inputs
- First Meta family built on a mixture-of-experts (MoE) architecture
- Scout: 17B active / 16 experts / 109B total, up to 10M-token context
- Maverick: 17B active / 128 experts / 400B total, up to 1M-token context
- Behemoth: 288B active teacher model (still unreleased)
- Open weights downloadable for self-hosting
- Multilingual text support
- Available across major cloud and inference providers
Who it’s for
Building multimodal chat assistants that accept text and images
Analyzing very long documents or codebases in a single context window
Self-hosted or private LLM deployments where data can't leave your infra
High-volume, cost-sensitive inference at scale
Pros & cons
The good
- Class-leading context windows (up to 10M tokens on Scout) plus native image understanding
- Genuinely open weights with very low per-token costs across many inference providers
- Efficient MoE design keeps only ~17B parameters active, cutting inference cost versus dense models
The catch
- Launch benchmarks used an experimental build, not the released weights, denting trust in the numbers
- No dedicated reasoning model, and the flagship Behemoth stayed shelved/unreleased
- Restrictive 'community' license bars EU users and 700M+ MAU companies, so it isn't OSI open source
Pricing
| Tier | Price | What you get |
|---|---|---|
| Self-host / open weights | Free | Free weights via llama.com and Hugging Face · Run on your own infrastructure · Community license restrictions apply (EU ban, 700M-MAU clause) |
| Scout via API providers | Custom | ~$0.08-0.15 per 1M input tokens (varies by host) · ~$0.30-0.60 per 1M output tokens · Up to 10M-token context |
| Maverick via API providers | Custom | ~$0.15-0.27 per 1M input tokens (varies by host) · ~$0.60-0.85 per 1M output tokens · Up to 1M-token context |
Free tier — Model weights are free to download and self-host under Meta's community license; Meta AI chat access is also free
Alternatives to Llama 4
More Chatbots & LLMs →Questions people ask
Is Llama 4 free to use?
The model weights are free to download and self-host under Meta's community license, and Meta AI chat access is free. Running it through third-party API providers is paid but inexpensive, typically a fraction of a dollar per million tokens.
What models are in the Llama 4 family?
Scout (17B active parameters, 16 experts, up to a 10M-token context), Maverick (17B active, 128 experts, up to 1M-token context), and Behemoth, a 288B-active-parameter teacher model that Meta has not publicly released.
Is Llama 4 actually open source?
It's open-weight, not OSI-approved open source. Meta's community license restricts use (EU-based users are prohibited, and companies over 700M monthly active users need a separate license), which the Open Source Initiative says disqualifies it as true open source.
How does Llama 4 compare to GPT-4o and DeepSeek?
Meta reported Maverick beating GPT-4o and matching DeepSeek v3 on several multimodal benchmarks at far lower cost, but those numbers came from an experimental build, and Llama 4 lacks a dedicated reasoning mode that rivals increasingly ship.
Can I use Llama 4 in the European Union?
No. The Llama 4 community license prohibits individuals and companies domiciled in the EU from using the multimodal models.

