On April 5, 2025, Meta released Llama 4 Scout and Llama 4 Maverick. These models use a Mixture-of-Experts (MoE) architecture to keep active parameter counts low while competing with proprietary models. Both are natively multimodal, handling text and images without extra adapters.
For companies running AI locally for GDPR compliance, cost control, or latency requirements, these models are a practical step. Scout offers a 10-million-token context window and, per Meta, fits on a single H100 with Int4 quantization. Maverick, which Meta reports ahead of GPT-4o on its launch benchmarks, has 128 experts and only 17 billion active parameters per forward pass.
We will break down the architecture, benchmarks, and practical deployment considerations.
Architecture: How MoE Changes the Game
Both Scout and Maverick use a Mixture-of-Experts architecture. Instead of activating every parameter for every token, MoE models route each input through a subset of specialized “expert” sub-networks. The result: massive total parameter counts with efficient inference costs.
flowchart TB
subgraph Input["Input Layer"]
T[Token / Image Patch]
end
subgraph Router["Gating Router"]
R[Expert Selection]
end
subgraph ScoutExperts["Scout: 16 Experts"]
direction LR
SE1[Expert 1]
SE2[Expert 2]
SE3["..."]
SE16[Expert 16]
end
subgraph MaverickExperts["Maverick: 128 Experts"]
direction LR
ME1[Expert 1]
ME2[Expert 2]
ME3["..."]
ME128[Expert 128]
end
subgraph Active["Active Parameters: 17B"]
AP[Selected Experts Process Token]
end
subgraph Output["Output"]
O[Combined Result]
end
T --> R
R -->|"Scout"| ScoutExperts
R -->|"Maverick"| MaverickExperts
ScoutExperts --> AP
MaverickExperts --> AP
AP --> O
style Input fill:#0B1628,stroke:#F5A623,color:#fff
style Router fill:#0B1628,stroke:#F5A623,color:#fff
style ScoutExperts fill:#1a2744,stroke:#4a90d9,color:#fff
style MaverickExperts fill:#1a2744,stroke:#4a90d9,color:#fff
style Active fill:#0B1628,stroke:#2ecc71,color:#fff
style Output fill:#0B1628,stroke:#F5A623,color:#fffThe key insight: both models activate only 17 billion parameters per forward pass, regardless of their total size. This is what makes MoE models practical for local deployment: you pay inference costs proportional to the active parameters, not the total.
| Specification | Scout | Maverick |
|---|---|---|
| Active parameters | 17B | 17B |
| Total experts | 16 | 128 |
| Total parameters | 109B | 400B |
| Context window | 10M tokens | 1M tokens |
| Training tokens | 40T | 22T |
| Modalities | Text + Image | Text + Image |
| Architecture | MoE | MoE |
| Release date | April 5, 2025 | April 5, 2025 |
Scout was trained on 40 trillion tokens, nearly double Maverick’s 22 trillion, which contributes to its strong performance on knowledge-heavy benchmarks despite having fewer experts. Maverick compensates with 8x more experts, giving it better specialization across diverse task types.
Benchmark Comparison: Where Each Model Excels
The scores below are Meta’s own, from the instruction-tuned Llama 4 model card (checked 2026-10-09). They are vendor numbers, not our measurements.
| Benchmark | Scout | Maverick |
|---|---|---|
| MMMU (image reasoning) | 69.4 | 73.4 |
| MathVista (visual math) | 70.7 | 73.7 |
| ChartQA (chart comprehension) | 88.8 | 90.0 |
| DocVQA (document QA) | 94.4 | 94.4 |
| LiveCodeBench (coding) | 32.8 | 43.4 |
| MMLU Pro (reasoning) | 74.3 | 80.5 |
| GPQA Diamond (reasoning) | 57.2 | 69.8 |
In its launch post, Meta says Maverick beats GPT-4o and Gemini 2.0 Flash “across a broad range of widely reported benchmarks”, and that Scout beats Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1. Treat those as vendor claims and test on your own documents.
The gap is widest on coding and reasoning, where Maverick’s extra experts show. On document QA (DocVQA) the two tie at 94.4, so for document processing on moderate hardware, Scout is the cheaper choice.
The 10-Million-Token Context Window
Scout’s 10-million-token context window is not a marketing number. It enables use cases that were previously impossible with open-source models:
- Full codebase analysis: Load an entire mid-size repository (100,000+ lines) into a single prompt
- Book-length document processing: Analyze complete legal contracts, technical manuals, or regulatory frameworks without chunking
- Multi-document synthesis: Cross-reference dozens of documents simultaneously
For compliance workloads (where you might need to analyze an entire regulatory text against your company’s documentation) this is transformative. No RAG pipeline, no chunking strategy, no information loss from retrieval. Just the full document in context.
Maverick’s 1-million-token context is still substantial, covering most production use cases. The tradeoff is clear: Scout for context-heavy work, Maverick for quality-heavy work.
What About Muse Spark?
Meta Superintelligence Labs released Muse Spark on April 8, 2026 (Meta AI blog), a multimodal reasoning model. It runs inside Meta AI and is offered as a private-preview API, not as downloadable weights. So for local deployment, Scout and Maverick remain Meta’s options.
Meta has not announced open weights for a successor, so Llama 4 is the newest Meta family you can run on your own hardware. Compare it with newer open models from other labs before you commit.
Deployment Considerations for Local Hardware
Running these models locally requires careful hardware planning:
Scout (109B total, 17B active): Meta says it fits on a single H100 with Int4 quantization. Ollama’s default llama4:scout build (Q4_K_M) is 67 GB, so it needs one 80 GB GPU or a machine with 96 GB+ of unified memory. Two 24 GB consumer cards (48 GB) are not enough. Long contexts add KV cache on top of the weights.
Maverick (400B total, 17B active): Despite 400B total parameters, only 17B are active per token, so per-token compute is similar to Scout. Memory is not: Ollama lists the FP16 build at 803 GB, Q8_0 at 428 GB and Q4_K_M at 245 GB (ollama.com/library/llama4, checked 2026-10-09). That means a multi-GPU server, and the comparison of local models shows smaller options for most SMEs.
Both models work with vLLM, llama.cpp and Ollama. The Ollama tags that exist are llama4:scout (alias llama4:16x17b, 67 GB) and llama4:maverick (alias llama4:128x17b, 245 GB):
# Scout via Ollama (Q4_K_M, 67 GB download)
ollama pull llama4:scout
# Maverick (Q4_K_M, 245 GB download: multi-GPU server only)
# ollama pull llama4:maverick
# Test with a simple prompt
curl http://localhost:11434/api/generate -d '{
"model": "llama4:scout",
"prompt": "Summarize the key requirements of the EU AI Act"
}'At VORLUX AI, we deploy these models on edge hardware for our SME clients, optimized for their specific workloads.
The Bottom Line
Llama 4 Scout and Maverick, released in April 2025, are Meta’s open-weight models for local deployment. Scout’s 10M context window and training on 40T tokens make it ideal for knowledge-heavy, document-intensive workloads. Maverick’s 128 experts and stronger coding and reasoning scores make it the right choice when quality is the priority.
Both models share the same fundamental advantage: they are open-source, run locally, and keep your data on your hardware. In a regulatory environment where GDPR fines total EUR 7.1 billion and the EU AI Act adds another penalty layer, that is not just a technical preference: it is a business requirement.
Want to deploy Llama 4 on your own infrastructure? Contact VORLUX AI for a hardware assessment and deployment plan tailored to your workload. We handle the optimization so you get production-grade performance on hardware you control.
Sources: Meta launch post, 5 April 2025 · Llama 4 model card · Ollama llama4 tags · Llama 4 Official (Meta) · Llama 4 on HuggingFace · Scout vs Maverick (RunPod)
Related reading
- Your First 3 AI Agents: A Local Deployment Guide for SMEs (2026)
- AI Agents for SMEs: Where the ROI Really Comes From
- Kit Digital and Spanish AI Grants Guide 2026: Fund Your AI Deployment for Free
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.