View all articles
modelsopen-sourcemultimodalreview

Llama 4 Scout and Maverick: A Practical Review for Local AI Deployment

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: nested glass spheres with an amber light at the centre. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:Llama 4 Scout y Maverick: análisis para despliegue local

On April 5, 2025, Meta released Llama 4 Scout and Llama 4 Maverick. These models use a Mixture-of-Experts (MoE) architecture to keep active parameter counts low while competing with proprietary models. Both are natively multimodal, handling text and images without extra adapters.

For companies running AI locally for GDPR compliance, cost control, or latency requirements, these models are a practical step. Scout offers a 10-million-token context window and, per Meta, fits on a single H100 with Int4 quantization. Maverick, which Meta reports ahead of GPT-4o on its launch benchmarks, has 128 experts and only 17 billion active parameters per forward pass.

We will break down the architecture, benchmarks, and practical deployment considerations.

Architecture: How MoE Changes the Game

Both Scout and Maverick use a Mixture-of-Experts architecture. Instead of activating every parameter for every token, MoE models route each input through a subset of specialized “expert” sub-networks. The result: massive total parameter counts with efficient inference costs.

flowchart TB
    subgraph Input["Input Layer"]
        T[Token / Image Patch]
    end

    subgraph Router["Gating Router"]
        R[Expert Selection]
    end

    subgraph ScoutExperts["Scout: 16 Experts"]
        direction LR
        SE1[Expert 1]
        SE2[Expert 2]
        SE3["..."]
        SE16[Expert 16]
    end

    subgraph MaverickExperts["Maverick: 128 Experts"]
        direction LR
        ME1[Expert 1]
        ME2[Expert 2]
        ME3["..."]
        ME128[Expert 128]
    end

    subgraph Active["Active Parameters: 17B"]
        AP[Selected Experts Process Token]
    end

    subgraph Output["Output"]
        O[Combined Result]
    end

    T --> R
    R -->|"Scout"| ScoutExperts
    R -->|"Maverick"| MaverickExperts
    ScoutExperts --> AP
    MaverickExperts --> AP
    AP --> O

    style Input fill:#0B1628,stroke:#F5A623,color:#fff
    style Router fill:#0B1628,stroke:#F5A623,color:#fff
    style ScoutExperts fill:#1a2744,stroke:#4a90d9,color:#fff
    style MaverickExperts fill:#1a2744,stroke:#4a90d9,color:#fff
    style Active fill:#0B1628,stroke:#2ecc71,color:#fff
    style Output fill:#0B1628,stroke:#F5A623,color:#fff
Diagram

The key insight: both models activate only 17 billion parameters per forward pass, regardless of their total size. This is what makes MoE models practical for local deployment: you pay inference costs proportional to the active parameters, not the total.

SpecificationScoutMaverick
Active parameters17B17B
Total experts16128
Total parameters109B400B
Context window10M tokens1M tokens
Training tokens40T22T
ModalitiesText + ImageText + Image
ArchitectureMoEMoE
Release dateApril 5, 2025April 5, 2025

Scout was trained on 40 trillion tokens, nearly double Maverick’s 22 trillion, which contributes to its strong performance on knowledge-heavy benchmarks despite having fewer experts. Maverick compensates with 8x more experts, giving it better specialization across diverse task types.

Benchmark Comparison: Where Each Model Excels

The scores below are Meta’s own, from the instruction-tuned Llama 4 model card (checked 2026-10-09). They are vendor numbers, not our measurements.

BenchmarkScoutMaverick
MMMU (image reasoning)69.473.4
MathVista (visual math)70.773.7
ChartQA (chart comprehension)88.890.0
DocVQA (document QA)94.494.4
LiveCodeBench (coding)32.843.4
MMLU Pro (reasoning)74.380.5
GPQA Diamond (reasoning)57.269.8

In its launch post, Meta says Maverick beats GPT-4o and Gemini 2.0 Flash “across a broad range of widely reported benchmarks”, and that Scout beats Gemma 3, Gemini 2.0 Flash-Lite and Mistral 3.1. Treat those as vendor claims and test on your own documents.

The gap is widest on coding and reasoning, where Maverick’s extra experts show. On document QA (DocVQA) the two tie at 94.4, so for document processing on moderate hardware, Scout is the cheaper choice.

The 10-Million-Token Context Window

Scout’s 10-million-token context window is not a marketing number. It enables use cases that were previously impossible with open-source models:

  • Full codebase analysis: Load an entire mid-size repository (100,000+ lines) into a single prompt
  • Book-length document processing: Analyze complete legal contracts, technical manuals, or regulatory frameworks without chunking
  • Multi-document synthesis: Cross-reference dozens of documents simultaneously

For compliance workloads (where you might need to analyze an entire regulatory text against your company’s documentation) this is transformative. No RAG pipeline, no chunking strategy, no information loss from retrieval. Just the full document in context.

Maverick’s 1-million-token context is still substantial, covering most production use cases. The tradeoff is clear: Scout for context-heavy work, Maverick for quality-heavy work.

What About Muse Spark?

Meta Superintelligence Labs released Muse Spark on April 8, 2026 (Meta AI blog), a multimodal reasoning model. It runs inside Meta AI and is offered as a private-preview API, not as downloadable weights. So for local deployment, Scout and Maverick remain Meta’s options.

Meta has not announced open weights for a successor, so Llama 4 is the newest Meta family you can run on your own hardware. Compare it with newer open models from other labs before you commit.

Deployment Considerations for Local Hardware

Running these models locally requires careful hardware planning:

Scout (109B total, 17B active): Meta says it fits on a single H100 with Int4 quantization. Ollama’s default llama4:scout build (Q4_K_M) is 67 GB, so it needs one 80 GB GPU or a machine with 96 GB+ of unified memory. Two 24 GB consumer cards (48 GB) are not enough. Long contexts add KV cache on top of the weights.

Maverick (400B total, 17B active): Despite 400B total parameters, only 17B are active per token, so per-token compute is similar to Scout. Memory is not: Ollama lists the FP16 build at 803 GB, Q8_0 at 428 GB and Q4_K_M at 245 GB (ollama.com/library/llama4, checked 2026-10-09). That means a multi-GPU server, and the comparison of local models shows smaller options for most SMEs.

Both models work with vLLM, llama.cpp and Ollama. The Ollama tags that exist are llama4:scout (alias llama4:16x17b, 67 GB) and llama4:maverick (alias llama4:128x17b, 245 GB):

bash
# Scout via Ollama (Q4_K_M, 67 GB download)
ollama pull llama4:scout

# Maverick (Q4_K_M, 245 GB download: multi-GPU server only)
# ollama pull llama4:maverick

# Test with a simple prompt
curl http://localhost:11434/api/generate -d '{
  "model": "llama4:scout",
  "prompt": "Summarize the key requirements of the EU AI Act"
}'

At VORLUX AI, we deploy these models on edge hardware for our SME clients, optimized for their specific workloads.

The Bottom Line

Llama 4 Scout and Maverick, released in April 2025, are Meta’s open-weight models for local deployment. Scout’s 10M context window and training on 40T tokens make it ideal for knowledge-heavy, document-intensive workloads. Maverick’s 128 experts and stronger coding and reasoning scores make it the right choice when quality is the priority.

Both models share the same fundamental advantage: they are open-source, run locally, and keep your data on your hardware. In a regulatory environment where GDPR fines total EUR 7.1 billion and the EU AI Act adds another penalty layer, that is not just a technical preference: it is a business requirement.


Want to deploy Llama 4 on your own infrastructure? Contact VORLUX AI for a hardware assessment and deployment plan tailored to your workload. We handle the optimization so you get production-grade performance on hardware you control.

Sources: Meta launch post, 5 April 2025 · Llama 4 model card · Ollama llama4 tags · Llama 4 Official (Meta) · Llama 4 on HuggingFace · Scout vs Maverick (RunPod)


Work with us

We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates