View all articles
modelsopen-sourceedge-aireview

Google Gemma 4: The Open Model Family That Changed Our Entire Stack

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: ribbons of amber and white light weaving in a gentle curve. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:Google Gemma 4: análisis de la familia de modelos abiertos

When we reviewed Gemma 2 9B a few weeks ago, we called it the best small model for European business AI. It ran our scheduling tasks, fit on modest hardware, and handled instructions with reliability. Then Google announced Gemma 4 on April 2, 2026, making Gemma 2 look like a rough draft.

This is the model family we have been waiting for. It is not perfect, but it closes the gap between what small open models can do and what businesses need them to do. We replaced Gemma 2 in production within 48 hours. Here is what happened, what impressed us, and where the limits are.

Open source AI model comparison

Four Variants, Four Use Cases

Gemma 4 is not a single model. It is a family of four, each designed for a different tier of hardware and workload. This is Google doing what Google does best: scaling a single architecture across wildly different resource budgets.

VariantTotal ParamsEffective / ActiveContextAudioOllama SizeArena AI Rank
E2B5.1B (w/ embeddings)2.3B effective128KYes7.2 GB—
E4B8B (w/ embeddings)4.5B effective128KYes9.6 GB—
26B MoE25.2B total3.8B active (8 of 128 experts + 1 shared)256KNo19 GB#6 open
31B Dense30.7B30.7B (all dense)256KNo20 GB#3 open

The E2B and E4B variants are multimodal: they accept text, images, and audio as input and produce text output. The 26B and 31B handle text and images only. All four were pre-trained on 140+ languages and support configurable thinking modes, native function calling, structured JSON output, and system instructions.

xychart-beta
    title "Gemma 4 Variants — Parameters vs Memory"
    x-axis ["E2B (2.3B)", "E4B (4B)", "26B MoE", "31B Dense"]
    y-axis "Memory Required (GB)" 0 --> 25
    bar [7.2, 9.6, 19, 20]
Diagram

That 31B Dense model sitting at #3 on the Arena AI leaderboard among all open models worldwide is not a typo. The 26B MoE variant holds #6. Google’s own claim that Gemma 4 “outcompetes models 20x its size” sounds like marketing until you see the benchmark numbers.

How to Run Each Variant with Ollama

Getting started takes one command per variant. Sizes below are the default Q4_K_M downloads listed on the Ollama library page:

bash
# E2B — ultralight, our scheduling workhorse
ollama pull gemma4:e2b

# E4B — medium-duty, content and briefings
ollama pull gemma4:e4b

# 26B MoE — heavy reasoning with sparse activation
ollama pull gemma4:26b

# 31B Dense — maximum quality, needs beefy hardware
ollama pull gemma4:31b

The E2B download is 7.2 GB and the E4B 9.6 GB: both fit a 16 GB machine, the E4B with little headroom, and are comfortable on 32 GB. The 26B and 31B variants need 19-20 GB of VRAM or unified memory, which puts them squarely in Mac Studio or dedicated GPU territory. For hardware guidance, see our edge AI hardware guide.

What We Actually Run in Production

At VORLUX AI, we run Gemma 4 E2B and E4B on a Mac Mini M4 as part of our local AI infrastructure. Here is how they fit:

Gemma 4 E2B is our primary scheduling model. It handles 58 orchestrator jobs: task routing, status updates, lightweight classification, and JSON-structured outputs for downstream agents. At 2.3B effective parameters, it is absurdly fast. Response times average under 800ms for typical scheduling prompts. It replaced Gemma 2 9B for these tasks.

Gemma 4 E4B is our medium-duty model for briefings, content drafting, and multi-step analysis. When a task needs more reasoning than E2B can provide but does not justify pulling in a 26B+ model, E4B handles it. The 128K context window means we can feed it full documents without chunking.

The Mac Mini M4 runs both simultaneously with room to spare. That would have been unthinkable a year ago.

What Gemma 4 Does Exceptionally Well

Function calling and structured output. Native support, not bolted on. We feed Gemma 4 a tool schema and it returns valid JSON function calls consistently. This matters enormously for agent orchestration, no more regex parsing of freeform text.

Instruction following. The configurable thinking modes let us toggle between fast responses (thinking off) and deliberate reasoning (thinking on) per request. For scheduling, we keep thinking off. For content analysis, we turn it on.

Multilingual performance. With 140+ languages, our Spanish and English workflows run on the same model without fine-tuning. For a consultancy based in Valencia serving Spanish SMEs, this is not a nice-to-have: it is essential.

Audio input on E2B/E4B. We have not deployed this in production yet, but the ability to process audio natively opens doors for meeting transcription, voice-driven workflows, and accessibility features without a separate speech-to-text pipeline.

Where It Falls Short: Honest Limits

Deep reasoning and complex coding. For multi-step mathematical proofs or competitive programming challenges, Gemma 4 31B is strong but still trails Llama 3.3 70B and Qwen 2.5 72B Coder. If your primary workload is code generation, Qwen remains the better specialized choice. Gemma 4 is a generalist that happens to code well: it is not a coding specialist.

The 26B MoE trade-off. The Mixture-of-Experts architecture is brilliant for efficiency: only 3.8B of the 25.2B parameters activate per token. But MoE models can be unpredictable on tasks that fall between expert boundaries. We have seen occasional inconsistency on hybrid tasks that the 31B Dense handles cleanly.

Text output only. Gemma 4 can understand images (and, on E2B/E4B, audio) as input, but it only generates text. If you need image generation or audio synthesis, you still need separate models.

VRAM pressure on larger variants. The 31B Dense at 20 GB is tight on a 32 GB Mac. Running it alongside other models requires careful memory management. Check our Q2 2026 model comparison for side-by-side VRAM budgets.

Gemma 4 vs Gemma 2: Is It Worth Upgrading?

Unequivocally yes. The E2B alone makes Gemma 2 9B redundant for most scheduling and classification tasks: it is faster, smaller, and more capable. The 128K context window (up from 8K on Gemma 2) eliminates the chunking workarounds we used to need. Function calling support means we deleted hundreds of lines of output-parsing code. And the multilingual quality jumped from “functional” to “genuinely good.”

If you are currently running Gemma 2, the migration path is straightforward: pull the new model, test your prompts, and switch. We did it in a weekend.

Who Should Use Which Variant?

  • E2B: Edge devices, scheduling, classification, IoT, mobile. Anything where speed and size matter more than depth.
  • E4B: SME workstations, content generation, briefings, customer support. The sweet spot for most business use cases.
  • 26B MoE: Research, analysis, long-document processing. Great when you need 256K context but want to keep memory reasonable.
  • 31B Dense: Maximum quality on demanding tasks. Translation, complex analysis, multi-turn reasoning. Worth it if you have the hardware.

Getting Started

Gemma 4 is available now on Ollama and Google’s official announcement has the full technical details. All variants are released under the Apache 2.0 license, which permits commercial use.

Note: on June 3, 2026 Google added a fifth size, Gemma 4 12B Unified, with native audio input. It is not covered in this review.

If you want help deploying Gemma 4 on your own hardware (whether that is a single Mac Mini or a fleet of edge devices) that is exactly what we do. We build local AI systems for European SMEs that keep data on-premises and costs predictable. See our services or get in touch for a free consultation.

The era of needing massive cloud budgets for capable AI is ending. Gemma 4 is proof.

Next steps

  • Test the E2B and E4B variants on your own hardware using the Ollama commands provided.
  • Evaluate the 128K context window for your specific classification and scheduling tasks.
  • Compare these results against our previous review of the multimodal Gemma 3 model.
  • Check the n8n workflow tutorial to see how to integrate structured JSON output into your automation.

Work with us

We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates