The open-source model landscape changed significantly in three months. Qwen 3 brought MoE to the masses, Gemma 4 improved small-model quality, and Llama 4 Scout increased the context window. Here is how they compare for local deployment, and which one you should pick.

flowchart TD
START["What is your primary task?"] --> CODE{"Code generation?"}
START --> OFFICE{"Office assistant\n(emails, docs, Q&A)?"}
START --> REASON{"Complex reasoning\nor math?"}
START --> DOCS{"Massive documents\n(contracts, research)?"}
START --> QUALITY{"Maximum quality\n(no hardware limits)?"}
CODE -->|Yes| CODER["Qwen 2.5 Coder 7B\n4.7 GB download"]
OFFICE --> LANG{"Need multilingual\n(Spanish, etc.)?"}
LANG -->|Yes| QWEN["Qwen 3 8B\n5.2 GB download"]
LANG -->|No| GEMMA["Gemma 4 E4B\n9.5 GB download"]
REASON -->|Yes| PHI["Phi-4 14B\n9.1 GB download"]
DOCS -->|Yes| LLAMA["Llama 4 Scout 109B\n67 GB download — 10M context"]
QUALITY -->|Yes| DS["DeepSeek V3 671B\n404 GB download"]
style START fill:#DBEAFE,stroke:#2563EB,color:#000
style CODER fill:#D1FAE5,stroke:#059669,color:#000
style QWEN fill:#D1FAE5,stroke:#059669,color:#000
style GEMMA fill:#D1FAE5,stroke:#059669,color:#000
style PHI fill:#FEF3C7,stroke:#F5A623,color:#000
style LLAMA fill:#FECACA,stroke:#B91C1C,color:#000
style DS fill:#FECACA,stroke:#B91C1C,color:#000The Contenders
| Model | Params | Ollama download (Q4) | Strength |
|---|---|---|---|
| Qwen 3 8B | 8B | 5.2 GB | Multilingual: 119 languages and dialects per the Qwen team |
| Gemma 4 E4B | ~4B effective | 9.5 GB | Strong quality for its size; text, image and audio |
| Phi-4 | 14B | 9.1 GB | Reasoning and math (Microsoft’s benchmarks) |
| Llama 4 Scout | 109B (17B active) | 67 GB | 10M token context window (Meta’s figure) |
| DeepSeek V3 | 671B (37B active) | 404 GB | Frontier-class open model; needs server hardware |
| Qwen 2.5 Coder 7B | 7.6B | 4.7 GB | Code generation |
Download sizes are from the Ollama library default tags. Add memory for the context window and your operating system. Speed depends on your hardware, so measure it with ollama run <model> --verbose. Everything down to Phi-4 runs on a 24 GB Mac mini; Llama 4 Scout needs a 96-128 GB machine and DeepSeek V3 a multi-GPU server.
Our Pick by Use Case
For a Spanish SME office assistant
Winner: Qwen 3 8B
Why: strong Spanish support (the Qwen team lists 119 languages and dialects), a 5.2 GB download that runs comfortably on 24GB hardware, Apache 2.0 license for commercial use. Handles email drafting, customer Q&A, document summaries, and internal queries without breaking a sweat.
ollama pull qwen3:8bFor code generation and technical work
Winner: Qwen 2.5 Coder 7B
Why: purpose-built for code, a 4.7GB download. Handles Python, JavaScript, TypeScript, SQL and many other languages. We measured it in our hands-on review.
ollama pull qwen2.5-coder:7bFor complex reasoning and analysis
Winner: Phi-4 (14B)
Why: Microsoft’s Phi-4 punches far above its weight: 80.4% on MATH in Microsoft’s own benchmarks, ahead of Llama 3.3 70B’s 66.3% in the same table. Needs 16GB RAM but delivers strong reasoning for strategy documents, legal analysis, and financial modeling.
ollama pull phi4Alternative: DeepSeek R1 14B (distilled). DeepSeek’s technical report gives the 14B distill 93.9% on MATH-500 (97.3% for the full R1). It is a 9 GB download and shows its reasoning step by step, which is slower but easier to check.
ollama pull deepseek-r1:14bFor maximum quality (when you have server hardware)
Winner: DeepSeek V3
Why: MoE architecture activates only 37B of 671B parameters per token, so it generates faster than its size suggests. But all 671B parameters must sit in memory: the Q4 download alone is 404 GB, which means a multi-GPU server or a machine with 512 GB of unified memory. Best for complex research, multi-step analysis, and content where quality matters more than speed.
For massive documents (contracts, research papers)
Winner: Llama 4 Scout
Why: Meta rates it for a 10 million token context window. In practice the memory for a long context comes on top of the 67 GB model, so you will run far shorter contexts locally. Needs a 96-128 GB machine.
Hardware Requirements at a Glance
| Your Hardware | Best Model | What You Can Do |
|---|---|---|
| 8GB RAM (Jetson Orin Nano) | Qwen 2.5 3B | Basic Q&A, classification |
| 24GB RAM (Mac Mini M4) | Qwen 3 8B or Gemma 4 E4B | Full office assistant |
| 48GB RAM (Mac Mini M4 Pro) | Phi-4 14B or Gemma 4 26B | Complex reasoning |
| 128GB RAM (DGX Spark / Jetson AGX Thor) | Llama 4 Scout 109B | Enterprise-grade |
Quick-Start Tip
If you’re deploying your first local model, start with Ollama: it handles downloading, quantization, and serving in a single command. Install it from ollama.com, then run ollama pull qwen3:8b. Within five minutes you’ll have a production-ready model answering queries on localhost:11434. From there, connect it to n8n for workflow automation or build a simple RAG pipeline for your internal documents.
The Bottom Line
For a typical SME office assistant, Qwen 3 8B on a Mac Mini M4 is the sweet spot: a one-time hardware cost plus electricity, with no per-query fees. Whether that beats a cloud API depends on your volume; our cost analysis shows how to work it out.
For routine business tasks the gap between local and cloud models is now small enough that data control and fixed costs often decide it. Test on your own documents first.
Related reading
- Cloud vs Local AI: Real Cost Analysis for Spanish SMEs in 2026
- DeepSeek R1: The Best Open-Source Reasoning Model You Can Run Locally
- Local AI Readiness Checklist: Is Your Business Ready to Run AI On-Premise?
Related resources
- 50 AI Models catalog: browse all models with VRAM and install commands
- Hardware catalog: 17 devices from EUR 200 to cloud GPU
- Software stack: Ollama, MLX, and the tools we use
- ROI Calculator: compare local vs cloud costs for your usage
- Contact: need help choosing and deploying?
Sources: Ollama Library · Qwen3 announcement · Phi-4 model card
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.