Updated 4 October 2026. This post is from April 2026 and names GPT-4o, Claude 3.5, Llama 3.1, Qwen 2.5, Qwen2.5 as current. The method and the advice still hold; today you would run it with:
- DeepSeek V4: Use this if you need even higher performance for complex reasoning, though memory requirements remain significant.
- Llama 4 Scout / Maverick: These models offer a more practical alternative for local deployment on smaller, more accessible hardware.
- Gemma 4: A smaller, efficient option for developers needing high-quality performance on constrained local infrastructure.
DeepSeek V3 changed the conversation about what open-weight models can achieve. Built by a Chinese AI lab, it was trained on 14.8 trillion tokens using 2.788 million H800 GPU hours. In DeepSeek’s own evaluation, this model matches or beats GPT-4o on several hard benchmarks, specifically coding and mathematics.
At 671 billion parameters, it is far too large for any local workstation. We should be upfront about what this model delivers, what it costs to run, and what European businesses should consider before adopting it.

flowchart LR
INPUT["Input Token"] --> ROUTER["Router Network"]
ROUTER --> E1["Expert 1"]
ROUTER --> E2["Expert 2"]
ROUTER -.->|inactive| E3["Expert 3"]
ROUTER -.->|inactive| E4["..."]
ROUTER -.->|inactive| E128["Expert 128"]
E1 --> COMBINE["Combine Outputs"]
E2 --> COMBINE
subgraph TOTAL ["671B Total Parameters"]
ROUTER
E1
E2
E3
E4
E128
end
subgraph ACTIVE ["37B Active per Token"]
E1
E2
end
COMBINE --> MTP["Multi-Token\nPrediction"]
MTP --> OUTPUT["Output Tokens"]
style INPUT fill:#DBEAFE,stroke:#2563EB,color:#000
style ROUTER fill:#FEF3C7,stroke:#F5A623,color:#000
style E1 fill:#D1FAE5,stroke:#059669,color:#000
style E2 fill:#D1FAE5,stroke:#059669,color:#000
style E3 fill:#FECACA,stroke:#B91C1C,color:#000
style E4 fill:#FECACA,stroke:#B91C1C,color:#000
style E128 fill:#FECACA,stroke:#B91C1C,color:#000
style COMBINE fill:#FEF3C7,stroke:#F5A623,color:#000
style MTP fill:#DBEAFE,stroke:#2563EB,color:#000
style OUTPUT fill:#D1FAE5,stroke:#059669,color:#000How Mixture-of-Experts works (in plain terms)
Most language models are “dense” — every parameter participates in every computation. DeepSeek V3 uses a different approach called Mixture-of-Experts (MoE). Think of it like a company with 671 billion employees, but for any given task, only 37 billion of them actually work on it. A routing network decides which “expert” sub-networks handle each token.
The result: you get the knowledge depth of a 671B model with the inference speed closer to a 37B dense model. DeepSeek also introduced Multi-Token Prediction (MTP) for faster generation and FP8 mixed-precision training to keep costs down. The model supports a 128K context window, which is generous for long-document tasks.
Real benchmarks — honest numbers
These numbers come directly from the official HuggingFace model card and the DeepSeek technical report.
| Benchmark | DeepSeek V3 | GPT-4o (0513) | Claude 3.5 Sonnet (1022) | Llama 3.1 405B | Qwen2.5 72B |
|---|---|---|---|---|---|
| MMLU (Chat) | 88.5% | 87.2% | 88.3% | 88.6% | 85.3% |
| MMLU (Base, 5-shot) | 87.1% | — | — | 84.4% | — |
| MMLU-Pro | 75.9% | 72.6% | 78.0% | 73.3% | 71.6% |
| GPQA Diamond | 59.1% | 49.9% | 65.0% | 51.1% | 49.0% |
| MATH-500 | 90.2% | 74.6% | 78.3% | 73.8% | 80.0% |
| AIME 2024 | 39.2% | 9.3% | 16.0% | 23.3% | 23.3% |
| HumanEval-Mul (code) | 82.6% | 80.5% | 81.7% | 77.2% | 77.3% |
| LiveCodeBench (Pass@1-COT) | 40.5% | 33.4% | 36.3% | 28.4% | 31.1% |
| Codeforces (percentile) | 51.6 | 23.6 | 20.3 | 25.3 | 24.8 |
| Arena Hard | 85.5 | 80.4 | 85.2 | 69.3 | 81.2 |
Sources: DeepSeek V3 on HuggingFace, DeepSeek technical report (arXiv). All columns are DeepSeek’s own measurements, published with the model in December 2024; newer versions of the other models score differently.
The standout results: DeepSeek V3 scores 90.2% on MATH-500 (well ahead of GPT-4o’s 74.6% in the same table), and 51.6 percentile on Codeforces versus GPT-4o’s 23.6. On LiveCodeBench, it hits 40.5% compared to GPT-4o’s 33.4%. On DeepSeek’s numbers, for coding and math-heavy workloads it outperformed the most widely used commercial models of its time.
On general knowledge (MMLU) and science reasoning (GPQA Diamond), it is competitive but not dramatically ahead — Claude 3.5 Sonnet leads it on GPQA (65.0% vs 59.1%), and MMLU scores sit within about three points across the models compared.
Hardware reality — this is NOT a local model
Let us be direct. DeepSeek V3 has 671 billion parameters. Even with MoE reducing active parameters to 37B per token, the full model weights must still be loaded into memory.
| Setup | VRAM Required | Feasibility |
|---|---|---|
| Full FP16 | ~1.3 TB | Server cluster only (16x A100 80GB minimum) |
| Q4 quantized | ~404 GB (Ollama download) | High-end multi-GPU server |
| Q2/Q3 aggressive quant | ~200 GB | Possible but significant quality loss |
| Cloud API | N/A | Most practical option for almost everyone |
If you have heard of running models locally via Ollama, DeepSeek V3 does have community-contributed quantizations, but you would need hardware that most businesses simply do not have:
# Available in Ollama, but the default Q4 download is 404 GB
ollama pull deepseek-v3
# For practical local coding work, use the smaller DeepSeek Coder V2 instead
ollama pull deepseek-coder-v2:16bFor most teams, the realistic path is the DeepSeek API, which uses an OpenAI-compatible format and costs significantly less than GPT-4o per token.
Geopolitical considerations for European businesses
DeepSeek is a Chinese AI company. For European businesses operating under GDPR, this raises legitimate questions:
- Data sovereignty: Queries sent to DeepSeek’s API travel to Chinese infrastructure. If you process personal data or confidential business information, this may conflict with your GDPR compliance posture.
- Regulatory uncertainty: The EU-China data transfer landscape is less settled than EU-US arrangements. There is no adequacy decision for China under GDPR.
- Code under MIT, weights under DeepSeek’s licence: The code itself is freely available under MIT license, and the model weights are under DeepSeek’s Model Agreement. If you self-host on European infrastructure, you control where data goes — but self-hosting requires the massive hardware described above.
The pragmatic approach for European SMEs: use DeepSeek V3 via API for non-sensitive tasks (public data analysis, coding assistance, research synthesis), and keep sensitive workloads on local models or European-hosted alternatives. For a deeper look at local deployment economics, see our cloud vs local cost analysis.
When to use DeepSeek V3 (and when not to)
Strong use cases:
- Complex coding and competitive programming tasks
- Advanced mathematics, data science, and quantitative analysis
- Research synthesis and scientific reasoning
- Batch processing of non-sensitive analytical tasks via API
Think twice when:
- Processing personal data subject to GDPR
- You need guaranteed uptime independent of Chinese infrastructure
- Local deployment is a hard requirement (the model is simply too large)
- You need the strongest scientific reasoning (Claude 3.5 Sonnet leads on GPQA Diamond in DeepSeek’s own table)
Related reading
- DeepSeek R1: The Best Open-Source Reasoning Model You Can Run Locally
- AESIA: What Every Spanish Business Deploying AI Must Know in 2026
- AI Evaluations: How to Test Your RAG Pipeline Before Going Live
Conclusion
DeepSeek V3 is a genuinely impressive achievement. It proves that open-weight models from outside the US-UK axis can compete at the frontier, and its MoE architecture is a masterclass in efficiency. The math and coding benchmarks speak for themselves — 90.2% on MATH-500 and 51.6 percentile on Codeforces were the best of any model in DeepSeek’s comparison table at release.
But for European SMEs, it is not a drop-in replacement for local AI. The 671B parameter count puts it firmly in cloud-or-cluster territory, and the Chinese origin adds a data governance layer that businesses must evaluate honestly. The smartest approach is hybrid: use DeepSeek V3 via API where it excels and data sensitivity allows, while keeping more practical models like Qwen 2.5 72B or Llama 3.3 70B for local deployment.
Sources: DeepSeek V3 HuggingFace model card, DeepSeek technical report, DeepSeek official site.
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.