View all articles
modelsopen-sourceaireview

DeepSeek V3: a 671B Open-Weight Giant for Code and Math

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: a crystal polyhedron refracting an amber beam into white rays. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:DeepSeek V3: un gigante open-weight de 671B para código

Updated 4 October 2026. This post is from April 2026 and names GPT-4o, Claude 3.5, Llama 3.1, Qwen 2.5, Qwen2.5 as current. The method and the advice still hold; today you would run it with:

  • DeepSeek V4: Use this if you need even higher performance for complex reasoning, though memory requirements remain significant.
  • Llama 4 Scout / Maverick: These models offer a more practical alternative for local deployment on smaller, more accessible hardware.
  • Gemma 4: A smaller, efficient option for developers needing high-quality performance on constrained local infrastructure.

DeepSeek V3 changed the conversation about what open-weight models can achieve. Built by a Chinese AI lab, it was trained on 14.8 trillion tokens using 2.788 million H800 GPU hours. In DeepSeek’s own evaluation, this model matches or beats GPT-4o on several hard benchmarks, specifically coding and mathematics.

At 671 billion parameters, it is far too large for any local workstation. We should be upfront about what this model delivers, what it costs to run, and what European businesses should consider before adopting it.

Open source AI model comparison
flowchart LR
    INPUT["Input Token"] --> ROUTER["Router Network"]
    ROUTER --> E1["Expert 1"]
    ROUTER --> E2["Expert 2"]
    ROUTER -.->|inactive| E3["Expert 3"]
    ROUTER -.->|inactive| E4["..."]
    ROUTER -.->|inactive| E128["Expert 128"]

    E1 --> COMBINE["Combine Outputs"]
    E2 --> COMBINE

    subgraph TOTAL ["671B Total Parameters"]
        ROUTER
        E1
        E2
        E3
        E4
        E128
    end

    subgraph ACTIVE ["37B Active per Token"]
        E1
        E2
    end

    COMBINE --> MTP["Multi-Token\nPrediction"]
    MTP --> OUTPUT["Output Tokens"]

    style INPUT fill:#DBEAFE,stroke:#2563EB,color:#000
    style ROUTER fill:#FEF3C7,stroke:#F5A623,color:#000
    style E1 fill:#D1FAE5,stroke:#059669,color:#000
    style E2 fill:#D1FAE5,stroke:#059669,color:#000
    style E3 fill:#FECACA,stroke:#B91C1C,color:#000
    style E4 fill:#FECACA,stroke:#B91C1C,color:#000
    style E128 fill:#FECACA,stroke:#B91C1C,color:#000
    style COMBINE fill:#FEF3C7,stroke:#F5A623,color:#000
    style MTP fill:#DBEAFE,stroke:#2563EB,color:#000
    style OUTPUT fill:#D1FAE5,stroke:#059669,color:#000
Diagram

How Mixture-of-Experts works (in plain terms)

Most language models are “dense” — every parameter participates in every computation. DeepSeek V3 uses a different approach called Mixture-of-Experts (MoE). Think of it like a company with 671 billion employees, but for any given task, only 37 billion of them actually work on it. A routing network decides which “expert” sub-networks handle each token.

The result: you get the knowledge depth of a 671B model with the inference speed closer to a 37B dense model. DeepSeek also introduced Multi-Token Prediction (MTP) for faster generation and FP8 mixed-precision training to keep costs down. The model supports a 128K context window, which is generous for long-document tasks.

Real benchmarks — honest numbers

These numbers come directly from the official HuggingFace model card and the DeepSeek technical report.

BenchmarkDeepSeek V3GPT-4o (0513)Claude 3.5 Sonnet (1022)Llama 3.1 405BQwen2.5 72B
MMLU (Chat)88.5%87.2%88.3%88.6%85.3%
MMLU (Base, 5-shot)87.1%——84.4%—
MMLU-Pro75.9%72.6%78.0%73.3%71.6%
GPQA Diamond59.1%49.9%65.0%51.1%49.0%
MATH-50090.2%74.6%78.3%73.8%80.0%
AIME 202439.2%9.3%16.0%23.3%23.3%
HumanEval-Mul (code)82.6%80.5%81.7%77.2%77.3%
LiveCodeBench (Pass@1-COT)40.5%33.4%36.3%28.4%31.1%
Codeforces (percentile)51.623.620.325.324.8
Arena Hard85.580.485.269.381.2

Sources: DeepSeek V3 on HuggingFace, DeepSeek technical report (arXiv). All columns are DeepSeek’s own measurements, published with the model in December 2024; newer versions of the other models score differently.

The standout results: DeepSeek V3 scores 90.2% on MATH-500 (well ahead of GPT-4o’s 74.6% in the same table), and 51.6 percentile on Codeforces versus GPT-4o’s 23.6. On LiveCodeBench, it hits 40.5% compared to GPT-4o’s 33.4%. On DeepSeek’s numbers, for coding and math-heavy workloads it outperformed the most widely used commercial models of its time.

On general knowledge (MMLU) and science reasoning (GPQA Diamond), it is competitive but not dramatically ahead — Claude 3.5 Sonnet leads it on GPQA (65.0% vs 59.1%), and MMLU scores sit within about three points across the models compared.

Hardware reality — this is NOT a local model

Let us be direct. DeepSeek V3 has 671 billion parameters. Even with MoE reducing active parameters to 37B per token, the full model weights must still be loaded into memory.

SetupVRAM RequiredFeasibility
Full FP16~1.3 TBServer cluster only (16x A100 80GB minimum)
Q4 quantized~404 GB (Ollama download)High-end multi-GPU server
Q2/Q3 aggressive quant~200 GBPossible but significant quality loss
Cloud APIN/AMost practical option for almost everyone

If you have heard of running models locally via Ollama, DeepSeek V3 does have community-contributed quantizations, but you would need hardware that most businesses simply do not have:

bash
# Available in Ollama, but the default Q4 download is 404 GB
ollama pull deepseek-v3

# For practical local coding work, use the smaller DeepSeek Coder V2 instead
ollama pull deepseek-coder-v2:16b

For most teams, the realistic path is the DeepSeek API, which uses an OpenAI-compatible format and costs significantly less than GPT-4o per token.

Geopolitical considerations for European businesses

DeepSeek is a Chinese AI company. For European businesses operating under GDPR, this raises legitimate questions:

  • Data sovereignty: Queries sent to DeepSeek’s API travel to Chinese infrastructure. If you process personal data or confidential business information, this may conflict with your GDPR compliance posture.
  • Regulatory uncertainty: The EU-China data transfer landscape is less settled than EU-US arrangements. There is no adequacy decision for China under GDPR.
  • Code under MIT, weights under DeepSeek’s licence: The code itself is freely available under MIT license, and the model weights are under DeepSeek’s Model Agreement. If you self-host on European infrastructure, you control where data goes — but self-hosting requires the massive hardware described above.

The pragmatic approach for European SMEs: use DeepSeek V3 via API for non-sensitive tasks (public data analysis, coding assistance, research synthesis), and keep sensitive workloads on local models or European-hosted alternatives. For a deeper look at local deployment economics, see our cloud vs local cost analysis.

When to use DeepSeek V3 (and when not to)

Strong use cases:

  • Complex coding and competitive programming tasks
  • Advanced mathematics, data science, and quantitative analysis
  • Research synthesis and scientific reasoning
  • Batch processing of non-sensitive analytical tasks via API

Think twice when:

  • Processing personal data subject to GDPR
  • You need guaranteed uptime independent of Chinese infrastructure
  • Local deployment is a hard requirement (the model is simply too large)
  • You need the strongest scientific reasoning (Claude 3.5 Sonnet leads on GPQA Diamond in DeepSeek’s own table)

Conclusion

DeepSeek V3 is a genuinely impressive achievement. It proves that open-weight models from outside the US-UK axis can compete at the frontier, and its MoE architecture is a masterclass in efficiency. The math and coding benchmarks speak for themselves — 90.2% on MATH-500 and 51.6 percentile on Codeforces were the best of any model in DeepSeek’s comparison table at release.

But for European SMEs, it is not a drop-in replacement for local AI. The 671B parameter count puts it firmly in cloud-or-cluster territory, and the Chinese origin adds a data governance layer that businesses must evaluate honestly. The smartest approach is hybrid: use DeepSeek V3 via API where it excels and data sensitivity allows, while keeping more practical models like Qwen 2.5 72B or Llama 3.3 70B for local deployment.


Sources: DeepSeek V3 HuggingFace model card, DeepSeek technical report, DeepSeek official site.


Work with us

We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates