View all articles
LLMOllamalocal AIHardwareComparison

Best local LLMs: which one to use for your hardware and task

JG
Jacobo González Jaspe
|
Abstract illustration: ribbons of amber and white light weaving in a gentle curve. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:Mejores LLM locales: cuál usar según tu hardware y tu tarea

By the end you will know which model to download for your task and your memory, and how to check it with your own numbers in ten minutes. You can start on any laptop with 8 GB of RAM. The rule underneath is simple: the smallest model that does the job well.

Last reviewed: 6 October 2026. Sizes taken from the Ollama library that day; speeds measured on our NVIDIA DGX Spark. We revise this table when a model on it changes.

What you need

  • A computer with 8 GB of memory or more. A GPU helps but is not required.
  • Ollama installed. It is free and needs no account.
  • Ten real prompts from your own work, saved in a text file.
  • Between 2 and 20 GB of free disk, depending on the column you pick.

Step 1: find out how much memory the model can use

The number that matters is the memory the model can reach. On a Mac or a laptop without a GPU it is total RAM, because it is shared. On a PC with a graphics card it is the card’s VRAM.

bash
# Linux with an NVIDIA GPU: total VRAM
nvidia-smi --query-gpu=memory.total --format=csv
# macOS: unified memory in bytes
sysctl hw.memsize

With shared memory, the operating system keeps a slice. Our VRAM explorer reserves 3 GB on machines up to 16 GB, 4 GB up to 32 GB and 6 GB above that. An 8 GB laptop leaves about 5 GB for the model.

Step 2: pick from the table

Each cell gives the tag for ollama pull and the download size its Ollama library page shows. Where you see a range, Ollama serves different files per platform.

Task8 GB16 GB24-32 GB64 GB or more
Chat and assistantqwen3:4b (2.5 GB)qwen3:8b (5.2 GB)gemma4:26b (16-19 GB)qwen3.6:35b (23-24 GB)
Codingqwen2.5-coder:3b (1.9 GB)qwen2.5-coder:7b (4.7 GB)qwen3-coder:30b (19 GB)qwen3.6:35b-coding (23-24 GB)
Documents and RAGbge-m3 (1.2 GB) + qwen3:4bbge-m3 + qwen3:8bbge-m3 + gemma4:26bbge-m3 + qwen3.6:35b
Reasoningqwen3:4b-thinking (2.5 GB)deepseek-r1:14b (9.0 GB)qwen3:30b-thinking (19 GB)qwen3:30b-thinking with long context
Vision (photos, scanned PDFs)qwen3-vl:2b (1.9 GB)qwen3-vl:8b (6.1 GB)qwen3-vl:30b (20 GB)qwen3-vl:30b; qwen2.5vl:72b (49 GB) for batch jobs only
Writing in Spanishqwen3:4bgemma4:e4b (6.6-9.5 GB)gemma4:26bqwen3.6:35b
Speech to text (outside Ollama)faster-whisper smallsmallmediumlarge-v3

How we decide that it fits: library size, plus the context cache, plus 1.5 GB of runtime. It is the same sum this site uses in src/lib/hardware-math.ts. The cache for a 3 to 8B model at 4,096 tokens is about 0.5 GB.

Some cells are tight. qwen3:4b needs about 4.5 GB on an 8 GB laptop, so close the browser. The 24-32 GB column needs 17 to 22 GB: comfortable on a 24 GB GPU or with 32 GB of shared memory, and at the limit on a 24 GB Mac.

Speech to text does not go through Ollama. Use faster-whisper: its README measures small on CPU at 1.5 GB of RAM in int8, and large-v2 on GPU at 2.9 GB in int8. large-v3 has the same parameter count as large-v2.

Why the 64 GB column repeats models

More memory does not mean a bigger model. Tags with a3b or a4b in them are mixture-of-experts models: they keep every parameter in memory but activate about 3 or 4 billion per token. That is why qwen3.6:35b generated faster than deepseek-r1:14b in our test.

With 64 GB or more, what you gain is long context and room for two models loaded at once, say one for chat and one for vision. A dense 70B model fits, but on our machine qwen2.5vl:72b ran at about 3 tokens per second. That suits overnight batches, not conversation.

Step 3: download and measure on your machine

Pull the tag from your cell and time it with the same call we use:

bash
ollama pull qwen3:8b
curl -s localhost:11434/api/generate -d '{
  "model": "qwen3:8b", "stream": false, "think": false,
  "prompt": "Write a short email moving a meeting from Thursday to Monday.",
  "options": {"num_predict": 200, "num_ctx": 4096}
}' | python3 -c "import json,sys; d=json.load(sys.stdin); print(round(d['eval_count']/d['eval_duration']*1e9,1),'tok/s')"
ollama ps   # the SIZE column is the memory it really uses

Expected output: one line with tokens per second and an ollama ps table. If PROCESSOR does not say 100% GPU on a GPU machine, the model did not fit whole. Move one cell left or lower num_ctx.

Below about 10 tokens per second, reading the answer starts to drag. To measure your own machine with --verbose, follow the Ollama install guide.

Step 4: test quality with your ten prompts

Speed takes a second to measure. Quality is yours to judge. Run your ten prompts through your memory’s cell and the one to its left, at temperature: 0, and put the answers side by side.

If you cannot tell them apart, keep the smaller model. You free memory and gain speed. For Spanish writing, check accents, agreement and whether English words slip in.

What we measured

NVIDIA DGX Spark (GB10, 128 GB unified memory), Ollama 0.30.10, 4,096-token context. The machine is shared with other services, so treat these figures as a floor.

ModelGeneration speedMemory loadedDate
qwen2.5-coder:7b41.0 tok/s6.6 GB2026-10-04
gemma4:26b60.7 tok/s17 GB2026-10-04
deepseek-r1:14b (thinking off)21.6 tok/s, one run9.5 GB2026-10-06
qwen2.5vl:7b (text prompt)42.3 tok/s, one runnot measured2026-10-06
qwen3.6:35b71.0 tok/s, one run23 GB2026-10-06
qwen2.5vl:72b2.7 to 3.2 tok/s49 GB at 8k context2026-09-09
faster-whisper small0.252 x real timenot measured2026-10-04
faster-whisper medium0.513 x real timenot measured2026-10-04

A factor of 0.252 means a 30-minute meeting is transcribed in about 7.5 minutes. A Spark is not your laptop. On a 16 GB machine without a dedicated GPU, expect much lower numbers, which is why Step 3 exists.

When a cloud API is the better call

  • Coding agents across a large repository. A 7 to 30B model completes functions and explains code. Planning a change across fifty files is still work for the large cloud models.
  • Occasional use with non-sensitive data. If you send twenty requests a month, paying per use is cheaper than any machine. Run the numbers with cloud vs local break-even.
  • Long reasoning where a mistake is expensive. A local 14B model reasons well on bounded problems. For delicate legal or financial analysis, compare against a large model before you trust it.

Local wins when the data must not leave your network, when volume is high and steady, or when you want a fixed cost. Electricity and your own time count as costs too.

Next steps

Work with us

If you want to see your own task running on real hardware before you buy anything, we measure the model and the machine with your prompts. Let’s talk for 15 minutes or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates