View all articles
ollamalocal-aillmtutorial

How to install Ollama and run your first local LLM

JG
Jacobo González Jaspe
|
Abstract illustration: a network of white nodes joined by amber lines, like a constellation. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:Cómo instalar Ollama y ejecutar tu primer LLM en local

By the end you will have a language model running on your own computer, with no account and no API key. You will be able to chat with it, measure its speed and call it from your own scripts. A laptop with 8 GB of RAM is enough.

What Ollama is

Ollama is an open-source program that downloads language models and runs them on your machine. Under the hood it uses the same GGUF files as llama.cpp, but it spares you the compiling and the configuration. Think of a waiter: you order a dish by name (llama3.1:8b) and the waiter deals with the kitchen. It also leaves a server listening on port 11434 of your own computer, so any program you write can use the model.

What you need

Your memory (RAM or VRAM)Model to start withDownload size
8 GBqwen2.5:3b or llama3.2:3b1.9 GB and 2.0 GB
16 GBllama3.1:8b4.9 GB
32 GB or more14B to 26B models9 to 17 GB

Download sizes come from ollama.com/library and our own ollama list (checked 2026-10-06). A loaded model takes more memory than its file, because it also holds the conversation. Before downloading anything, try your case in the VRAM calculator.

System requirements from the official documentation (checked 2026-10-06):

  • Windows: Windows 10 22H2 or newer. With an NVIDIA card, drivers 551.61 or newer. About 4 GB for the program itself (Windows docs).
  • macOS: Sonoma (14) or newer. On an M-series Mac it uses the GPU; on an Intel Mac it runs on the CPU only (macOS docs).
  • Linux: any modern distribution. With NVIDIA you need CUDA and nvidia-smi must see the card (Linux docs).

If you do not have the machine yet, see the hardware we recommend at each budget.

Step 1: Install Ollama

The commands come from the official README on GitHub (checked 2026-10-06).

Windows. Download OllamaSetup.exe and run it. It does not need administrator rights. If you prefer PowerShell:

PowerShell
irm https://ollama.com/install.ps1 | iex

macOS. Download the .dmg from ollama.com/download, open it and drag Ollama into Applications. On first launch it asks to create the ollama command for the terminal.

Linux. One command installs the program and sets it up as a systemd service:

bash
curl -fsSL https://ollama.com/install.sh | sh

Check that it answers:

bash
ollama --version

Expected: a line like ollama version is 0.30.10. Our two machines run different versions today (0.30.10 and 0.34.0), and every command in this post works the same on both.

If it fails: command not found usually means your terminal does not see the new PATH yet. Open a new terminal. On Linux, sudo systemctl status ollama tells you whether the service is running.

Step 2: Run your first model

With 16 GB or more:

bash
ollama run llama3.1:8b

With 8 GB:

bash
ollama run qwen2.5:3b

The first time, it downloads the model and then opens a chat in the terminal. Type a question and press Enter. To leave, type /bye.

If the download stops, run the same command again: it resumes where it left off. The first answer takes longer because the model is loading into memory. On our machine, that first load of llama3.1:8b took 5.8 seconds and later ones 0.21 seconds.

Step 3: Measure speed with —verbose

bash
ollama run llama3.1:8b --verbose

After each answer, Ollama prints its timings. The line that matters is eval rate, in tokens per second:

text
load duration:        184.728152ms
prompt eval rate:     518.39 tokens/s
eval count:           8 token(s)
eval rate:            44.63 tokens/s

That output is real, from our machine, for an answer only 8 tokens long. Longer answers bring the figure down a little, as the table below shows. To compare several models, ask each one the same question with --verbose and compare their eval rate.

Step 4: The commands you will use every day

CommandWhat it does
ollama listLists downloaded models and their size.
ollama psShows which models are loaded in memory and where they run.
ollama stop llama3.1:8bUnloads the model from memory right now.
ollama rm llama3.1:8bDeletes the model from disk.
ollama pull qwen2.5:3bDownloads a model without opening the chat.
ollama show llama3.1:8bShows parameters, context length and quantization.

When you stop using a model, Ollama unloads it after 5 minutes (official FAQ). ollama stop does it at once.

Step 5: Call the model from your code

The Ollama server listens on localhost:11434. This is the shape of the chat API:

bash
curl http://localhost:11434/api/chat -d '{
  "model": "llama3.1:8b",
  "stream": false,
  "messages": [
    {"role": "user", "content": "Explain in three sentences what a local language model is."}
  ]
}'

The response is JSON with the text in message.content. It also carries eval_count (tokens generated) and eval_duration (in nanoseconds). Divide the first by the second and multiply by 10⁹ to get tokens per second.

On Windows, PowerShell handles quotes differently. Use curl.exe with the JSON in a file, or run the example in WSL. For more ways to connect a local model to your tools, the developers page collects them.

By default the server only accepts connections from your own machine. Before opening it to your network, read our Ollama security checklist.

What we measured

MachineModelGenerationLoaded memoryDate
NVIDIA DGX Spark (GB10, 128 GB unified), Ollama 0.30.10llama3.1:8b40.5 / 40.9 / 40.8 tok/s over three /api/chat calls9.2 GB, 100% GPU2026-10-06
Same machine, published reference runllama3.1:8b41.3 tok/s9.2 GB2026-10-04

The 9.2 GB includes a 32,768-token context window, which is how our server is configured. With Ollama’s default (4,096 tokens) the same file took 5.3 GB in our quantization test. A laptop will be slower than this machine, which is why you should measure yours.

Where models live and how to free disk

From the official FAQ:

  • macOS: ~/.ollama/models
  • Linux: /usr/share/ollama/.ollama/models
  • Windows: C:\Users\%username%\.ollama\models

To free space, run ollama list and then ollama rm on whatever you do not use. To keep models on another drive, set the OLLAMA_MODELS environment variable and restart Ollama. Our own model folder holds 159 GB (measured 2026-10-06). Nobody planned that; it accumulates.

Common problems

  1. The model will not load and mentions memory. Pick a smaller model (from 8B down to 3B) or close programs that use a lot of RAM. A model that does not fit is not slow: it barely runs.
  2. The GPU is not used. Run ollama ps with the model loaded and read the PROCESSOR column. 100% GPU is what you want. 100% CPU or a split such as 48%/52% CPU/GPU means the model does not fit on the card or the driver is missing. Try a smaller model and check the driver (nvidia-smi on Linux and Windows).
  3. It is very slow on an Intel Mac. That is expected: Ollama uses only the CPU on those Macs. Stay with 3B models.
  4. Something fails and you cannot tell what. On Linux, journalctl -e -u ollama shows the service log. On macOS it lives in ~/.ollama/logs.

When not to bother

  • You have 4 GB of RAM or less. The models that fit answer worse than any free cloud chat.
  • You need top quality on long reasoning. An 8B model does well on summaries, drafts and classification. It makes more mistakes on maths, regulation and long chains of steps. For that, a large cloud model is the honest choice.
  • You use it once a month and your data is not sensitive. Local wins when the data cannot leave your machine, when the volume is high, or when you want to learn how it works inside.

Next steps

Work with us

If you want this running for your team, with the model and the machine measured on your real tasks, we do that with you. Let’s talk for 15 minutes or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates