By the end you will have a language model running on your own computer, with no account and no API key. You will be able to chat with it, measure its speed and call it from your own scripts. A laptop with 8 GB of RAM is enough.
What Ollama is
Ollama is an open-source program that downloads language models and runs them on your machine. Under the hood it uses the same GGUF files as llama.cpp, but it spares you the compiling and the configuration. Think of a waiter: you order a dish by name (llama3.1:8b) and the waiter deals with the kitchen. It also leaves a server listening on port 11434 of your own computer, so any program you write can use the model.
What you need
| Your memory (RAM or VRAM) | Model to start with | Download size |
|---|---|---|
| 8 GB | qwen2.5:3b or llama3.2:3b | 1.9 GB and 2.0 GB |
| 16 GB | llama3.1:8b | 4.9 GB |
| 32 GB or more | 14B to 26B models | 9 to 17 GB |
Download sizes come from ollama.com/library and our own ollama list (checked 2026-10-06). A loaded model takes more memory than its file, because it also holds the conversation. Before downloading anything, try your case in the VRAM calculator.
System requirements from the official documentation (checked 2026-10-06):
- Windows: Windows 10 22H2 or newer. With an NVIDIA card, drivers 551.61 or newer. About 4 GB for the program itself (Windows docs).
- macOS: Sonoma (14) or newer. On an M-series Mac it uses the GPU; on an Intel Mac it runs on the CPU only (macOS docs).
- Linux: any modern distribution. With NVIDIA you need CUDA and
nvidia-smimust see the card (Linux docs).
If you do not have the machine yet, see the hardware we recommend at each budget.
Step 1: Install Ollama
The commands come from the official README on GitHub (checked 2026-10-06).
Windows. Download OllamaSetup.exe and run it. It does not need administrator rights. If you prefer PowerShell:
irm https://ollama.com/install.ps1 | iexmacOS. Download the .dmg from ollama.com/download, open it and drag Ollama into Applications. On first launch it asks to create the ollama command for the terminal.
Linux. One command installs the program and sets it up as a systemd service:
curl -fsSL https://ollama.com/install.sh | shCheck that it answers:
ollama --versionExpected: a line like ollama version is 0.30.10. Our two machines run different versions today (0.30.10 and 0.34.0), and every command in this post works the same on both.
If it fails: command not found usually means your terminal does not see the new PATH yet. Open a new terminal. On Linux, sudo systemctl status ollama tells you whether the service is running.
Step 2: Run your first model
With 16 GB or more:
ollama run llama3.1:8bWith 8 GB:
ollama run qwen2.5:3bThe first time, it downloads the model and then opens a chat in the terminal. Type a question and press Enter. To leave, type /bye.
If the download stops, run the same command again: it resumes where it left off. The first answer takes longer because the model is loading into memory. On our machine, that first load of llama3.1:8b took 5.8 seconds and later ones 0.21 seconds.
Step 3: Measure speed with —verbose
ollama run llama3.1:8b --verboseAfter each answer, Ollama prints its timings. The line that matters is eval rate, in tokens per second:
load duration: 184.728152ms
prompt eval rate: 518.39 tokens/s
eval count: 8 token(s)
eval rate: 44.63 tokens/sThat output is real, from our machine, for an answer only 8 tokens long. Longer answers bring the figure down a little, as the table below shows. To compare several models, ask each one the same question with --verbose and compare their eval rate.
Step 4: The commands you will use every day
| Command | What it does |
|---|---|
ollama list | Lists downloaded models and their size. |
ollama ps | Shows which models are loaded in memory and where they run. |
ollama stop llama3.1:8b | Unloads the model from memory right now. |
ollama rm llama3.1:8b | Deletes the model from disk. |
ollama pull qwen2.5:3b | Downloads a model without opening the chat. |
ollama show llama3.1:8b | Shows parameters, context length and quantization. |
When you stop using a model, Ollama unloads it after 5 minutes (official FAQ). ollama stop does it at once.
Step 5: Call the model from your code
The Ollama server listens on localhost:11434. This is the shape of the chat API:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.1:8b",
"stream": false,
"messages": [
{"role": "user", "content": "Explain in three sentences what a local language model is."}
]
}'The response is JSON with the text in message.content. It also carries eval_count (tokens generated) and eval_duration (in nanoseconds). Divide the first by the second and multiply by 10⁹ to get tokens per second.
On Windows, PowerShell handles quotes differently. Use curl.exe with the JSON in a file, or run the example in WSL. For more ways to connect a local model to your tools, the developers page collects them.
By default the server only accepts connections from your own machine. Before opening it to your network, read our Ollama security checklist.
What we measured
| Machine | Model | Generation | Loaded memory | Date |
|---|---|---|---|---|
| NVIDIA DGX Spark (GB10, 128 GB unified), Ollama 0.30.10 | llama3.1:8b | 40.5 / 40.9 / 40.8 tok/s over three /api/chat calls | 9.2 GB, 100% GPU | 2026-10-06 |
| Same machine, published reference run | llama3.1:8b | 41.3 tok/s | 9.2 GB | 2026-10-04 |
The 9.2 GB includes a 32,768-token context window, which is how our server is configured. With Ollama’s default (4,096 tokens) the same file took 5.3 GB in our quantization test. A laptop will be slower than this machine, which is why you should measure yours.
Where models live and how to free disk
From the official FAQ:
- macOS:
~/.ollama/models - Linux:
/usr/share/ollama/.ollama/models - Windows:
C:\Users\%username%\.ollama\models
To free space, run ollama list and then ollama rm on whatever you do not use. To keep models on another drive, set the OLLAMA_MODELS environment variable and restart Ollama. Our own model folder holds 159 GB (measured 2026-10-06). Nobody planned that; it accumulates.
Common problems
- The model will not load and mentions memory. Pick a smaller model (from 8B down to 3B) or close programs that use a lot of RAM. A model that does not fit is not slow: it barely runs.
- The GPU is not used. Run
ollama pswith the model loaded and read thePROCESSORcolumn.100% GPUis what you want.100% CPUor a split such as48%/52% CPU/GPUmeans the model does not fit on the card or the driver is missing. Try a smaller model and check the driver (nvidia-smion Linux and Windows). - It is very slow on an Intel Mac. That is expected: Ollama uses only the CPU on those Macs. Stay with 3B models.
- Something fails and you cannot tell what. On Linux,
journalctl -e -u ollamashows the service log. On macOS it lives in~/.ollama/logs.
When not to bother
- You have 4 GB of RAM or less. The models that fit answer worse than any free cloud chat.
- You need top quality on long reasoning. An 8B model does well on summaries, drafts and classification. It makes more mistakes on maths, regulation and long chains of steps. For that, a large cloud model is the honest choice.
- You use it once a month and your data is not sensitive. Local wins when the data cannot leave your machine, when the volume is high, or when you want to learn how it works inside.
Next steps
- Compare cost and speed against the cloud: cloud vs local AI cost benchmarks.
- Pick the right file for your memory: GGUF quantization.
- Practise with guided exercises in the playground.
Work with us
If you want this running for your team, with the model and the machine measured on your real tasks, we do that with you. Let’s talk for 15 minutes or see how we work in consulting.