By the end of this post you will know what local AI is, which pieces you need to try it this afternoon, and how to decide whether it beats a cloud API for you. You can start with the laptop you already have. With 8 GB of memory a small model already runs.
What local AI is
Local AI means the model runs on hardware you control. That can be your laptop, a mini PC in the office or a server in your rack. The prompt and the answer never leave that machine.
One distinction saves a lot of confusion: we are almost always talking about inference, not training. Inference is using a model that is already trained to get answers. Training a model from scratch takes data centres and is not a project for a small business. Adapting an existing model to your data (fine-tuning) is possible locally, but it is a second step and most people do not need it.
If you come from hospitality, think of it this way. A cloud API is eating out: you pay per plate and the kitchen is not yours. Local AI is your own kitchen: you buy the stove once and decide what comes through the door.
The three pieces you need
- A model. Open-weight models (Llama, Qwen, Gemma, Mistral, DeepSeek) are free to download. They usually come as GGUF files, and the file type decides how much memory they take. We explain it in GGUF quantization: pick the right file.
- A runtime, the program that loads the model and answers. Ollama is the simplest: free, no account, with a local API. llama.cpp is the engine underneath and gives you more control.
- Enough memory. The model file has to fit in RAM or VRAM, plus room for the conversation context. If it does not fit, it does not run slowly; it barely runs at all.
How much memory depends on model size and quantization. These figures come from the llama.cpp quantize README and bartowski’s table, collected in our GGUF guide (September 2026):
| Model | Q4_K_M file (4-bit) | F16 file (original) |
|---|---|---|
| Llama 3.1 8B | 4.58 GiB | 14.96 GiB |
| Llama 3.3 70B | 42.5 GB | ~141 GB |
The rule of thumb that follows: at 4 bits, a little over half a GB per billion parameters, plus context. Before you download anything, work out your case in the VRAM explorer.
Try it in ten minutes
On Linux, Ollama installs with one line. On macOS and Windows, download it from ollama.com/download. The full step-by-step guide, with the usual problems: how to install Ollama and run your first local LLM.
curl -fsSL https://ollama.com/install.sh | sh
ollama run llama3.1:8b "Explain local AI in three sentences."
ollama ps # shows how much memory the loaded model usesIf your machine has 8 GB, swap llama3.1:8b for llama3.2:3b. If the download fails, check you have about 5 GB of free disk.
What we measured on our machine
These figures come from our own workstation, not a vendor sheet. It is far more machine than you need to start; it is here so you have a reference point.
| Model (Ollama) | Memory loaded | Generation | Machine and date |
|---|---|---|---|
| llama3.1:8b | 9.2 GB | 41.3 tok/s | DGX Spark GB10, 128 GB, 2026-10-04 |
| qwen2.5-coder:7b | 6.6 GB | 41.0 tok/s | DGX Spark GB10, 128 GB, 2026-10-04 |
| gemma4:26b | 17 GB | 60.7 tok/s | DGX Spark GB10, 128 GB, 2026-10-04 |
| faster-whisper small | n/a | 30 min of audio in about 7.5 min | DGX Spark GB10, 128 GB, 2026-10-04 |
Look at the third row. The 26B model generated faster than the 7B and 8B ones. Parameter count alone does not predict speed, which is why you measure. To measure it on your machine, use --verbose as the Ollama install guide explains.
What it is good at today, and what not
It works well on bounded, repetitive tasks:
- Summarising, classifying and extracting data from documents.
- Drafting emails, product sheets or reports.
- Answering questions over your own documents with RAG (retrieval plus generation).
- Completing and reviewing code with a 7B model.
- Transcribing meetings with Whisper without uploading the audio anywhere.
When not to do this, or not only locally: hard multi-step reasoning, long-running agents and very specialised knowledge. The best cloud models are still ahead there. Locally, maintenance is also yours: updates, backups and security. The pattern that works best is hybrid: routine and private work on your machine, and a spend-capped cloud key for the exceptions.
Privacy and GDPR, in one paragraph
If the prompt and the answer stay inside your perimeter, there is no outside provider processing that data and no international transfer to justify. That fits the data protection by design principle in GDPR Article 25. It does not make you compliant on its own: you still need a legal basis, a record of processing and security measures. We cover it in GDPR Article 25: local AI inference is privacy by design and GDPR and AI in 2026: local deployment is the clean answer.
Local versus a cloud API
| Criterion | Local AI | Cloud API |
|---|---|---|
| Where the data lives | On your machine | On the provider’s servers |
| Cost shape | Hardware once, plus electricity and your time | Variable, per token used |
| Latency | No network round trip; depends on your hardware | Depends on the network and provider load |
| Maintenance | Yours: updates, monitoring, backups | The provider’s |
| Quality ceiling | The best open model that fits your memory | The most capable models on the market |
| Getting started | Install Ollama and download a model | Account, card and API key |
When it pays off
The cloud charges per token and your own machine is a fixed cost. So the answer almost always comes down to volume.
- It pays off when several people use it daily, when volume is steady and when data must not leave the building.
- It does not pay off on cost if you send a few short requests a day and a small cloud model handles the task. Then choose local only if privacy or predictability matter more.
- Electricity counts. Worked example: a 30 W machine left on all month uses 30 W × 24 h × 30 days = 21.6 kWh. Multiply by your own tariff.
To put in your own numbers, use the script in cloud vs local AI: work out your break-even. For a full business case, continue with the local AI ROI framework. The hardware page includes a break-even calculator.
Verdict. Buy if you have daily volume or sensitive data. Wait if you do not yet know how many tokens you use: measure for two weeks. Skip if you ask ten questions a week.
Where to start, by profile
- Curious. Install Ollama on your laptop with the block above and ask it questions from your work. When you want to understand the pieces, the playground has short in-browser exercises on the context window, text similarity and whether a model fits in memory.
- Developer. Start at the developer hub, pick your file with the GGUF guide and measure speed with
ollama run <model> --verbose. - Company. See what each budget buys in the local AI hardware catalogue. Then work out the savings with the ROI framework.
Model reviews to read next
- Qwen 2.5 Coder 7B on your own machine in 20 minutes: coding on a laptop.
- Llama 3.3 70B on your own hardware: what it takes to run it.
- Qwen2.5-72B-Instruct locally against a faster 35B.
- Google Gemma 4: the open model family review.
- DeepSeek R1: open-source reasoning you can run locally.
Work with us
This for your company? We run the numbers on your real volumes before recommending any hardware, and we tell you when the cloud is enough. Let’s talk for 15 minutes or see how our consulting works.