Raspberry Pi + Hailo · Board / edge
RPi 5 + AI HAT+ 2
Education, prototyping, small models
- 16 GBunified memory
- 7–8Bfits in memory (Q4, 8k)
- 3Bpractical (estimate)
- 15 Wmax draw
- ~€200indicative price
What can it run?
Pick a job or a model size. The bar compares what it needs with this machine's usable memory.
- Model weights
- Context cache
- Runtime + extras (Whisper, embeddings)
How we calculate this
Weights = parameters × bits per weight / 8 (Q4_K_M ≈ 4.85; Q8_0 ≈ 8.5). Cache = context tokens × the reference model's per-token size (layers × KV heads × head size × 2 × 2 bytes). Plus 1.5 GB of runtime. Usable memory = total minus what the system keeps (3–6 GB on shared memory, 0.5 GB on a GPU). Comfortable = fits in 85%. It is a ±20% estimate: measure it on your machine before you decide. How to pick the GGUF file
When does it beat the cloud?
Set your usage. We compare the purchase plus electricity with what you would pay a cloud API.
| Month | Own hardware | Cloud |
|---|
- Own hardware (purchase + electricity)
- Cloud API
Assumptions: €0.25/kWh (Spain average, editable), USD 1 = €0.92, 3 input tokens per output token, and the machine at maximum draw for every hour it is on (worst case). API prices verified 2026-09-09. Your time and maintenance are not included. Note: the cloud side is a frontier model and here you would run a smaller open one; this compares cost, not quality.
Learn with this machine
- Local AI Hardware Catalogue 2026: What to Buy at Each BudgetA buying catalogue for local AI by budget tier, with third-party benchmarks for every device and five questions that tell you which tier you actually need.
- Your First 3 AI Agents: A Local Deployment Guide for SMEs (2026)Deploy three AI agents on hardware you own with Ollama and MCP, in order of safety and impact, starting with a read-only digest on an EUR 80 board.
- Edge AI in Manufacturing: How Spanish Factories Deploy ItPredictive maintenance, visual quality control and operator support with AI on the factory floor: use cases, a starter hardware kit and how a Spanish SME starts.
- GGUF Quantization: Pick the Right File for Your MachineHow to read a GGUF file name, choose between Q4_K_M, Q5 and Q8 for the memory you have, and measure speed, memory and quality yourself with three commands.
Skip it if…
- you want 7–8B models to run comfortably: here they are tight or do not fit.
Specs and where the numbers come from
| CPU | 4-core Arm Cortex-A76 |
|---|---|
| GPU | VideoCore VII |
| NPU | 40 TOPS (Hailo) |
| Memory | 16 GB (unified memory); usable by the model ≈ 13 GB |
| Speed | 8 tok/s with 4B vendor or community estimate, not measured by us |
| Price | ~€200 indicative, checked 2026-09-19; check the live price in the shop |
We only call something "measured" when it ran on our machines. Editorial policy