Updated 9 October 2026. This post is from April 2026 and uses llama3.1:8b. The method still holds; with 16 GB, today you would run it with:
- qwen3:8b (5.2 GB) to generate the answers, our pick for a Spanish-language assistant in the memory-and-task table.
- bge-m3 (1.2 GB) for embeddings if your documents mix languages.
By the end you will have an assistant that answers questions about your PDFs, contracts and manuals, citing the source document, without a single byte leaving your network. It runs on a 16 GB Mac mini M4 or any PC with 16 GB of RAM. The cheapest machine to try it on is the laptop you already have.
What RAG Is in Two Sentences
A language model knows what it saw during training; it does not know your vacation policy or your deployment checklist. RAG (Retrieval-Augmented Generation) adds a step first: it searches your documents for the relevant passages and passes them to the model as context, so the answer comes from your real documentation and not from the model’s memory.
graph LR
A[Documents<br/>PDF, DOCX, TXT] --> B[Chunk<br/>500 tokens]
B --> C[Embeddings<br/>nomic-embed-text]
C --> D[Vector DB<br/>ChromaDB]
E[Question] --> F[Question<br/>embedding]
F --> G[Search<br/>5 chunks]
D --> G
G --> H[Prompt<br/>context + question]
H --> I[Ollama<br/>llama3.1:8b]
I --> J[Answer<br/>with sources]
style A fill:#0B1628,color:#FAFAFA
style J fill:#F5A623,color:#0B1628What You Need
| Component | Purpose | Install |
|---|---|---|
| n8n | Orchestrates the workflow | docker run -d --name n8n -p 5678:5678 -v n8n_data:/home/node/.n8n --add-host=host.docker.internal:host-gateway docker.n8n.io/n8nio/n8n |
| Ollama | Local LLM and embeddings | curl -fsSL https://ollama.com/install.sh | sh (Linux) or brew install ollama (macOS) |
| ChromaDB | Vector database | pip install chromadb |
| llama3.1:8b | Generates the answers | ollama pull llama3.1:8b (4.9 GB) |
| nomic-embed-text | Turns text into vectors | ollama pull nomic-embed-text (274 MB) |
Hardware: 16 GB of RAM. Time: an afternoon. No account or API key needed. If you are unsure which machine to buy, see our edge AI hardware guide.
One detail that saves an hour: if n8n runs in Docker and Ollama on the host, the Ollama URL inside n8n is http://host.docker.internal:11434, and Ollama must start with OLLAMA_HOST=0.0.0.0. The examples use localhost for clarity.
Step 1: Ingest, Chunk and Embed the Documents
Create an n8n workflow with a Schedule Trigger every 15 minutes followed by a Read/Write Files from Disk node that reads the documents folder:
{
"nodes": [
{
"name": "Every 15 minutes",
"type": "n8n-nodes-base.scheduleTrigger",
"parameters": { "rule": { "interval": [{ "field": "minutes", "minutesInterval": 15 }] } }
},
{
"name": "Read documents",
"type": "n8n-nodes-base.readWriteFile",
"parameters": { "operation": "read", "fileSelector": "/data/company-docs/*.pdf" }
}
]
}Add an Extract from File node (PDF operation) to get the text and a Code node that splits it into chunks of about 500 tokens with a 50-token overlap. The overlap keeps a sentence that falls on the cut from being lost.
For each chunk, an HTTP Request node calls Ollama’s embeddings endpoint:
{
"url": "http://localhost:11434/api/embed",
"method": "POST",
"body": { "model": "nomic-embed-text", "input": "{{ $json.chunk_text }}" }
}The response carries embeddings: [[...]], a vector of 768 numbers. Store it in ChromaDB with the text and metadata (file, page, chunk index):
import chromadb
client = chromadb.PersistentClient(path="/data/chromadb")
col = client.get_or_create_collection("company_docs")
col.add(
ids=[f"{filename}-{i}"],
embeddings=[vector],
documents=[text],
metadatas=[{"file": filename, "page": page, "chunk": i}],
)Check it works before moving on:
curl -s localhost:11434/api/embed -d '{"model":"nomic-embed-text","input":"test"}' | head -c 120It should return {"model":"nomic-embed-text","embeddings":[[0.0123,.... If it returns 404, your Ollama is older than 0.3: update it or use the old /api/embeddings endpoint with the prompt field.
Step 2: Query, Search and Answer
When someone asks, the flow repeats three moves: embed the question, find the five most similar chunks and pass them to the model.
{
"url": "http://localhost:11434/api/embed",
"method": "POST",
"body": { "model": "nomic-embed-text", "input": "What is our vacation policy?" }
}results = col.query(
query_embeddings=[question_embedding],
n_results=5,
include=["documents", "metadatas", "distances"],
){
"url": "http://localhost:11434/api/chat",
"method": "POST",
"body": {
"model": "llama3.1:8b",
"stream": false,
"messages": [
{ "role": "system", "content": "Answer using ONLY the provided context. If the context does not contain the answer, say so. Cite the source document." },
{ "role": "user", "content": "Context:\n{{ $json.retrieved_chunks }}\n\nQuestion: What is our vacation policy?" }
]
}
}The instruction “if the context does not contain the answer, say so” is the most important line in the system: it turns a hallucination into “I don’t know, ask HR”.
Step 3: Try It with a Real Question
An employee types: “What is our vacation policy?”
- nomic-embed-text embeds the question.
- ChromaDB returns five chunks from
HR-Policy-2026.pdf. - n8n builds the prompt with those chunks.
- llama3.1:8b answers: “According to the HR Policy 2026 (section 4.2), you have 23 working days of paid vacation per year. Requests go in 15 days in advance through the HR portal. Unused days can be carried over to the first quarter of the following year.”
On our machine, with five chunks (5,199 tokens of context), the full answer took about 6 seconds. Cost per query: zero.
What We Measured
The rows for our workstation (NVIDIA GB10, 128 GB unified memory, shared with other services) were measured on 2026-09-09 with curl against the Ollama API.
| Hardware | Model | Task | Result |
|---|---|---|---|
| GB10 | nomic-embed-text | Embed one question | 18 to 28 ms warm; 3.6 s on the first call (model load) |
| GB10 | llama3.1:8b | Read 5,199 tokens of context | 2.5 s |
| GB10 | llama3.1:8b | Generate the answer (8K window) | 26.6 tokens/s; 37.8 tokens/s with the default window (2026-09-08) |
| Mac mini M4, 24 GB | Qwen 2.5 7B | Generation | ~35 tokens/s (Compute Market) |
A larger context window costs speed: if your chunks are long, measure with the real num_ctx, not the default.
What It Costs Against the Cloud
A typical RAG query sends about 3,000 tokens (question plus five chunks) and gets 300 back. With OpenAI’s prices (GPT-5 mini $0.25/$2 and GPT-5 $1.25/$10 per million input/output tokens) and Anthropic’s (Claude Sonnet 5, $2/$10), checked on 2026-10-09:
| Option | Cost per query | 1,000 queries/day, 22 days |
|---|---|---|
| Local (Ollama + ChromaDB on a Mac mini) | 0 | ~EUR 2.40 of electricity (assumption: 30 W average around the clock, EUR 0.15/kWh) |
| GPT-5 mini | ~$0.0014 | ~$30 |
| GPT-5 | ~$0.007 | ~$150 |
| Claude Sonnet 5 | ~$0.009 | ~$200 |
For the cloud, add vector-database hosting and a GDPR processor agreement. Honestly: at a hundred queries a day, GPT-5 mini is hard to beat on cost; there you choose local because HR documents and contracts should not leave the building. The full calculation is in cloud vs local AI: find your break-even.
Three Settings That Change Quality
- Chunk size. 500 tokens works for most documents. Drop to 300 for dense technical manuals and go up to 800 for conversational text.
- Overlap. 10% of the chunk size prevents losing information at the cuts.
- Number of chunks retrieved. Start with 5. Go up to 8 or 10 for questions that span several sections.
Limits worth knowing: an 8B model summarises and locates well, but reasons worse than a frontier model over long contracts; scanned PDFs need OCR before ingestion; and tables inside PDFs often come out scrambled. To choose a model, the comparison of local models covers Llama, Mistral and Qwen.
Where It Fits
The real value appears when you connect the flow to the tools people already use. n8n can trigger the query from a message in a Slack #ask-hr channel, an intranet form, an email to a set address, or a weekly digest of the ten most repeated questions. More patterns in the n8n AI automation tutorial.
Next steps
- Apply the same pattern to code: Automate code reviews with n8n and Ollama.
- Give the assistant tools: n8n + MCP to connect agents to your systems.
- Reference docs: Ollama API, AI in n8n and ChromaDB.
Work with us
We build this pipeline for Spanish SMEs with their real documents and measure answer quality before we hand it over. If you want to do it with us, book a 15-minute call or see how we work in consulting.