View all articles
tutorialn8nragollama

Local RAG over Company Docs with n8n and Ollama

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: ribbons of amber and white light weaving in a gentle curve. Models
Illustration generated with AI on our own machine.
This article is also available in Spanish:RAG local con n8n y Ollama sobre tus documentos

Updated 9 October 2026. This post is from April 2026 and uses llama3.1:8b. The method still holds; with 16 GB, today you would run it with:

  • qwen3:8b (5.2 GB) to generate the answers, our pick for a Spanish-language assistant in the memory-and-task table.
  • bge-m3 (1.2 GB) for embeddings if your documents mix languages.

By the end you will have an assistant that answers questions about your PDFs, contracts and manuals, citing the source document, without a single byte leaving your network. It runs on a 16 GB Mac mini M4 or any PC with 16 GB of RAM. The cheapest machine to try it on is the laptop you already have.

What RAG Is in Two Sentences

A language model knows what it saw during training; it does not know your vacation policy or your deployment checklist. RAG (Retrieval-Augmented Generation) adds a step first: it searches your documents for the relevant passages and passes them to the model as context, so the answer comes from your real documentation and not from the model’s memory.

graph LR
    A[Documents<br/>PDF, DOCX, TXT] --> B[Chunk<br/>500 tokens]
    B --> C[Embeddings<br/>nomic-embed-text]
    C --> D[Vector DB<br/>ChromaDB]
    E[Question] --> F[Question<br/>embedding]
    F --> G[Search<br/>5 chunks]
    D --> G
    G --> H[Prompt<br/>context + question]
    H --> I[Ollama<br/>llama3.1:8b]
    I --> J[Answer<br/>with sources]
    style A fill:#0B1628,color:#FAFAFA
    style J fill:#F5A623,color:#0B1628
Diagram

What You Need

ComponentPurposeInstall
n8nOrchestrates the workflowdocker run -d --name n8n -p 5678:5678 -v n8n_data:/home/node/.n8n --add-host=host.docker.internal:host-gateway docker.n8n.io/n8nio/n8n
OllamaLocal LLM and embeddingscurl -fsSL https://ollama.com/install.sh | sh (Linux) or brew install ollama (macOS)
ChromaDBVector databasepip install chromadb
llama3.1:8bGenerates the answersollama pull llama3.1:8b (4.9 GB)
nomic-embed-textTurns text into vectorsollama pull nomic-embed-text (274 MB)

Hardware: 16 GB of RAM. Time: an afternoon. No account or API key needed. If you are unsure which machine to buy, see our edge AI hardware guide.

One detail that saves an hour: if n8n runs in Docker and Ollama on the host, the Ollama URL inside n8n is http://host.docker.internal:11434, and Ollama must start with OLLAMA_HOST=0.0.0.0. The examples use localhost for clarity.

Step 1: Ingest, Chunk and Embed the Documents

Create an n8n workflow with a Schedule Trigger every 15 minutes followed by a Read/Write Files from Disk node that reads the documents folder:

JSON
{
  "nodes": [
    {
      "name": "Every 15 minutes",
      "type": "n8n-nodes-base.scheduleTrigger",
      "parameters": { "rule": { "interval": [{ "field": "minutes", "minutesInterval": 15 }] } }
    },
    {
      "name": "Read documents",
      "type": "n8n-nodes-base.readWriteFile",
      "parameters": { "operation": "read", "fileSelector": "/data/company-docs/*.pdf" }
    }
  ]
}

Add an Extract from File node (PDF operation) to get the text and a Code node that splits it into chunks of about 500 tokens with a 50-token overlap. The overlap keeps a sentence that falls on the cut from being lost.

For each chunk, an HTTP Request node calls Ollama’s embeddings endpoint:

JSON
{
  "url": "http://localhost:11434/api/embed",
  "method": "POST",
  "body": { "model": "nomic-embed-text", "input": "{{ $json.chunk_text }}" }
}

The response carries embeddings: [[...]], a vector of 768 numbers. Store it in ChromaDB with the text and metadata (file, page, chunk index):

Python
import chromadb
client = chromadb.PersistentClient(path="/data/chromadb")
col = client.get_or_create_collection("company_docs")
col.add(
    ids=[f"{filename}-{i}"],
    embeddings=[vector],
    documents=[text],
    metadatas=[{"file": filename, "page": page, "chunk": i}],
)

Check it works before moving on:

bash
curl -s localhost:11434/api/embed -d '{"model":"nomic-embed-text","input":"test"}' | head -c 120

It should return {"model":"nomic-embed-text","embeddings":[[0.0123,.... If it returns 404, your Ollama is older than 0.3: update it or use the old /api/embeddings endpoint with the prompt field.

Step 2: Query, Search and Answer

When someone asks, the flow repeats three moves: embed the question, find the five most similar chunks and pass them to the model.

JSON
{
  "url": "http://localhost:11434/api/embed",
  "method": "POST",
  "body": { "model": "nomic-embed-text", "input": "What is our vacation policy?" }
}
Python
results = col.query(
    query_embeddings=[question_embedding],
    n_results=5,
    include=["documents", "metadatas", "distances"],
)
JSON
{
  "url": "http://localhost:11434/api/chat",
  "method": "POST",
  "body": {
    "model": "llama3.1:8b",
    "stream": false,
    "messages": [
      { "role": "system", "content": "Answer using ONLY the provided context. If the context does not contain the answer, say so. Cite the source document." },
      { "role": "user", "content": "Context:\n{{ $json.retrieved_chunks }}\n\nQuestion: What is our vacation policy?" }
    ]
  }
}

The instruction “if the context does not contain the answer, say so” is the most important line in the system: it turns a hallucination into “I don’t know, ask HR”.

Step 3: Try It with a Real Question

An employee types: “What is our vacation policy?”

  1. nomic-embed-text embeds the question.
  2. ChromaDB returns five chunks from HR-Policy-2026.pdf.
  3. n8n builds the prompt with those chunks.
  4. llama3.1:8b answers: “According to the HR Policy 2026 (section 4.2), you have 23 working days of paid vacation per year. Requests go in 15 days in advance through the HR portal. Unused days can be carried over to the first quarter of the following year.”

On our machine, with five chunks (5,199 tokens of context), the full answer took about 6 seconds. Cost per query: zero.

What We Measured

The rows for our workstation (NVIDIA GB10, 128 GB unified memory, shared with other services) were measured on 2026-09-09 with curl against the Ollama API.

HardwareModelTaskResult
GB10nomic-embed-textEmbed one question18 to 28 ms warm; 3.6 s on the first call (model load)
GB10llama3.1:8bRead 5,199 tokens of context2.5 s
GB10llama3.1:8bGenerate the answer (8K window)26.6 tokens/s; 37.8 tokens/s with the default window (2026-09-08)
Mac mini M4, 24 GBQwen 2.5 7BGeneration~35 tokens/s (Compute Market)

A larger context window costs speed: if your chunks are long, measure with the real num_ctx, not the default.

What It Costs Against the Cloud

A typical RAG query sends about 3,000 tokens (question plus five chunks) and gets 300 back. With OpenAI’s prices (GPT-5 mini $0.25/$2 and GPT-5 $1.25/$10 per million input/output tokens) and Anthropic’s (Claude Sonnet 5, $2/$10), checked on 2026-10-09:

OptionCost per query1,000 queries/day, 22 days
Local (Ollama + ChromaDB on a Mac mini)0~EUR 2.40 of electricity (assumption: 30 W average around the clock, EUR 0.15/kWh)
GPT-5 mini~$0.0014~$30
GPT-5~$0.007~$150
Claude Sonnet 5~$0.009~$200

For the cloud, add vector-database hosting and a GDPR processor agreement. Honestly: at a hundred queries a day, GPT-5 mini is hard to beat on cost; there you choose local because HR documents and contracts should not leave the building. The full calculation is in cloud vs local AI: find your break-even.

Three Settings That Change Quality

  1. Chunk size. 500 tokens works for most documents. Drop to 300 for dense technical manuals and go up to 800 for conversational text.
  2. Overlap. 10% of the chunk size prevents losing information at the cuts.
  3. Number of chunks retrieved. Start with 5. Go up to 8 or 10 for questions that span several sections.

Limits worth knowing: an 8B model summarises and locates well, but reasons worse than a frontier model over long contracts; scanned PDFs need OCR before ingestion; and tables inside PDFs often come out scrambled. To choose a model, the comparison of local models covers Llama, Mistral and Qwen.

Where It Fits

The real value appears when you connect the flow to the tools people already use. n8n can trigger the query from a message in a Slack #ask-hr channel, an intranet form, an email to a set address, or a weekly digest of the ten most repeated questions. More patterns in the n8n AI automation tutorial.

Next steps

Work with us

We build this pipeline for Spanish SMEs with their real documents and measure answer quality before we hand it over. If you want to do it with us, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates