View all articles
transparencyopen-sourcestacklocal-ai

The VORLUX AI Stack: Every Tool We Use, Nothing Hidden

JG
Jacobo González Jaspe
|

Reviewed:

Abstract illustration: pale and navy geometric planes fitting together, with an amber glow at the joins. Deployment
Illustration generated with AI on our own machine.

When we tell clients their AI will run locally with no cloud dependency, the natural follow-up is: “Okay, but what exactly are you running?” Fair question. If we are asking you to trust us with your infrastructure, you deserve to see everything under the hood.

This post is our technology disclosure: every component, what it does, and why we chose it. The model list and counts below come off our own machines, with the date and the command that produced them.

The Core Components

Here’s what powers VORLUX AI, from inference to interface:

LayerTechnologyRoleWhy this one
InferenceOllamaServes most of our modelsSimple model management, GPU offload, OpenAI-style API
InferencevLLMServes Qwen2.5-7B-Instruct on the second SparkHigher throughput for concurrent requests
RoutingLiteLLMOne gateway in front of every modelSwap models without touching callers
APIFastAPI + PythonOrchestrator on port 8091Fast, typed, async-native
DatabaseSQLitePersistence for the orchestratorZero config, zero network
SearchQdrant, FAISS + BM25RAG retrievalVector + keyword hybrid search
Embeddingsbge-m3, nomic-embed-textTurn documents into vectorsRun locally in Ollama
Public siteAstrovorluxai.comStatic-first, fast
Automationn8nWorkflow automationVisual workflows, self-hosted
Schedulingsystemd timers + APSchedulerRecurring jobsSurvive reboots, logged by the OS
Hardware2x NVIDIA DGX SparkPrimary and assistant server128 GB unified memory each, one GB10 GPU

Every component runs on our hardware or on the client’s. Inference never leaves the machine.

How It All Fits Together

flowchart TB
    subgraph CLIENT["Client Layer"]
        SITE["Astro Site<br/>vorluxai.com"]
        N8N["n8n Workflows<br/>:5678"]
    end

    subgraph API_LAYER["API & Orchestration"]
        ORCH["FastAPI Orchestrator<br/>:8091"]
        GW["LiteLLM Gateway<br/>:4000"]
    end

    subgraph INFERENCE["Inference Layer"]
        OLLAMA["Ollama<br/>:11434"]
        VLLM["vLLM on Spark 2<br/>:8200"]
        RAG["Qdrant + FAISS/BM25<br/>RAG Search"]
    end

    subgraph DATA["Data Layer"]
        SQLITE[("SQLite<br/>Orchestrator DBs")]
    end

    N8N --> ORCH
    ORCH --> GW
    GW --> OLLAMA
    GW --> VLLM
    ORCH --> RAG
    ORCH --> SQLITE

    style CLIENT fill:#0B1628,color:#FAFAFA
    style INFERENCE fill:#059669,color:#fff
    style DATA fill:#F5A623,color:#0B1628
Diagram

The Models We Run

Not every task needs the same model. We keep a library on disk and route each request to the right one through LiteLLM. On 2026-10-09, ollama list on our main DGX Spark (spark-43d5) showed 20 models, including:

ModelDownload sizeWhat we use it for
qwen3.6:35b23 GBGeneral reasoning and Spanish writing
gemma4:26b17 GBSecond opinion, long-form drafts
qwen2.5vl:72b48 GBReading scanned documents and images
deepseek-r1:14b9.0 GBStep-by-step reasoning
llama3.1:8b4.9 GBFast, light tasks
qwen2.5-coder:7b4.7 GBCode generation and review
bge-m3, nomic-embed-text1.2 GB, 274 MBEmbeddings for search

The rest are our own fine-tunes (the apprendere and j4sgon-finance families) and small vision and avatar models. They are not all loaded at once. At the same moment, ollama ps showed three models in memory: llama3.1:8b, qwen2.5-coder:7b and bge-m3. Ollama loads a model on first request and unloads it after a timeout, so memory goes to what is in use.

A 72B model at 48 GB would not fit next to others on a 32 GB laptop. It fits here because each Spark has 128 GB of unified memory shared by CPU and GPU.

What Runs on a Schedule

The system does not only answer requests. On 2026-10-09, systemctl --user list-timers --all on the main Spark listed 68 timers. They cover:

  • Content: research, draft, review and publish steps, each with a human gate before anything goes live
  • Quality: test runs, link checks on the built site, knowledge-base updates
  • Monitoring: health checks and watchdogs that restart a service after repeated failed probes
  • Backups: daily git bundles of every repo and database snapshots

A watchdog restarts the orchestrator only after three failed health probes in a row, so a single slow response does not kill a working service. We cover the deployment side in our local deployment guide.

Why Open-Source Matters

Every component in our stack is either open-source or built by us in-house. This isn’t ideological: it’s practical.

  1. No licence fees: our clients don’t pay software licences for the stack. Hardware and our time are the costs.
  2. No vendor lock-in: if Ollama disappears tomorrow, we switch to llama.cpp or vLLM. Same models, different runtime.
  3. Auditability: regulated clients can inspect the code that touches their data. That supports the data-protection-by-design duty in GDPR Article 25, though the duty itself is about your whole process, not your licence.
  4. Mature projects: Ollama, vLLM, n8n and Astro are widely used, actively maintained projects, not experiments.

Compared to Cloud-Dependent Stacks

AspectVORLUX AI (local)Typical cloud stack
Data locationYour hardwareThe provider’s data centres
Running costElectricity and maintenancePer-token or per-seat fees
Internet required for inferenceNoYes
AI vendor as GDPR processorNoYes, needs an Art. 28 contract
Model switchingChange one LiteLLM routeDepends on the provider’s catalogue
Uptime dependencyYour power and hardwareTheir SLA
Audit trailFull local logsProvider-dependent

The cloud stack isn’t wrong for everyone. For businesses processing sensitive data under European regulation, local deployment removes the AI vendor from the data flow. We explored the cost side in our cost analysis.

What This Means for You

When we deploy AI for your business, you get this same architecture, sized to your hardware and workloads. A small office does not need two DGX Sparks: a mini PC or a Mac with enough memory for one 7-8B model covers many tasks. We size it by testing your workload, not by guessing.

See It in Action

We run live demos of this stack during our free assessment calls. No slides, no mockups: the actual system, running actual models, on your sample queries.

Book your free 15-minute assessment and see what local AI looks like in practice.


This is post 2 of our Launch Week series. Yesterday: Local AI Readiness Checklist. Tomorrow: Our Services and Pricing.

External references: Ollama | vLLM | LiteLLM | n8n | Astro | GDPR on EUR-Lex


Next steps

  • Check which of the models above fits your own hardware
  • Map one recurring task to a scheduled job
  • Compare the running cost of a local stack with your current cloud bills

Work with us

We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.

Diagram
Share: LinkedIn X
Veredicto semanal

Get new guides before anyone else

Subscribe and we tell you when new guides, templates and workflows go up. One email a week, no spam.

Already published: 69 guides and 25 templates. All free, no signup.

Bonus: the local-AI starter pack PDF when you subscribe
Once a week No spam Unsubscribe anytime

See what you get

The EU AI Act now applies: a checklist you can complete

Tell us what you want to run

Tell us what you want to run and on what budget. We will tell you which hardware you need, which model fits, and what to expect from it, before you spend anything.

First call free, 15 min Local-first: your data stays on your network Open tools and guides

69 free guides · 17 compliance templates