When we tell clients their AI will run locally with no cloud dependency, the natural follow-up is: “Okay, but what exactly are you running?” Fair question. If we are asking you to trust us with your infrastructure, you deserve to see everything under the hood.
This post is our technology disclosure: every component, what it does, and why we chose it. The model list and counts below come off our own machines, with the date and the command that produced them.
The Core Components
Here’s what powers VORLUX AI, from inference to interface:
| Layer | Technology | Role | Why this one |
|---|---|---|---|
| Inference | Ollama | Serves most of our models | Simple model management, GPU offload, OpenAI-style API |
| Inference | vLLM | Serves Qwen2.5-7B-Instruct on the second Spark | Higher throughput for concurrent requests |
| Routing | LiteLLM | One gateway in front of every model | Swap models without touching callers |
| API | FastAPI + Python | Orchestrator on port 8091 | Fast, typed, async-native |
| Database | SQLite | Persistence for the orchestrator | Zero config, zero network |
| Search | Qdrant, FAISS + BM25 | RAG retrieval | Vector + keyword hybrid search |
| Embeddings | bge-m3, nomic-embed-text | Turn documents into vectors | Run locally in Ollama |
| Public site | Astro | vorluxai.com | Static-first, fast |
| Automation | n8n | Workflow automation | Visual workflows, self-hosted |
| Scheduling | systemd timers + APScheduler | Recurring jobs | Survive reboots, logged by the OS |
| Hardware | 2x NVIDIA DGX Spark | Primary and assistant server | 128 GB unified memory each, one GB10 GPU |
Every component runs on our hardware or on the client’s. Inference never leaves the machine.
How It All Fits Together
flowchart TB
subgraph CLIENT["Client Layer"]
SITE["Astro Site<br/>vorluxai.com"]
N8N["n8n Workflows<br/>:5678"]
end
subgraph API_LAYER["API & Orchestration"]
ORCH["FastAPI Orchestrator<br/>:8091"]
GW["LiteLLM Gateway<br/>:4000"]
end
subgraph INFERENCE["Inference Layer"]
OLLAMA["Ollama<br/>:11434"]
VLLM["vLLM on Spark 2<br/>:8200"]
RAG["Qdrant + FAISS/BM25<br/>RAG Search"]
end
subgraph DATA["Data Layer"]
SQLITE[("SQLite<br/>Orchestrator DBs")]
end
N8N --> ORCH
ORCH --> GW
GW --> OLLAMA
GW --> VLLM
ORCH --> RAG
ORCH --> SQLITE
style CLIENT fill:#0B1628,color:#FAFAFA
style INFERENCE fill:#059669,color:#fff
style DATA fill:#F5A623,color:#0B1628The Models We Run
Not every task needs the same model. We keep a library on disk and route each request to the right one through LiteLLM. On 2026-10-09, ollama list on our main DGX Spark (spark-43d5) showed 20 models, including:
| Model | Download size | What we use it for |
|---|---|---|
| qwen3.6:35b | 23 GB | General reasoning and Spanish writing |
| gemma4:26b | 17 GB | Second opinion, long-form drafts |
| qwen2.5vl:72b | 48 GB | Reading scanned documents and images |
| deepseek-r1:14b | 9.0 GB | Step-by-step reasoning |
| llama3.1:8b | 4.9 GB | Fast, light tasks |
| qwen2.5-coder:7b | 4.7 GB | Code generation and review |
| bge-m3, nomic-embed-text | 1.2 GB, 274 MB | Embeddings for search |
The rest are our own fine-tunes (the apprendere and j4sgon-finance families) and small vision and avatar models. They are not all loaded at once. At the same moment, ollama ps showed three models in memory: llama3.1:8b, qwen2.5-coder:7b and bge-m3. Ollama loads a model on first request and unloads it after a timeout, so memory goes to what is in use.
A 72B model at 48 GB would not fit next to others on a 32 GB laptop. It fits here because each Spark has 128 GB of unified memory shared by CPU and GPU.
What Runs on a Schedule
The system does not only answer requests. On 2026-10-09, systemctl --user list-timers --all on the main Spark listed 68 timers. They cover:
- Content: research, draft, review and publish steps, each with a human gate before anything goes live
- Quality: test runs, link checks on the built site, knowledge-base updates
- Monitoring: health checks and watchdogs that restart a service after repeated failed probes
- Backups: daily git bundles of every repo and database snapshots
A watchdog restarts the orchestrator only after three failed health probes in a row, so a single slow response does not kill a working service. We cover the deployment side in our local deployment guide.
Why Open-Source Matters
Every component in our stack is either open-source or built by us in-house. This isn’t ideological: it’s practical.
- No licence fees: our clients don’t pay software licences for the stack. Hardware and our time are the costs.
- No vendor lock-in: if Ollama disappears tomorrow, we switch to llama.cpp or vLLM. Same models, different runtime.
- Auditability: regulated clients can inspect the code that touches their data. That supports the data-protection-by-design duty in GDPR Article 25, though the duty itself is about your whole process, not your licence.
- Mature projects: Ollama, vLLM, n8n and Astro are widely used, actively maintained projects, not experiments.
Compared to Cloud-Dependent Stacks
| Aspect | VORLUX AI (local) | Typical cloud stack |
|---|---|---|
| Data location | Your hardware | The provider’s data centres |
| Running cost | Electricity and maintenance | Per-token or per-seat fees |
| Internet required for inference | No | Yes |
| AI vendor as GDPR processor | No | Yes, needs an Art. 28 contract |
| Model switching | Change one LiteLLM route | Depends on the provider’s catalogue |
| Uptime dependency | Your power and hardware | Their SLA |
| Audit trail | Full local logs | Provider-dependent |
The cloud stack isn’t wrong for everyone. For businesses processing sensitive data under European regulation, local deployment removes the AI vendor from the data flow. We explored the cost side in our cost analysis.
What This Means for You
When we deploy AI for your business, you get this same architecture, sized to your hardware and workloads. A small office does not need two DGX Sparks: a mini PC or a Mac with enough memory for one 7-8B model covers many tasks. We size it by testing your workload, not by guessing.
See It in Action
We run live demos of this stack during our free assessment calls. No slides, no mockups: the actual system, running actual models, on your sample queries.
Book your free 15-minute assessment and see what local AI looks like in practice.
This is post 2 of our Launch Week series. Yesterday: Local AI Readiness Checklist. Tomorrow: Our Services and Pricing.
External references: Ollama | vLLM | LiteLLM | n8n | Astro | GDPR on EUR-Lex
Next steps
- Check which of the models above fits your own hardware
- Map one recurring task to a scheduled job
- Compare the running cost of a local stack with your current cloud bills
Related reading
- Best Local LLM Models for Q2 2026: Practical Comparison for SMEs
- Cloud vs Local AI: Real Cost Analysis for Spanish SMEs in 2026
- Cloud vs Local AI Cost Benchmarks
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.