AI agents are past the demo stage. They read documents, call tools, and hand work back to people every day. Finding an honest answer to what they will return for your company is harder. Most ROI percentages online come from vendors selling agent platforms. This guide explains what agents do, where the return actually comes from, and how to measure it on your own work, with the data staying on your hardware.

What AI agents actually do
An AI agent isn’t a chatbot. A chatbot responds to questions. An agent plans, executes and iterates: it breaks a task into steps, uses tools to complete each step, evaluates the results and adjusts its approach.
flowchart LR
TASK["Business Task"] --> PLAN["Agent Plans<br/>Steps"]
PLAN --> TOOL1["Uses Tool 1<br/>(Database)"]
PLAN --> TOOL2["Uses Tool 2<br/>(Email)"]
PLAN --> TOOL3["Uses Tool 3<br/>(CRM)"]
TOOL1 --> EVAL["Evaluates<br/>Results"]
TOOL2 --> EVAL
TOOL3 --> EVAL
EVAL -->|"Not done"| PLAN
EVAL -->|"Complete"| OUTPUT["Delivers<br/>Result"]
style TASK fill:#1E293B,color:#FAFAFA
style PLAN fill:#F5A623,color:#0B1628
style EVAL fill:#059669,color:#FAFAFA
style OUTPUT fill:#059669,color:#FAFAFAPractical examples:
- Invoice processing agent: receives an invoice PDF → extracts vendor, amount and date → validates against the PO → flags discrepancies → routes for approval → updates the accounting system
- Customer support agent: reads a ticket → searches the knowledge base → drafts a response → escalates if confidence is low → logs the resolution
- Lead qualification agent: receives an inquiry → researches the company → scores fit → drafts a personalised response → schedules follow-up
About those ROI headlines
You will see figures such as “171% average ROI” or “280-520% first-year ROI for SMEs” quoted for agentic AI. We used to repeat some of them here. We no longer do: the ones we traced lead to blog posts by companies that sell agent platforms or consulting, not to a published study with a method, a sample and a definition of ROI. A number like that tells you what a vendor wants you to expect, not what your invoices will do.
The return from an agent comes from three places, and all of them can be measured in your own business within weeks:
- Hours no longer spent on a repetitive task (minutes per item × items per month).
- Errors caught earlier, such as duplicate invoices or wrong VAT codes, which have a cost you can look up.
- Response time, where it drives revenue, such as first reply to a sales inquiry.
Against that, count the setup effort, the hardware, the time a person spends reviewing the agent’s output, and the cases it gets wrong.
Measure it on one workflow first
Pick one high-volume, low-risk workflow and record a baseline for two weeks before the agent touches it:
| Metric | How to measure | Why it matters |
|---|---|---|
| Time per item | Timestamp start and finish on 20 real items | The core of the saving |
| Volume | Items per month from your existing system | Multiplies the saving |
| Error rate | Items corrected after the fact | Quality, before and after |
| Human review rate | Share of agent outputs a person has to fix | Hidden cost of the agent |
| Cost | Hardware, setup, electricity or API fees | The denominator |
Run the agent alongside the manual process for a few weeks, compare, and only then extend it. Our local deployment guide for your first three agents includes the targets we use for review and override rates.
How we run it ourselves
VORLUX AI is a small company that runs its own content, research and quality-control pipelines as agents on local models, on a workstation in our office. The stack is open source: no cloud APIs, no per-query costs, no data leaving the building. The same architecture is what we deploy for clients.
Why local agents beat cloud agents
| Factor | Cloud agents | Local agents |
|---|---|---|
| Cost per query | Charged per token | Close to zero after the hardware |
| Data privacy | Data sent to the provider | Data stays local |
| GDPR compliance | Requires a DPA and transfer assessment | Privacy by design is simpler |
| Availability | Depends on API uptime | Runs on your network |
| Customisation | Limited to API options | Full fine-tuning possible |
| Vendor lock-in | High | None (open-weight models) |
To see where local hardware becomes cheaper than paying per token for your volume, use current API prices and measured local speeds in our break-even guide.
Under the EU AI Act, automated decision-making requires transparency about how decisions are made. Local agents with chain-of-thought reasoning help: every reasoning step can be logged on your hardware.
Getting started: 3 agent patterns for SMEs
Pattern 1: Document intelligence
Best for: law firms, consulting, accounting
Deploy DeepSeek R1 14B with an n8n workflow that reads incoming documents, extracts key data, classifies by type and routes to the right person.
# Install the model and test document classification
ollama pull deepseek-r1:14b
curl http://localhost:11434/api/generate -d '{
"model": "deepseek-r1:14b",
"prompt": "Classify this document: Invoice from Supplier X, EUR 5,400, payment due 30 days. Return JSON: {type, priority, department}"
}'Hardware: Mac mini M4 with 16 GB or more Measure: minutes per document before and after, and the share of documents routed correctly
Pattern 2: Customer response
Best for: e-commerce, services, hospitality
Deploy Qwen 2.5 7B with a RAG pipeline over your FAQ or product database. The agent drafts first-level answers and escalates complex cases with full context.
Hardware: a Mac mini or a PC with a 12 GB GPU (small boards such as the Jetson Orin Nano fit 3B models, not 7B) Measure: share of tickets answered without rework, and time to first reply
Pattern 3: Business intelligence
Best for: any SME with data
Deploy Gemma 3 4B (multimodal) to process reports, invoices and images. The agent generates weekly summaries, flags anomalies and suggests actions.
Hardware: Mac mini M4 with 16 GB or more Measure: hours of manual reporting per week, before and after
Funding
Spanish SMEs may be able to cover part of the cost with public programmes. The conditions and deadlines change, so check our Kit Digital and grants guide for what is open now.
Ready to deploy AI agents in your business? Schedule a free 15-minute assessment and we’ll help you pick the workflow to measure first.
Related: n8n + MCP Tutorial | DeepSeek R1 Review | Cloud vs Local Break-Even | VORLUX AI Stack
Next steps
- Pick one high-volume, low-risk workflow to audit.
- Record a baseline for two weeks before any deployment.
- Review our guide on local deployment for SMEs.
- Test DeepSeek R1 with n8n for document intelligence.
Related reading
- Your First 3 AI Agents: A Local Deployment Guide for SMEs (2026)
- Kit Digital and Spanish AI Grants: 2026 Guide
- Llama 4 Scout and Maverick: A Practical Review for Local AI Deployment
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.