Every week another company announces it is investing in AI. Most of them just mean they signed up for cloud API subscriptions. Some are asking a different question: would running the same work on hardware they own cost less?
The public case studies are often quoted with big numbers attached. Here is what they actually say, who published them, and how much weight each can bear.
The enterprise evidence, read carefully
Dell AI Factory with NVIDIA: a vendor-commissioned model. Enterprise Strategy Group (ESG) produced an economic analysis commissioned by Dell, which sells the hardware in question. It modelled a hypothetical mid-to-large organisation spending USD 1.96 million on the platform and estimated USD 25.9 million in “savings and benefits” over four years, an ROI of 1,225%. Two things to keep in mind before quoting it: the result is a model, not a measured deployment (Dell itself notes that actual results will vary), and most of the modelled benefit comes from productivity, time to value and risk reduction, not from a smaller cloud bill. It is useful as a list of benefit categories to consider, not as a forecast for your company.
37signals: a company’s own disclosure. The makers of Basecamp and HEY published their numbers themselves. They spent USD 3.2 million on cloud in 2022, and CTO David Heinemeier Hansson estimated that moving the non-storage part (about USD 2.3 million a year) to their own servers, at roughly USD 840,000 a year all-in, would save about USD 7 million over five years. This is the strongest public cloud-exit case because the company that paid the bills is the one reporting them. It is also general web infrastructure, not AI inference, and at a scale far above a typical SME, so it shows the mechanism (steady, predictable workloads are cheaper on owned hardware) rather than a percentage you can copy.
What neither case gives you is your own break-even. That depends on your volume, the model tier you actually need and the hardware that fits it.
Why SMEs look at local AI at all
The reason is structural, not a headline ROI figure. Cloud APIs charge per token, so the bill rises with every document processed and every question asked. A machine you own has a fixed purchase price plus electricity, so the marginal cost of one more query is close to zero. Whether that fixed cost pays off depends on how much you use it:
- Low, occasional use: pay-per-token is usually cheaper. A local box sits idle.
- Steady daily use (document processing, internal knowledge search, support drafts): the fixed cost is spread over many queries and local becomes competitive.
- Sensitive data: even where cost is a draw, keeping data inside your perimeter can decide the question on its own.
Work out your own break-even
Rather than borrowing someone else’s percentage, use current API prices and measured local speeds. Our cloud vs local break-even guide lists today’s per-token prices, measured tokens per second on hardware from a EUR 250 board to a Mac mini, and a 12-line script that returns the month in which owned hardware becomes cheaper for your workload. The ROI framework turns that number into a business case.
Why cloud repatriation keeps coming up
The trend has a name: cloud repatriation. Some companies that moved everything to the cloud are bringing steady workloads back to their own hardware. For AI, three reasons recur:
-
Cost predictability. Cloud AI pricing changes as providers launch and retire models. Owned hardware is a fixed cost.
-
Data sovereignty. Under GDPR and the EU AI Act, in general application since 2 August 2026 (Regulation (EU) 2024/1689; see the calendar after the Omnibus), sending proprietary data to cloud APIs creates compliance work. Local deployment keeps it inside your security perimeter. We covered this in our GDPR and AI convergence analysis.
-
Performance. Local inference has no network round trip, no rate limits and no provider outages. It is slower per token than frontier cloud models on small hardware, so match the model to the task.
What VORLUX AI deploys
A typical small deployment is a Mac mini M4 or similar box running open-weight models through Ollama. Hardware, model selection, integration and training are scoped and quoted per project.
For Spanish businesses, regional and national programmes can fund part of an AI project when a call is open; Kit Digital’s calls for beneficiaries have ended. See our grants guide for what is open now.
The hardware options and model recommendations are in our local LLM comparison guide.
Related reading
- Local AI ROI Framework: How to Calculate Cloud Savings for Your SME in 2026
- Your First 3 AI Agents: A Local Deployment Guide for SMEs (2026)
- AI Agents for SME Automation: Where the ROI Really Comes From
The bottom line
The best public evidence for owning your infrastructure comes from companies that disclosed their own bills, like 37signals, and it shows a mechanism rather than a universal percentage. Vendor-commissioned models, like Dell’s, are worth reading for what they count as benefits, not for their headline ROI. The only number that should drive your decision is your own break-even, and it takes about fifteen minutes to work out.
Want help working out your break-even? Request a free assessment: we will model your usage and compare cloud and local costs for your case.
Work with us
We size the model and the machine by measuring, not by guessing. If you want to see your own task running on real hardware, book a 15-minute call or see how we work in consulting.