We also sell on-premise AI. So what follows works against us, and it is worth saying that up front rather than pretending to be neutral.
But we have seen enough projects stall in month six to prefer an honest conversation at the start over an awkward one later. This article collects the numbers that usually never make it into the meeting room.
The figure that cuts against the story
Read the trade press and the narrative is that open models have caught up, that private AI is the future, and that companies are bringing everything in-house.
Market data says otherwise. According to Menlo Ventures' research on enterprise generative AI adoption, the share of enterprise workloads running on open-weight models fell from 19% to 11% in a year.
The interesting part is why. It is not that open models got worse: on benchmarks they very much caught up. The reasons cited are three others, and all of them are boring: commercial licence constraints, real total cost once every piece is counted, and the difficulty of verifying a model's provenance.
Those are precisely the three things that never show up in a demo.
The break-even, translated into something you can picture
The economic threshold above which self-hosting beats cloud provider APIs sits, in the available analyses, at around a billion tokens a month, with payback over roughly thirty months. At ten billion a month, payback drops to two.
How much is a billion tokens a month?
A token is about three quarters of a word. So we are talking about roughly 750 million words a month — 25 million words a day. At five hundred words a page, that is fifty thousand pages of text processed every single day, holidays included.
If you run a company of 50, 100 or 300 people, you will not get there. Not remotely, not even if everyone did nothing else.
Which leads to the uncomfortable conclusion: if your argument for bringing AI in-house is saving money, the argument does not hold. There are excellent reasons to do it, but that is not one of them, and building a business case on it means explaining to the board a year from now why the numbers do not add up.
The costs that appear after the machine
The hardware quote is the easy part, because it is a single visible number. The rest arrives afterwards, in instalments.
Somebody to run it. An AI server is not an appliance. Someone has to update drivers, manage the inference stack, watch memory, work out why it is slower this morning. If that person does not exist in-house, the cost is a managed service contract; if they do, the cost is the time they stop spending on something else. Either way it recurs, and it is not small relative to the hardware.
Model updates. Open models ship new versions every few months. Staying behind means working with a worse tool than competitors using pay-as-you-go APIs. Updating means re-testing, re-validating, sometimes retraining your customisations. It is ongoing work, not an installation.
Quality evaluation. This is the line item that is almost always missing. How do you know the model is answering well? You need a test set built on your real documents with known correct answers, and you need to re-run it on every change. Building it takes days of work from somebody who knows the domain — not the IT person, the head of purchasing. Without it you are going on impressions, and after three months impressions become "I think it gets things wrong a lot" and the project dies.
Security and backup. That machine holds, in indexed form, a substantial share of your institutional knowledge. It needs protecting, segmenting, backing up and restoring like any other critical system. If it is under a desk, it is not.
Obsolescence. AI hardware depreciates fast and model requirements keep growing. A machine bought today holds up well for three years, then it is a fresh decision.
None of these is insurmountable. But together they change the total substantially, and a comparison against APIs based on hardware alone is a rigged comparison — even when the person making it means well.
The middle option almost nobody proposes
The discussion is usually framed as binary: all cloud or all in-house. That is a false choice, and in many cases the right answer sits in between.
The reasoning is simple: not all your data carries the same confidentiality. A client's tender documents, a mechanical design, a patient record, a commercial strategy are one thing. Summarising a public article, translating marketing copy or drafting an email are another.
A small model in-house for data that must not leave, APIs for everything else. An 8-14 billion parameter model runs on a card costing a few thousand euro, and is more than enough to classify, extract from and search across confidential documents. For the rest, use the best available tool and pay by consumption, without tying up capital.
This cuts the initial investment by an order of magnitude, removes the risk of buying the wrong machine before you understand the use case, and keeps control exactly where it is needed. It also earns us less than selling you a workstation, which may be why you rarely hear it proposed.
Questions to ask before signing
Five questions. If you do not have all five answers, it is not time to buy yet.
What is the task, in one sentence, and how often does it happen. If you cannot write it in a line, the project is not ready. "We want to use AI" is not a task.
Who notices if it stops working, and how bad is that. This determines whether you need redundancy, and redundancy doubles the bill.
Who decides whether the answers are good. By name. If the answer is "we'll see", the project has no owner and will end up in a drawer.
What happens when the model is plausibly wrong. Not when it crashes: when it gives a credible, incorrect answer. If the downstream process has no human check, you have built a fast error generator.
What would you do if the same thing cost a tenth via API a year from now. That is the most likely scenario, and it belongs in the plan now.
The cost of doing nothing
All that said, there is an opposite mistake that costs just as much: postponing indefinitely because the perfect case never arrives.
Italian data shows AI adoption among SMEs remains low — the national statistics office puts it around 16% of companies with ten or more employees, against 53% for large firms — and that roughly three quarters of Italian SMEs have neither invested nor made plans to.
Meanwhile people use it anyway. Circulating estimates suggest nearly one in two Italian workers uses AI tools in their daily work, and that the large majority of them do so with tools their employer has not approved. We found those last figures only in secondary sources and report them as indicative rather than settled — but the order of magnitude matches what we see walking into companies.
Which means the real question, for many businesses, is not "should we bring AI in-house". It is: your quotes, your price lists and your client emails are already passing through tools you did not choose, and you do not know about it.
Answering that question costs far less than a workstation. It takes knowing which tools are in use, deciding which are acceptable, offering a usable company alternative and making it the path of least resistance. That is governance, not hardware. And it almost always comes first.
Questions we get asked
Our vendor says running AI in-house saves on licences. Is that true?
Below a certain volume, no. Economic break-even sits around a billion tokens a month and no SME reaches it. The saving exists in very specific cases; the reason to do it is usually something else.
A billion tokens a month — what does that look like?
Roughly fifty thousand pages of text processed every single day, holidays included. If that figure sounds absurd for your company, it is: which means saving money is not your argument.
If I buy now, will the machine be worthless in a year?
Not worthless, but overtaken. AI hardware holds up well for about three years, then it becomes a fresh decision. Worth putting in the depreciation plan rather than discovering it later.
Can we keep part in-house and part in the cloud?
Yes, and it is nearly always the right answer. A small model in-house for documents that must not leave, APIs for everything else: it cuts the initial investment by an order of magnitude and keeps control where it actually matters.
Who looks after it once you have installed it?
Somebody has to, and it is the line item missing from almost every quote. Drivers, model updates, checking answer quality, backups. If that person is not in-house it is a managed service contract; if they are, it is time taken from something else. Either way it belongs in the budget up front.
In short
Self-hosting does not justify itself on cost per token below a billion tokens a month, and no SME gets there. The share of enterprise workloads on open models is falling, for practical rather than technical reasons. Recurring costs — operations, updates, quality evaluation, security — weigh as much as the hardware or more. A hybrid setup solves most cases at a tenth of the investment. And the more urgent problem, in a great many companies, is not infrastructure: it is that everyone is already using AI, without rules.
If you want an assessment that is allowed to conclude "don't do it", get in touch. It is an answer we give fairly regularly.
Related articles: