There is a conversation happening in German industry that almost never reaches the conference stage. It goes like this: we would use AI properly, but our data cannot leave the building.
It comes from defence suppliers, from Mittelstand manufacturers with process knowledge worth more than their machinery, from hospitals, from anyone under sector regulation. That is not paranoia. They have read their own risk position correctly.
The public conversation has almost nothing to offer them. There is enthusiasm for local AI, plenty of it, and there is a great deal of advocacy about sovereignty. What there is not, in any usable form, is a straight answer to the question these companies actually ask: what hardware, what will it cost, and what will it realistically do?
What follows is my attempt at that answer, and at the more awkward question underneath it.
The debate is usually presented as local versus cloud. That framing is wrong and it produces bad decisions in both directions.
Almost no organisation should run everything locally. The frontier models are still ahead of anything you can host yourself, and for a lot of everyday work, drafting, summarising public material, general research, the sensitivity simply does not justify the cost.
Equally, almost no organisation in a regulated or IP-sensitive position should run everything in a public cloud, whatever the contractual assurances.
The professional answer is routing. Sensitive workloads stay inside your boundary. Everything else goes to whichever model is best for the task. The architectural work is not choosing a side; it is defining the boundary and enforcing it automatically, so the decision is not left to whoever is typing.
Before any hardware discussion, one thing must be established: which workloads genuinely cannot leave?
In my experience the honest answer is smaller than the initial one. Companies begin by saying "all of it" and end up with something defensible: contract and pricing data, engineering documents, personnel matters, anything covered by customer confidentiality clauses. In some sectors, anything touching a defined category of protected information.
Everything outside that category can use the best available model. That distinction determines the size of the machine you need, and it is the difference between a workstation and a rack.
Get this wrong in the cautious direction and you buy hardware you did not need. Get it wrong the other way and no amount of hardware saves you. I have seen both, and the second one is worse.
Local inference comes down to one thing above all others: memory. A model has to fit into it. Speed, concurrency and quality all follow from what you managed to fit.
This gives four practical tiers.
--> A single power user. One workstation with a current high-end consumer GPU, or an Apple machine with large unified memory. Runs a capable mid-sized open-weight model at usable speed for one person. Suitable for an executive, a legal team of two, an R&D lead. This is the tier where people wrongly conclude local AI does not work. Because they tried it on a laptop.
--> A department, roughly ten regular users. A single server with professional-grade accelerators and substantially more memory. Runs a larger model with acceptable concurrency. This is where most Mittelstand deployments actually land, and where the economics start looking sensible against per-seat cloud pricing.
--> A business unit, fifty or more users. Multiple accelerators, proper inference serving, queuing, and someone whose job includes keeping it running. At this point it is infrastructure, not a project.
--> Organisation-wide, or air-gapped. A rack, redundancy, formal operations. Reserved for critical infrastructure, defence and public sector, where the requirement is absolute rather than economic.
I have deliberately not printed prices here. Hardware pricing in this category moves fast enough that a published figure is misleading within a quarter, and the configuration that is correct for you depends on which models you actually need to run. What I would insist on is that any advisor gives you named configurations and current quotes for your case, in writing, before you commit. An adviser who will not put a specific machine on paper is not advising you.
Local infrastructure is a capital cost with low marginal usage cost. Cloud is the reverse. The crossover therefore depends almost entirely on how heavily you actually use it.
Light, occasional use across a small team rarely justifies buying hardware on economics alone. Heavy sustained use across a department frequently does, and the payback period is often shorter than people expect.
But economics is usually the second argument. The first is that certain data cannot go out, in which case the cloud comparison is not a comparison at all. Be clear which argument you are making. If sovereignty is the driver, say so and stop pretending the spreadsheet decides. If cost is the driver, then the spreadsheet must actually be run, with realistic usage assumptions rather than optimistic ones.
There is also a third option that is often overlooked: hosted inference with a European provider under European jurisdiction. It does not satisfy an air-gap requirement, but for a large share of cases it satisfies the actual legal requirement at a fraction of the operational burden.
A word of caution about the emerging market around this.
Installing a local model is not a service. The tooling has got good enough that it is an afternoon's work, and anyone charging serious money for it is selling you something you could have had for nothing.
What is actually difficult, and worth paying for, is everything around it.
Deciding which workloads must stay inside, with reasoning that survives a legal review. Evaluating models against your own tasks rather than against public benchmarks, because a leaderboard position tells you nothing about whether the thing reads your specifications correctly. Designing the routing between local and external. Access control and audit logging, so you can show afterwards who asked what. And then operations, which is the part everyone forgets: models change, quality drifts, capacity gets exceeded, and someone has to notice before the users do.
That is an architecture and operations engagement. The installation is an afternoon inside it.
The strongest open-weight model families now come substantially from Chinese laboratories, and increasingly so does the cost-effective inference hardware.
For a European company with a sovereignty requirement this creates a awkward question, and it deserves a serious answer rather than a reflex in either direction. An open-weight model that runs entirely on your own hardware, disconnected, does not transmit anything to anyone. The origin of the weights and the security of the deployment are separate questions. But they are both real questions, and the evaluation has to cover licence terms, provenance, and behaviour under your own testing.
What I would not accept is an advisor who dismisses the entire category without having evaluated it, or who recommends it without being able to explain the licence. Assessing both ecosystems honestly is not a political position. It is simply the job.
If this is a live question in your organisation, the first step is not procurement. It is a short, bounded assessment that produces three things: the defensible list of workloads that must stay inside, a named hardware configuration with a current quote, and a routing architecture that makes the boundary automatic rather than a matter of discipline.
Everything after that is implementation, and implementation is the easy part.
Christian Rose (罗仕) is the Founder and Managing Director of Q-Bridges GmbH, a Berlin-based strategy, AI and technology advisory firm. He has 25 years of consulting experience and has worked as a Managing Director and Partner in both Germany and China. Q-Bridges is vendor-neutral: it sells no licences, resells no hardware and takes no commissions.
→ www.q-bridges.com · LinkedIn
Sovereign AI covers the sovereignty assessment, model evaluation against your own tasks, named hardware configurations and the routing architecture.
See Sovereign AI