Skip to content

Strategy · 10 August 2026

Running AI in-house: what it costs and when it pays off

A 96 GB graphics card works out at roughly CHF 5,600 a year over three years; the same volume of work through a Swiss endpoint costs under CHF 400. Where the tipping point sits.

Author

Reto Lutz

Founder & AI trainer

In brief

  • An NVIDIA RTX PRO 6000 Blackwell with 96 GB starts at CHF 10,808 in Switzerland (as of 10 August 2026); written off over three years and with power and cooling, that is at least CHF 5,581 a year.
  • Ten people working intensively with AI generate a volume that costs around CHF 347 a year through a Swiss endpoint; the card only pays off from roughly 160 such people.
  • None of the Swiss statutes and supervisory documents I checked requires local hosting - the Swiss Bar Association lists on-premise as one of three permissible routes.
  • What running it yourself changes is the distribution of obligations, not their number. Logging and the processing policy under the data protection ordinance become your job.

As of 10 August 2026. All prices were retrieved on that day, and all legal references are to the versions in force then. Both move: the memory market is pushing hardware prices, and one of the quoted prices changes on 1 September, when cloudscale.ch raises its GPU rates. I am not a lawyer; this is not legal advice. Sources with retrieval dates at the end.

Utility room in a small Swiss company: a tower computer stands on a metal shelf with its side panel removed, a large graphics card visible inside. Coiled network cables sit beside it, a half-open cardboard box on the floor, a drain and a mop by the wall.
What AI in-house actually looks like - not a data centre, but a shelf in the utility room.

The question usually arrives in this shape: what does it cost to run our own top-tier AI in-house, instead of sending data to an American provider? The question contains a category error. A frontier AI cannot be bought. What can be bought is hardware that runs a good open-weight model - and the gap between that model and the best closed one is the real price, not the graphics card.

A frontier AI is not for sale

An open-weight model is a language model whose trained parameters are released for download, so that it runs on your own hardware without any connection to the vendor. Disclosed training data and training code are explicitly not part of that - those are what the Open Source Initiative’s definition of open source AI requires.

The gap to the top has narrowed, and it can be quantified, though only with a date attached. On the Artificial Analysis Intelligence Index, the best open model stood at 60 points on 10 August 2026 (Kimi K3), the best closed one at 63. At the end of April 2026 the difference was six points; a year before that, around thirteen.

Those three points obscure the distribution, though. On the hardest exams the gap remains clear: on Humanity’s Last Exam the three leading open models reached 34 to 36 per cent at the end of April, against 44 per cent for the closed model then in front; on agentic coding, 43 to 46 against 61 per cent. For summarising and research an open model is good enough. For the tasks that separate models, it is not.

For Swiss organisations, Apertus is the obvious candidate: the open language model from EPFL, ETH Zurich and the national supercomputing centre CSCS, available in version 1.5 since 24 July 2026, licensed under Apache 2.0 and therefore usable commercially. It is only released after you accept a usage policy, which among other things requires you to indemnify ETH and EPFL against third-party claims and to refresh a filter file every six months. Meta’s Llama 4 is not under an open-source licence at all, but under a community licence with an attribution requirement. Anyone bringing a model in-house reads the licence, not the label.

What the hardware costs

The currency here is graphics memory. A model has to fit, or it does not run - or it runs so slowly that nobody uses it. How much memory a model needs depends on quantisation, that is, how far the weights have been shrunk. Take a 27-billion-parameter model from Qwen: in 4-bit it occupies 16.7 GB, in 8-bit 28.6 GB, unquantised 53.8 GB.

The benchmark for the calculation below is the middle class: OpenAI’s gpt-oss-120b downloads 65 GB and, according to the vendor, needs a card with 80 GB - in Switzerland that means the RTX PRO 6000 Blackwell with 96 GB.

01 · up to 24 GB

Small model

A 27B model in 4-bit takes 16.7 GB; Gemma 4 with 31 billion parameters downloads 20 GB. Fits on a GeForce RTX 5090 with 32 GB.

from CHF 3,900

02 · up to 96 GB

Mid-size model

gpt-oss-120b downloads 65 GB and needs an 80 GB card according to OpenAI. That means the RTX PRO 6000 Blackwell Workstation with 96 GB.

from CHF 10,808

03 · over 140 GB

Large model

Apertus 1.5 with 70 billion parameters is 144.6 GB of files - even an H200 NVL with 141 GB is not enough. It takes several cards or a quantised build.

CHF 29,999 per card

Memory requirements and the matching card. Entry prices from individual retailers via Toppreise.ch, retrieved 10 August 2026, including Swiss VAT. The model sizes are download sizes of quantised standard builds, not memory use in operation.

For data-centre cards the Swiss market is thin. The H200 NVL with 141 GB is listed at CHF 29,999 - by exactly one retailer. There is no regular Swiss listing for the H100 at all, and no public price for the newer B200 cards. The direction is not encouraging either: NVIDIA raised the list price of the RTX PRO 6000 to USD 13,250 in June 2026, 55 per cent above its March 2025 launch price, and a European reseller is signalling a further 10 to 20 per cent for the DGX Station.

At the same time a popular argument has fallen away: the Mac Studio as the cheap route to a lot of memory. Apple dropped the 512 GB option in March 2026; on 10 August 2026 the Swiss configurator offered only 36, 64 and 96 GB. If you want more than 96 GB of unified memory, you end up with an NVIDIA DGX Spark at 128 GB from CHF 4,520 - or straight at the DGX Station with 748 GB for around USD 100,000 excluding VAT. Right now there is nothing in between.

The calculation: when the card pays off

Against an open model on a Swiss endpoint, your own card only pays off at a volume equivalent to roughly 160 people working intensively with AI. How that number comes about is set out here in the open - it is a calculation with declared assumptions, not a measurement.

Hardware side, for one RTX PRO 6000 Blackwell Workstation with 96 GB: CHF 10,808 to buy, written off over three years, is CHF 3,603 a year. This variant draws 600 watts according to NVIDIA; running continuously that is 5,256 kilowatt hours. The EKZ business tariff for 2026 is 24.44 centimes in the winter half-year and 20.14 in the summer half-year; I use the higher figure throughout, which gives around CHF 1,285 - averaged across the year it would be about CHF 1,170. For cooling and infrastructure there is no Swiss SME figure; as a stand-in I use the global data-centre average PUE of 1.54, so around CHF 693. Together that is CHF 5,581 a year for a single card - without the computer around it, the room, or anyone to look after it.

01 · purchase

CHF 3,603

CHF 10,808 for the card, spread over three years. Amazon shortened the useful life of part of its server fleet from six years to five in 2025, citing the pace of AI.

02 · electricity

CHF 1,285

600 watts around the clock is 5,256 kWh. Calculated with the higher EKZ winter rate for business customers, 24.44 centimes per kWh excluding VAT.

03 · cooling

CHF 693

Roughly half a watt of cooling and infrastructure per watt of compute - a borrowed average from the Uptime Institute's global data-centre survey 2025.

CHF 5,581 a year for one card. Not included: the computer around it, the room, maintenance, and the person who runs it.

The imprecision runs in both directions, but not equally far. On electricity I deliberately calculate at the top end, which is a good hundred francs too much; continuous operation is an assumption too, since an SME machine sits idle most of the time - which lowers the power bill but not the depreciation. Against that stand items missing entirely: for staff time and maintenance contracts I found no citable source, only blog posts without method. Anyone who prices them in ends up above CHF 5,581, not below.

Now the other side. Assume one person works intensively with AI: 50 requests per working day, 2,000 tokens in and 700 tokens out each. Over 220 working days that is 22 million input and 7.7 million output tokens a year; for ten such people, 220 and 77 million. Through Infomaniak, which by its own account runs Apertus 1.5 on servers in Switzerland, that costs CHF 0.70 and CHF 2.50 per million tokens respectively, both excluding VAT - CHF 347 a year for the whole team.

For the single card to cost the same as that endpoint, you would need around 4.8 billion tokens a year. That is sixteen times what the modelled ten-person team produces. The tipping point therefore sits at roughly 160 people working intensively - and whether one card can actually serve 160 people is a separate question.

Both numbers hold, for different questions. Compare open model against open model and you land at 160 people. Set your own card against a frontier model through an endpoint - the same volumes cost around USD 3,000 a year with Claude Opus 5 at USD 5 and 25 per million tokens - and you land at just under twenty, but the card then buys you the weaker model. That second calculation is also rough: the dollar prices are net, the Swiss figures partly include and partly exclude VAT, and the newer Anthropic models produce around 30 per cent more tokens for the same text according to the vendor.

The reason for the imbalance is utilisation. Low per-token costs only arise when a card runs at full load and handles many requests at once. An SME uses a few per cent of it - and pays for all of it. If you pick a provider instead of a card, factor in vendor lock-in as part of that decision.

The middle path: Swiss infrastructure without your own rack

Between the machine in the basement and the US cloud sits a layer that is almost always missing from this debate: Swiss providers running open models on Swiss servers. Infomaniak is the most transparent - the catalogue is on the website with prices, and the endpoint follows the OpenAI standard, so moving an existing application over is manageable. Exoscale runs managed inference endpoints and explicitly names its Geneva zone for them, but only releases GPUs after a manual review. cloudscale.ch bills by the second and is raising prices on 1 September 2026. And CSCS, co-developer of Apertus, sells compute time to private companies too, from a minimum of CHF 3,000.

So hardware in Switzerland is transparently priced and operating services almost never are: Swisscom, Green and several smaller providers advertise AI infrastructure without naming figures. If you want to compare, you have to ask.

And if you want processing kept in Switzerland with one of the large providers, you pay for it with a model lag. In the Azure region Switzerland North, Microsoft’s own overview lists only gpt-4.1, gpt-4.1-mini and gpt-4o for regional deployments; gpt-5.1 is available in Europe only in Sweden Central. That is the same trade-off as with the business plans of ChatGPT, Claude and Copilot - one step further along.

What Swiss data protection law actually requires

I went looking specifically for the Swiss rule that mandates local hosting - in the revDSG (the revised Swiss Federal Act on Data Protection) and its ordinance, in the criminal code, at the FDPIC (the federal data protection authority), at the Bar Association and at FINMA. I found none; what stands there throughout are conditions, not location requirements.

The Swiss Bar Association’s AI guidance lists running the system inside the firm’s own network as one of three permissible routes, on equal footing with sourcing it from a provider under the outsourcing rules and with the informed client’s consent. A legal opinion from the University of Zurich, commissioned by the association, holds that Art. 321 of the criminal code implies no duty to choose Swiss providers only; what matters is proportionality to the risk. FINMA’s supervisory notice 08/2024 covers governance, testing and outsourcing controls - it prescribes no location. And the FDPIC considers high-risk AI processing permissible, provided appropriate safeguards are in place and a data protection impact assessment has been carried out.

What running it yourself changes is the distribution of obligations, not their number. Two questions fall away: if you run it yourself you have no processor, so Art. 9 revDSG does not arise - and no data is disclosed abroad, so neither do Art. 16 and 17. That follows from the statutory definitions rather than from any statement by an authority; it is my reading, and nobody told me so.

What remains is more work, not less. Art. 8 revDSG requires data security proportionate to the risk, regardless of who operates the system. The ordinance spells that out: determine the protection required, control access and users, keep a processing policy where sensitive personal data is processed extensively - and log activity, in wording revised as of 1 December 2025: instead of “reading”, the ordinance now speaks of “accessing” the data. Logs must be kept for at least a year, separately from the system, and failing the minimum requirements carries a fine of up to CHF 250,000. Whatever you previously left to a provider, you now do yourself.

For those bound by professional secrecy there is one more point: Art. 321 of the criminal code names lawyers, notaries, doctors and auditors bound to confidentiality under the Code of Obligations - but not fiduciaries or bookkeepers. That does not leave them free; they remain bound by the revDSG and by contract. How this plays out for law firms is set out in the article on professional secrecy, and the groundwork in the overview of the revised data protection act.

One distinction while we are here: moving to on-premise does nothing against manipulated training corpora, because a locally run model carries the same data behind it. Against data leaving your organisation it helps a great deal. Two different risks, two different answers.

When running it yourself does pay off

Four situations in which the calculation flips. First, genuine network isolation: if a machine or a dataset must have no outside connection, the alternative disappears. Second, sustained load - process large volumes of documents in batches and keep the card genuinely busy, and the tipping point moves down. Third, a contractual requirement from a client that permits no third party; that is a contract problem rather than a legal one, but it binds just as tightly. Fourth, fine-tuning on your own data.

For everything else, running it yourself is the more expensive route to the weaker model.

Frequently asked questions

What does it cost to run an AI model in-house?

At least around CHF 5,581 a year for a single graphics card with 96 GB: CHF 3,603 in depreciated purchase cost, CHF 1,285 in electricity and CHF 693 in cooling, as of 10 August 2026. Not included are the computer, the room, maintenance and staff time. The same volume of work costs a ten-person team around CHF 347 a year through a Swiss endpoint.

Is a Mac Studio enough for a large model?

Not for new purchases in Switzerland: on 10 August 2026 the configurator offered at most 96 GB, and a 70-billion-parameter model such as Apertus 1.5 occupies 144.6 GB unquantised. For smaller models up to around 30 billion parameters a Mac Studio still works.

What does Apertus offer a Swiss SME?

Apertus is a Swiss-developed language model licensed under Apache 2.0 that you can either run yourself or obtain from a Swiss provider - at Infomaniak for CHF 0.70 and CHF 2.50 per million tokens. On how version 1.5 performs against other models I cannot say anything reliable: the model card contains charts, not figures.

Does any Swiss supervisory authority require local hosting?

None of the statutes and supervisory documents I checked requires it. The Bar Association’s AI guidance names three routes of equal standing, and FINMA prescribes no location. What is required before any use is clarity on who has access to the data you enter and where it is cached.

In practice

The question of running your own AI in-house breaks into three answers. Technically you do not get a frontier model but an open one - enough for most office work, not enough for the hardest. Economically a card at roughly CHF 5,600 a year only pays off at a volume equivalent to some 160 people working intensively. And legally, in the sources I checked, neither statute nor supervisor requires running it yourself - it merely shifts which obligations you carry, and there are rather more of them.

The middle path is where most organisations should land: an open model on Swiss servers, taken as a service. The price is the return of the processor - Art. 9 revDSG and its duty of assurance apply again, and the model remains an open one with the gap described above. Testing that costs nothing: at Infomaniak the first requests against Apertus run on free credits, with no contract. That settles in an afternoon whether an open model solves your tasks at all - the cost question only arises afterwards. Which tasks it should carry in your organisation is what AI training for teams clarifies before the tooling question.


Sources (checked 10 August 2026):

Tags

#ai-strategy #sme #switzerland #data-protection #apertus #on-premise
Book an intro call