Skip to content

// on_prem · own_server

On-premise LLM: run the model on your own servers

A language model does not have to run in someone else's cloud. Open-weights models can be downloaded and run on a server you control: in your server room, in your factory, or on an EU box we operate for you. The data never leaves the building.

This is the right call for some companies and the wrong one for others. We tell you which on the first call. For many, an EU-hosted API with a solid agreement is enough, and then you should not be buying GPUs. For others, with drawings, patient data or source code that cannot leave the premises, your own server is the only thing that passes review.

We build both versions. This page is about the local one: which models handle Swedish, what hardware it takes, roughly how the cost works, and what you get from us.

What hardware is needed?
01

When on-prem is the right call, and when it is not

The right call when the data may not leave the building. Engineering drawings, source code, patient records, defence-related material, customer data under strict contracts. If your DPO, your security lead or your biggest customer requires that nothing goes to an external server, your own server is the answer. Also right when volume is high and steady: a model running around the clock can be cheaper on your own hardware than paying per call.

The wrong call when you just want to try. A three-month pilot should not start with buying a GPU server. Also the wrong call when the task needs the strongest model there is. The best open models are good, but the largest closed models like Claude are still stronger at hard reasoning. If that is the level you need, and the data allows it, we run Claude through an EU-hosted setup instead.

Many end up in the middle. Then we build a setup where sensitive documents stay on a local model and the rest goes to a larger model inside the EU. It is not either or.

02

Open-weights models that handle Swedish

Swedish was a problem for open models for a long time. It is not anymore. These are the families we use today, and we test them against your real documents before choosing:

Mistral. European, fast, good at Nordic languages in the larger variants. Often our first choice for Swedish text on your own hardware.

Llama from Meta. A broad family in many sizes with a large tooling ecosystem. Handles Swedish well in the larger models.

Qwen. Strong models at every size, particularly good at code and structured tasks. Swedish is fine and improves with every release.

Gemma from Google. Small, efficient models that fit when hardware is limited.

GPT-SW3 from AI Sweden. Trained specifically on Swedish and the Nordic languages. Older than the others, but worth knowing about for purely Swedish tasks.

We do not throw benchmark numbers at you. They do not tell you how a model handles your quotes or manuals. We run a sample of your documents through two or three candidates, show you the result, and choose with you.

03

Hardware in plain terms

Model size is counted in billions of parameters, and that decides what hardware you need.

A 7-8 billion parameter model runs on a single good GPU in a normal workstation. That is enough for internal search, summarisation, email classification and simpler agents. For many smaller companies this is the whole solution: one machine in the server room.

Models in the 30 billion class need a heavier card or two, and fit when the answers need to be noticeably better at Swedish and at longer documents.

Models in the 70 billion class and up need several datacenter GPUs working together. That is a real server with cooling and power to plan for, not a box under the desk. This is the level for high volume or genuinely hard tasks.

We spec the machine to your need, not to what looks impressive. If you do not want to own the hardware, we rent a dedicated GPU server in the EU for you, which we operate and which runs only your model.

04

How the cost works

We give no prices here, because they depend entirely on your volume and your model. But here is how the drivers work, so you can do the math yourself.

Your own server costs hardware once, then electricity and space every month, plus our time to set up and maintain. The cost is roughly the same whether you run a thousand queries a day or a hundred thousand. That is the point: high, steady volume makes your own server cheap per query.

An API costs per token, meaning per amount of text in and out. No hardware, no electricity, no setup. The cost grows in a straight line with usage. Low or uneven volume makes an API cheap overall.

The break-even point differs between companies. We calculate it with you. Often the answer is a mix: local model for the sensitive and the frequent, API for the heavy and the rare.

05

What we deliver

A model idling on a server is not a solution. This is what a finished setup from us includes:

Model serving

The model runs as a service with an API your systems can talk to. We handle loading, memory, parallel calls and updates to new versions.

Retrieval over your documents

Manuals, contracts, drawings, email and tickets are indexed locally so the model answers from your content and shows the source. The index lives on the same server as the model.

Agents on top

A support agent, quoting assistant or internal helper that uses the local model and your tools. Everything runs inside your firewall.

Monitoring and operations

We see when the model responds slowly, when the disk fills up and when it is time to change version. You get alerts, logs and a clear picture of what is being used.

06

The data never leaves the building

With a local model no external party sees your prompts, documents or answers. No vendor, no cloud operator, no sub-processor. That is not a promise in a contract but a physical boundary: the cable does not go out.

If you choose an EU box we operate instead of your own hardware, the boundary is our server in Germany or Finland. Only your model runs on it, we never train on your data, and we sign a DPA that makes that clear. It is the next best thing after your own server and good enough for most.

Whichever setup, we log every call locally so you can show afterwards what the system did. That is what your DPO is going to ask for.

// faq

Frequently asked questions

Can a local model really handle Swedish?

Yes, the larger open models do today. Mistral and Llama in their larger variants write and understand Swedish well enough for support, summarisation and search. Very small models are weaker at Swedish, so there we test carefully before promising anything. We always run a sample of your own documents before choosing a model.

Do we need to buy a GPU server?

Not necessarily. A smaller model runs on a workstation with a good GPU that you might already have. If you do not want to buy anything at all, we rent a dedicated GPU server in the EU for you that we operate. Hardware on your premises is only necessary when the data absolutely cannot leave the building.

How do we update the model when new versions come out?

We build so the model is swappable. When a better open model arrives, we test it against the same documents as last time, and if the result holds up we swap. Your indexed documents, your agents and your integrations are not affected. It is part of operations if you have that agreement with us.

Is a local model as good as Claude?

No, not on the hardest tasks. The largest closed models are still ahead on long reasoning and complicated instructions. But for search, summarisation, classification and most support questions the difference is small or none. We say honestly when your task needs the bigger model.

Can we start in the cloud and move in-house later?

Yes, and it is often the smartest route. Start on an EU server we operate, see what gets used and how much, and move the model to your own hardware when volume or requirements justify it. We build from the start so the move is an operations task, not a new project.

// read_next

Got data that cannot leave the building?

Tell us what it is and how much you run. We tell you whether your own server is right, which model fits, and what machine it takes. Honestly, even if the answer is that an EU API is enough.