Skip to main content
NexPatch
Knowledge · Comparison

Run a language model yourself or rent it?

Whether self-hosting a language model or paying per request makes more sense comes down to request volume. At low volume, pay-per-request is usually cheaper, since there's no fixed cost for hardware and staff. At high, sustained volume, that flips, because self-hosting drives down the ongoing cost per request.

Self-hosted
API
from 100,000 requests a month

Two cost models with different logic

Whether a company uses AI through a cloud API or runs models on its own GPUs is often decided by gut feeling. Yet both options follow a clear logic that can be calculated.

With pricing per token, you pay for every request. Inputs and outputs are measured in tokens, meaning fragments of words, and billed at a fixed price per million tokens. For getting started, this model is ideal: no investment, no operations, no waiting for hardware. You can start the same day. With every new use case, however, the bill grows linearly. An assistant in sales, a document review in purchasing and an agent in customer service add up month after month.

When you run AI on your own GPUs, the logic is reversed. Hardware, power, cooling, maintenance and staff cause fixed costs, regardless of how heavily the system is used. Each additional request, on the other hand, costs almost nothing as long as the existing capacity is sufficient. Running models yourself rewards utilization; pricing per token rewards restraint.

This leads to the central question of every cost comparison: how many tokens does your company consume today, and how many will it be in one to two years?

Where the cost lines cross

If you plot both models against monthly volume, you get a rising straight line for pricing per token and an almost flat line for running models on your own GPUs. The intersection is the break even point. Our analysis in the article "Das lokale KI Betriebssystem" (The local AI operating system) arrives at the following thresholds:

•
Compared with premium APIs, meaning the most powerful models from large cloud providers, running AI on your own GPUs pays off from about 5 to 10 million tokens per month.
•
Compared with budget APIs with significantly lower prices per token, the threshold is about 50 to 100 million tokens per month.

Looking at the full term is revealing. Calculated over 36 months, running models on your own GPUs comes to about USD 7.15 per million tokens, while using OpenAI comes to about USD 6.90. The two options are therefore surprisingly close. The price per token alone rarely settles a decision. What tips the balance is volume, utilization, data protection and the question of who is responsible for operations.

For your own situation, it is worth running the numbers with real consumption data. Our cost comparison calculator shows you where your company sits on the curve.

Comparison table: own GPUs vs pricing per token

CriterionRunning AI on your own GPUsPricing per token via cloud API
Cost structureMostly fixed: hardware, power, maintenance, staffVariable, grows linearly with every request
Startup effortProcurement, setup, security hardeningLow, usable immediately
Cost per additional requestAlmost zero up to the capacity limitConstant per million tokens
Economical fromAbout 5 to 10 million tokens per month compared with premium APIs, about 50 to 100 million compared with budget APIsBelow these thresholds
Data controlData never leaves your own networkDepends on provider, contract and jurisdiction
Operational responsibilityInternally or with an operations partnerWith the API provider
Adaptation to domain languageFine tuning and custom models possibleLimited, depending on the provider
PredictabilityHigh, costs known over the full termFluctuates with usage and price changes
DependencyOn hardware and operational expertiseOn the provider\'s prices, model versions and availability

The items missing from the calculation

Many cost comparisons set the price of a GPU against the price per token and stop there. That leaves out exactly the items that make the difference in daily use. Who monitors the system and notices when answers get worse? Who installs updates for drivers, runtime environment and models, and tests them beforehand? Who responds at night or on the weekend when the service is down? Who maintains permissions, logs and documentation for the next audit?

On premise operation moves operational risk into your own company. That is no argument against running AI on your own GPUs, though it is an argument for a complete calculation. Besides hardware, the fixed costs include staff time for operations and on call duty, redundancy or spare parts, energy and cooling, security updates and the maintenance of the applications built on top of the model. Anyone who leaves these items out is comparing an incomplete bill with a complete one.

On the other side, many comparisons also leave out items: data protection reviews for every new use case, contracts with subcontractors and the risk that prices or model versions change. Then there is the legal situation. The US CLOUD Act applies regardless of server location. A European data center operated by a US provider therefore does not automatically protect against access by US authorities.

For companies without their own operations team, there is a middle path. The hardware sits in your own data center or with a German hosting provider, and operations are contractually assigned to a partner. Your data stays within the company without your IT having to be on call at night. We describe how we take over this part under AI operations.

Utilization and software count too

Whether running AI on your own GPUs is economical depends not only on the hardware but on how many requests it processes in a given time. The software used to serve the models plays a bigger role than many expect. A recent scientific measurement shows around 800 tokens per second with vLLM compared with around 150 tokens per second with Ollama at ten concurrent users.

For the cost calculation, this means the same GPU can carry several times the load depending on the software. A tool that is convenient for testing on a single workstation is not automatically suitable for production use with many simultaneous users. Anyone calculating the cost of running models on their own GPUs should therefore not just count GPUs but plan the entire serving setup: runtime environment, queues, load balancing and a gateway that routes requests to suitable models.

A gateway like this also enables hybrid routing. Simple requests run on a small local model, confidential data always stays within your own network, and only selected tasks involving public data go to an external API when needed. This way you use the strengths of both cost models without committing permanently to either one.

Our recommendation in three tiers

From the thresholds and our project experience, we derive a simple tiered approach:

1
Below 5 million tokens per month: Start in the cloud, provided your data allows it. Use the time to measure consumption, use cases and quality.
2
From 10 million tokens per month or with regulated data: Switching to your own GPUs pays off. With personal, confidential or regulated data, it can be the right choice even below this threshold.
3
From 50 million tokens per month: Plan for multiple GPUs, load balancing and hybrid routing so that capacity and costs grow with your use cases.

Between 5 and 10 million tokens lies a gray zone in which the nature of your data, expected growth and internal expertise decide. What matters is not to make the decision once but to monitor consumption continuously. Many companies underestimate how quickly demand rises as soon as the first use case is in production and other departments follow.

The cost question is rarely decided by the price per token. It is decided by volume and by who carries the responsibility for operations.

Frequently asked questions

Sources

  1. NexPatch AI: "Das lokale KI Betriebssystem" (The local AI operating system), https://nexpatch.ai/de/blog/lokales-ki-betriebssystem-eigene-server-infrastruktur (break even compared with premium and budget APIs, 36 month TCO)
  2. arXiv 2511.17593, throughput measurement of vLLM and Ollama at ten concurrent users, https://arxiv.org/abs/2511.17593
  3. US CLOUD Act (Clarifying Lawful Overseas Use of Data Act)

We use cookies

We use cookies and similar technologies to enhance your browsing experience, analyze site traffic, and personalize content. You can choose which categories to accept.

Learn more in our Privacy Policy and Imprint.