Skip to main content
NexPatch
Knowledge · Comparison · Self-hosting vs API

Run a language model yourself or rent it?

Whether self-hosting a language model or paying per request makes more sense comes down to request volume. At low volume, pay-per-request is usually cheaper, since there's no fixed cost for hardware and staff. At high, sustained volume, that flips, because self-hosting drives down the ongoing cost per request.

Self-hosted
API
from 100,000 requests a month

The question that decides the purchase

Behind the question "run it yourself or rent it" is really a simple calculation: how many requests does your company process over what period, and how stable is that volume? With occasional, fluctuating demand, pay-per-request through a hosted interface usually wins, because there's no fixed cost for unused capacity. With high, steady volume, that calculation tips toward self-hosting, because the fixed costs of hardware and staff spread across more requests and the cost per request falls.

This page answers the question in prose, for anyone who wants to understand the mechanics before plugging in their own numbers. If you'd rather run your own figures directly, the cost-comparison calculator answers the same question for your specific situation.

Two examples show how differently the answer can land depending on the situation. A company with a small, internal use case and a few hundred requests a month is almost always better off with pay-per-request, because hardware and staff for self-hosting barely pay for themselves. A company with a customer-facing application and high, steady request volume all year round, on the other hand, often reaches a point where self-hosting on private AI infrastructure becomes both the cheaper and the more controlled option.

What self-hosting actually costs

Self-hosting means a company runs or reserves dedicated infrastructure for a language model itself, rather than paying an external provider per request. The breakdown below is our own model calculation with disclosed assumptions, not a universal figure - actual costs vary by model size, provider and contract terms.

Cost blockWhat it's forNote
Hardware or dedicated capacityCompute to run the modelFixed cost, independent of actual request volume
StaffSetup, monitoring, tuningOngoing effort, continuing after rollout
OperationsMonitoring, maintenance, securityRecurring effort for stable, ongoing operation

The assumption behind this model calculation is that self-hosting becomes economically worthwhile once request volume is high and stable over time, so the fixed costs of hardware and staff spread across many requests. At low or heavily fluctuating volume, too few requests carry the same fixed costs, so cost per request rises. A detailed breakdown of pricing models is available on the pricing page, and how private AI infrastructure specifically enables this kind of self-hosting is covered on the related solutions page.

What pay-per-request actually costs

With pay-per-request, a company pays only for requests actually used at an external provider, without keeping its own hardware or dedicated staff for running the model. The advantage is that no fixed costs accrue when volume is low or irregular. The downside shows up as volume grows: because every single request is billed, total cost scales linearly with usage, without the cost-per-request reduction self-hosting gets from spreading fixed costs.

There's also an aspect beyond pure cost: with pay-per-request, data typically leaves your own infrastructure and runs through the external provider's systems. For companies with strict requirements on data residency and traceability, that's an additional factor that feeds into the decision beyond cost alone.

Where the break-even point typically sits

There's no honest, fixed percentage or request count above which self-hosting pays off across the board - too many factors vary, including model size, contract terms and staff costs. As a general tendency, though: the higher and more stable request volume is over time, the more the advantage shifts from pay-per-request toward self-hosting, because self-hosting's fixed costs spread across more and more requests.

Rather than a blanket percentage, an individual calculation using your company's actual figures is the more reliable basis for a decision. The cost-comparison calculator runs exactly this calculation based on your own assumptions about volume, model size and time period, and shows which side of the break-even point your company currently sits on.

The comparison at a glance

The table below summarises the key differences between the two models, as a starting point for the conversation inside your own company.

CriterionSelf-hostingPay-per-request
Cost structureFixed costs for hardware and staffVariable cost per request used
Best suited toHigh, stable request volumeLow or heavily fluctuating volume
Data residencyInside your own, private infrastructureTypically on the external provider's infrastructure
Lead timeRequires building out infrastructureImmediate start, no infrastructure of your own needed

Other factors beyond pure cost

Beyond the pure cost calculation, other factors feed into the decision. One is how much independence from a single external provider matters to your company. Another is regulatory requirements on data residency and traceability, which in some industries tip the decision toward self-hosting regardless of cost. How predictable your own request volume is over the coming years also belongs in this weighing, since switching from pay-per-request to self-hosting is possible but takes additional lead time.

Some companies combine both models - running a base volume on self-hosted private AI infrastructure, for example, and absorbing load spikes through pay-per-request. That kind of hybrid approach pays off especially when request volume fluctuates strongly across the year but a reliable baseline demand is clearly identifiable.

How the decision plays out in practice

In practice, the decision rarely starts with cost alone. First comes the question of which use case is actually planned and how it's likely to develop over the coming years. Only after that does it make sense to look at expected request volume, followed by how much data residency and independence from a single provider matter to your company. Working through these three steps in this order usually leads to a more robust decision than looking only at the current per-request price.

For companies already running a pilot on private AI infrastructure, it's also worth reading the AI implementation timeline article, since cost and timeline planning tend to influence each other in practice.

Frequently asked questions