Skip to main content
NexPatch
Glossary3 min read

Token-based pricing explained: how AI costs arise

A token is the smallest unit of text into which a language model splits text, usually a word or part of a word. With token-based pricing you pay for every token the model reads and writes. Input and output are usually priced separately, so costs grow with usage.

Collage: an open book on a keyboard with single words in red labels floating above it.

Token-based pricing explained: how AI costs arise

A token is the smallest unit of text into which a language model splits text, usually a word or part of a word. With token-based pricing you pay for every token the model reads and writes. Input and output are usually priced separately, so costs grow with usage.

Collage: an open book on a keyboard with single words in red labels floating above it.
Image: Teresa Berndtsson / Letter Word Text Taxonomy / Licenced by CC-BY 4.0

What a token is

Language models do not calculate with letters or whole sentences but with tokens. A short, common word is often a single token, while a long compound technical term is split into several. How many tokens a text produces depends on the model and on the language. A word count is therefore not enough for a cost estimate. Only a measurement with the model you actually use will do.

How token-based pricing works

With every request the provider counts two quantities: the input tokens, meaning everything sent to the model, and the output tokens, meaning the answer it generates. Both are counted separately and are often priced differently. The monthly bill is requests times input tokens times input price, plus requests times output tokens times output price.

The underestimated item is context length. The input includes not only the question but everything the model needs to see for its answer: system instructions, the conversation so far and inserted documents. All of this is counted again with every request.

Cost driverEffect on the billLever
Number of requestsgrows linearly with usagescope the use case narrowly
Context length of the inputis charged again with every requestsend only relevant passages
Length of the outputcounts per generated tokenspecify the answer format
Multi-step workflowsevery intermediate step is a separate requestcount the steps before you build
Model sizelarger models usually cost more per tokenchoose the smallest suitable model

How it differs from self hosting at fixed cost

With self hosting the model runs on your own or dedicated reserved hardware. There is no bill per token, but fixed costs for hardware, staff and operations, spread across all requests.

CriterionToken-based pricingSelf hosting at fixed cost
Cost structurevariable per tokenfixed, independent of volume
Predictabilityfluctuates with usageknown in advance
Suited tolow or highly fluctuating volumehigh, stable volume
Where data staysusually with the providerin your own infrastructure
Lead timeimmediate startsetup required

How common both routes are is shown by the 2026 business survey of the ifo Institute, the Munich based economic research institute (ifo Konjunkturumfrage): 22.5 % of AI users rely exclusively on free systems, 18.7 % run their own AI systems, and the rest buy in from outside. Where the break even point lies cannot be stated in general terms. The detailed comparison is in our article on self hosting a language model versus renting one.

Practical example from our work

We see a typical pattern with knowledge assistants. An employee's question is short, yet several relevant passages from manuals are sent along with each question so that the answer is backed by evidence. The input is then often many times the size of the actual question. If usage grows from one department to several, the bill grows at the same rate. With AI in your own data centre this cost effect disappears, because the capacity is paid for at a fixed price.

Frequently asked questions

Why does this page not state any prices per token?
Prices change frequently and differ by provider and model. Any specific figure would soon be out of date. Only a calculation with current prices and your own volumes is reliable.

How can I estimate my token consumption in advance?
Measure, using a sample of real requests, how many input and output tokens arise, including context. Multiplied by the expected monthly volume, this gives you a first order of magnitude.

Full-Stack Development

From Idea to Production in Weeks

Next.js, React, TypeScript - co-founded products with government funding support.

See Our Work

Related Articles

RETURN TO BLOG

We use cookies

We use cookies and similar technologies to enhance your browsing experience, analyze site traffic, and personalize content. You can choose which categories to accept.

Learn more in our Privacy Policy and Imprint.