Skip to main content
NexPatch
Knowledge · Guide

From pilot project into operations.

Most AI pilots don't fail on the technology. They fail on what comes after: data that isn't mature enough, no fit into daily work, and no clear ownership once the pilot wraps up. Closing those three gaps deliberately before the pilot ends is what reliably turns it into stable, monitored operations rather than an unused demo.

Stall
Data Care
Integration
Access
Costs
Accountability

Why classic SLAs are not enough for AI systems

What happens when your AI system gives wrong answers on a Friday evening? If your contract has no answer to that, you do not have a service level, you have a hope.

In classic IT, service level agreements, or SLAs, have been standard for years. They govern availability, response times and maintenance windows. For a mail server or an ERP system that is largely sufficient, because these systems behave deterministically: if the service is running, it delivers the expected result.

AI systems behave differently. A language model can be reachable, respond quickly and still deliver content that is factually wrong. A forecasting model can run on time every morning and still drift further and further away from actual sales, because the product range, prices or customer behavior have changed. This phenomenon is called model drift. It occurs without a single change to the code.

An SLA that only governs availability therefore fails to cover exactly the kind of failure that is typical for AI: the system is running, yet it is not delivering.

The components of an AI service level

ComponentWhat is agreedHow you can verify it
AvailabilityOperating target, measurement method, maintenance windowsRegular report on availability
QualityTarget metric, baseline, thresholds for review and actionOngoing measurement against a fixed test catalog
Response by severity levelDefinition of the levels, reporting channel, contact person, escalationLog of every report with a timestamp
CostsUpper limits, alerts for unusual consumptionConsumption per use case in the report
RetrainingTriggers, regular review cycle, approvalDocumented training runs with results before and after
UpdatesMandatory testing before rollout, path back to the previous versionChange log with test results
ReportsContent, frequency, recipients in the business unit and ITA report the business unit can read as well
CooperationData access, contact persons, approvalsNamed people on both sides
ExitHandover of data, models and configurationAgreed format and handover plan

A complete service level for AI consists of several components. Not every use case needs all of them in the same depth. None of them should be missing, though.

For availability, we set 99.9% as the operating target. What matters is that the contract states how this figure is measured: at which point, over which period and with which exceptions for planned maintenance. An availability figure without a measurement method can neither be verified nor enforced.

The cost component is just as important. Unlike classic software, the running costs of an AI system depend directly on usage, especially with pricing per token. A new use case, a faulty loop in an agent or unexpectedly high demand can cause consumption to jump sharply. Agree on upper limits per use case and an alert before these limits are reached. The report should show which use case causes which costs.

Reports themselves are often underestimated. A report that only shows server uptime and response times helps IT, not the business unit. A good report answers three questions in language that management and the business unit understand as well: Did the system run as agreed? Has quality changed compared with the baseline? Which actions are next?

Making quality measurable: baseline and thresholds

The most important difference from a classic SLA is the quality component. It requires that a baseline was set before launch: How good was the process without AI, and how good was the system at acceptance?

For language models, a fixed catalog of test questions with verified answers from your business unit works well. For forecasting models, the deviation between forecast and actual value works well, assessed against the business target metric. In demand forecasting, that is the difference in tied up capital at the same service level.

On this basis, you agree on two thresholds. The first triggers a review: the operations team analyzes the cause and informs the business unit. The second triggers an action, such as retraining, an adjustment to the knowledge base or, in extreme cases, a temporary shutdown. Who decides on a shutdown should be specified by name, not just as a role.

Three severity levels and what must be defined for each

Severity levelTypical exampleWhat must be defined
1: The system is downApplication not reachable, interface to the ERP system has failedReporting channel outside business hours as well, named contact person, response time, notification of the business unit and IT, fallback process for the duration of the outage
2: Noticeably wrong resultsAnswers deviate systematically from the baseline, forecasts are clearly offWho determines the loss of quality, response time, right to a temporary shutdown, root cause analysis, decision on retraining or a return to the previous version
3: A single case behaves unexpectedlyA single answer is inappropriate, a rare document is read incorrectlyReport with an example, response time, addition to the test catalog, bundling into the next regular improvement

Three severity levels are enough in most cases. More levels make classification harder in an emergency, fewer levels blur the line between an outage and a quality problem.

Which response time is appropriate for which level depends on the use case. An assistant for internal research has different requirements than a system that answers customer inquiries or triggers orders. What matters is that every level has a response time, a contact person and an escalation path, and that both sides use the same definition. We set these values together with you based on the specific use case.

A joint test run also helps. Before launch, walk through one case per severity level and check whether the reporting channel, contact person and escalation work in practice. Gaps become visible before they cost time in an emergency.

Retraining, updates and the way back

Retraining belongs in the SLA as a planned process, not as an emergency measure. Three questions need to be answered: What triggers retraining, for example exceeding a quality threshold or a new product range? How often is it regularly checked whether retraining is needed? Who approves the new model before it goes into production?

The same logic applies to updates of models and software. A new base model can be better overall and still answer worse at exactly the point your process depends on. That is why the SLA should include mandatory testing against the agreed test catalog before every rollout and a documented path back to the previous version. Updates without this path back are a risk you should not carry.

This component also covers the question of what happens at the end of the collaboration. You should be able to take your data, adapted models, prompts and configurations with you in an open format. We describe what matters here under Exit.

The other side: what your company contributes

A service level does not run in only one direction. It only works when it is clear what the company itself contributes. This usually includes:

•
Access to the data needed for monitoring and retraining
•
a named contact person in the business unit who can assess quality issues on a subject matter level
•
a contact person in IT for interfaces, permissions and infrastructure
•
timely approvals when retraining or updates are due
•
reports of anomalies through the agreed channel

Without these contributions, even the best operator cannot keep its commitments. A sound SLA therefore names obligations on both sides. Please coordinate the contractual details of your specific case with your legal department or your law firm. This guide describes the substantive content and does not replace a legal review.

With us, these components are part of operations and not an optional extra. We show how we set up service levels in practice on the Service Level page.

Frequently asked questions

We use cookies

We use cookies and similar technologies to enhance your browsing experience, analyze site traffic, and personalize content. You can choose which categories to accept.

Learn more in our Privacy Policy and Imprint.