Skip to main content
NexPatch
Knowledge · Glossary

AI in your own house, in one definition.

An AI system runs in your own house when model, data and operations sit inside your own infrastructure or a named zone, and no request goes to an outside service. Everything else is an interface.

Model
Data
Operations

What model drift is

Your AI system worked well at launch. A few months later, the first users start to complain. Nobody has changed the code. The phenomenon has a name: model drift.

An AI model learns from data covering a specific period. It reflects the relationships contained in that data. As soon as the world the model works in changes, those relationships no longer fully hold. The model stays the same, the data does not.

The literature distinguishes two forms:

•
Data drift: The input data changes. A document recognition model suddenly receives forms in a new layout, a forecasting model sees order quantities it does not know from training.
•
Concept drift: The relationship between the input and the correct result changes. After a price change, the same customer inquiry calls for a different answer than before.

Language models used through a cloud API add a related effect: the provider swaps the model version, and the behavior changes without you having done anything. Strictly speaking this is not drift, yet the effect on your users is the same.

How drift shows up in daily work

Drift rarely arrives all at once. It is a slow loss that nobody notices until trust is gone. That is exactly what makes it treacherous: individual wrong results are first dismissed as outliers, until the business unit has gotten used to checking every result itself. From that point on, most of the value of the system is lost, even though it runs flawlessly from a technical standpoint.

In a demand forecast, drift shows up as a growing gap between the forecast and actual sales. The economic effect is considerable: well calibrated demand forecasts typically reduce inventory by 10 to 50%. Schneider Electric reduced its inventory by 10%, and Unilever improved forecast accuracy at its Hefei site by 39%. These effects only last as long as the model fits the current situation. If it drifts, safety stock and tied up capital creep back up.

In a language model with a knowledge base, drift shows up in answers that increasingly miss reality: outdated prices, superseded process descriptions, products that no longer exist.

In AI agents, it shows up in workflows that fail at interfaces because fields or status values in the ERP or CRM have changed.

Causes and countermeasures

CauseHow you recognize itCountermeasure
New products, prices or processesAnswers or forecasts refer to outdated informationUpdate the knowledge base continuously, define changes in the master data process as a trigger for a review
Changing customer behavior, for example due to seasons or a crisisSystematic deviation in one direction, not just scatterModel seasonality, retrain at fixed intervals and after exceptional events
New technical terms, forms or document templatesRising share of documents that are not recognized or are misclassifiedAdd new templates to the test cases, the business unit actively reports new formats
Changes in upstream systems such as ERP or CRMMissing fields or fields filled differently, workflows that break offMonitor interfaces, coordinate changes to source systems with the team operating the AI
Model version change at the providerBehavior changes without any change on your sidePin the model version, test new versions against the baseline before rolling them out

The most common causes are easy to pin down. Each one has a matching countermeasure.

Measuring drift: baseline, thresholds, alerts

Drift can only be detected if it is clear how good the system was at the start. That is why every project begins with a baseline: a target metric the business understands, measured over a meaningful period. For a demand forecast, this is the gap between forecast and sales at the same service level; for a language model, it is the hit rate on a fixed set of test questions.

On this basis, quality is measured continuously. Three methods of measurement complement each other:

•
Automated tests against a fixed evaluation dataset, regularly and after every change
•
Statistical comparisons of the input data with the training data, to spot data drift early
•
Spot checks and feedback from the business unit, to capture what tests do not cover

You define thresholds for each metric. A warning threshold triggers a review, an alert threshold triggers a decision. Where these thresholds lie depends on the use case and is set together with the business unit.

It is important that the measurement results reach the right people. A report that only the development team understands does not help the business unit. Good reports show how quality develops over time, flag deviations from the baseline and state which action was taken. This keeps it clear why a system keeps running or is adjusted.

Retraining as a process, not an emergency

Retraining should be a planned activity, not a hectic reaction to a complaint. Depending on the system, retraining means different things: updating the knowledge base, adjusting prompts and rules, repeating fine tuning with new examples, or retraining a forecasting model on current data.

Four rules make the process reliable. Retraining is triggered at fixed intervals and additionally whenever a threshold is crossed. Every new version is tested against the baseline before it is rolled out. There is always a way back to the last working version. A named person decides whether a model keeps running, is retrained or is switched off.

Data protection and documentation are part of this as well. If you retrain with new data, you should record which data was used, on what legal basis and with which approval. This documentation makes changes to the model auditable, for your IT as well as for your data protection officer.

Why monitoring belongs to operations

Many AI projects end with a handover. Quality monitoring often falls between the cracks: the project team has moved on, and IT monitors servers, not answers. For operations, we set a target of 99.9% availability. Availability alone, however, says nothing about quality. A system can be reachable and still deliver wrong results.

That is why, at NexPatch, quality monitoring is part of operations and not of the project. We describe how quality, retraining and responsibilities are governed by contract under Service level.

An AI model does not age. The world around it does.

Frequently asked questions

Sources

  1. NexPatch AI: "Glossar Bedarfsprognose" (Glossary: demand forecasting), nexpatch.ai/de/blog/bedarfsprognose (typical inventory reduction of 10 to 50%, Schneider Electric inventory down 10%, Unilever Hefei forecast accuracy up 39%).

We use cookies

We use cookies and similar technologies to enhance your browsing experience, analyze site traffic, and personalize content. You can choose which categories to accept.

Learn more in our Privacy Policy and Imprint.