AI in your own house, in one definition.
An AI system runs in your own house when model, data and operations sit inside your own infrastructure or a named zone, and no request goes to an outside service. Everything else is an interface.
What model drift is
Your AI system worked well at launch. A few months later, the first users start to complain. Nobody has changed the code. The phenomenon has a name: model drift.
An AI model learns from data covering a specific period. It reflects the relationships contained in that data. As soon as the world the model works in changes, those relationships no longer fully hold. The model stays the same, the data does not.
The literature distinguishes two forms:
Language models used through a cloud API add a related effect: the provider swaps the model version, and the behavior changes without you having done anything. Strictly speaking this is not drift, yet the effect on your users is the same.
How drift shows up in daily work
Drift rarely arrives all at once. It is a slow loss that nobody notices until trust is gone. That is exactly what makes it treacherous: individual wrong results are first dismissed as outliers, until the business unit has gotten used to checking every result itself. From that point on, most of the value of the system is lost, even though it runs flawlessly from a technical standpoint.
In a demand forecast, drift shows up as a growing gap between the forecast and actual sales. The economic effect is considerable: well calibrated demand forecasts typically reduce inventory by 10 to 50%. Schneider Electric reduced its inventory by 10%, and Unilever improved forecast accuracy at its Hefei site by 39%. These effects only last as long as the model fits the current situation. If it drifts, safety stock and tied up capital creep back up.
In a language model with a knowledge base, drift shows up in answers that increasingly miss reality: outdated prices, superseded process descriptions, products that no longer exist.
In AI agents, it shows up in workflows that fail at interfaces because fields or status values in the ERP or CRM have changed.
Causes and countermeasures
| Cause | How you recognize it | Countermeasure |
|---|---|---|
| New products, prices or processes | Answers or forecasts refer to outdated information | Update the knowledge base continuously, define changes in the master data process as a trigger for a review |
| Changing customer behavior, for example due to seasons or a crisis | Systematic deviation in one direction, not just scatter | Model seasonality, retrain at fixed intervals and after exceptional events |
| New technical terms, forms or document templates | Rising share of documents that are not recognized or are misclassified | Add new templates to the test cases, the business unit actively reports new formats |
| Changes in upstream systems such as ERP or CRM | Missing fields or fields filled differently, workflows that break off | Monitor interfaces, coordinate changes to source systems with the team operating the AI |
| Model version change at the provider | Behavior changes without any change on your side | Pin the model version, test new versions against the baseline before rolling them out |
The most common causes are easy to pin down. Each one has a matching countermeasure.
Measuring drift: baseline, thresholds, alerts
Drift can only be detected if it is clear how good the system was at the start. That is why every project begins with a baseline: a target metric the business understands, measured over a meaningful period. For a demand forecast, this is the gap between forecast and sales at the same service level; for a language model, it is the hit rate on a fixed set of test questions.
On this basis, quality is measured continuously. Three methods of measurement complement each other:
You define thresholds for each metric. A warning threshold triggers a review, an alert threshold triggers a decision. Where these thresholds lie depends on the use case and is set together with the business unit.
It is important that the measurement results reach the right people. A report that only the development team understands does not help the business unit. Good reports show how quality develops over time, flag deviations from the baseline and state which action was taken. This keeps it clear why a system keeps running or is adjusted.
Retraining as a process, not an emergency
Retraining should be a planned activity, not a hectic reaction to a complaint. Depending on the system, retraining means different things: updating the knowledge base, adjusting prompts and rules, repeating fine tuning with new examples, or retraining a forecasting model on current data.
Four rules make the process reliable. Retraining is triggered at fixed intervals and additionally whenever a threshold is crossed. Every new version is tested against the baseline before it is rolled out. There is always a way back to the last working version. A named person decides whether a model keeps running, is retrained or is switched off.
Data protection and documentation are part of this as well. If you retrain with new data, you should record which data was used, on what legal basis and with which approval. This documentation makes changes to the model auditable, for your IT as well as for your data protection officer.
Why monitoring belongs to operations
Many AI projects end with a handover. Quality monitoring often falls between the cracks: the project team has moved on, and IT monitors servers, not answers. For operations, we set a target of 99.9% availability. Availability alone, however, says nothing about quality. A system can be reachable and still deliver wrong results.
That is why, at NexPatch, quality monitoring is part of operations and not of the project. We describe how quality, retraining and responsibilities are governed by contract under Service level.
An AI model does not age. The world around it does.
Frequently asked questions
Sources
- NexPatch AI: "Glossar Bedarfsprognose" (Glossary: demand forecasting), nexpatch.ai/de/blog/bedarfsprognose (typical inventory reduction of 10 to 50%, Schneider Electric inventory down 10%, Unilever Hefei forecast accuracy up 39%).