From pilot into operations, in six steps.
Why pilots stall, data upkeep, integration, access and logging, cost, one accountable owner. The guide that turns an experiment into a system.
Why AI projects often cannot prove their value
Almost every AI project is asked after a year what it has delivered. Almost none can answer with a number that everyone involved accepts. The model is rarely the reason. The reason is that before the start, nobody recorded how well the process performed without AI.
The market situation makes the problem worse. Analyses by Deloitte and by Capgemini and Microsoft from 2025 show that only 6% of all AI initiatives pay for themselves in less than a year. Individual use cases typically need 3 to 12 months to deliver a return, large platforms 2 to 4 years. A 2025 MIT study concludes that 95% of generative AI pilots show no measurable return. The study's methodology is limited, so the figure should be read as a direction rather than an exact value. Still, the direction is clear: many projects cannot prove their value, even if that value may well exist.
Without a baseline, every attempt to measure success becomes a matter of opinion. The business unit feels relief, controlling mainly sees costs, and IT sees additional operational effort. Decisions to expand or shut down are then made on gut feeling or on next year's budget.
What a baseline is and what it is not
A baseline, sometimes also called a reference value, describes the measurable state of a process before AI is introduced. It consists of three parts: the metric itself, the period over which it was measured, and the conditions under which the value was produced.
A baseline is not a model metric. The accuracy of a model or the hit rate of a search are technical measures. They help the development team, but say little about whether a process has improved. A forecasting model can become more accurate without any less capital being tied up in the warehouse. A language model can phrase better answers without a case being closed any faster.
Nor is a baseline an estimate from memory. Statements such as "processing takes about half a day" are a start, but not a measurement. The value only becomes reliable when it comes from system data, from time tracking or from a structured sample.
Four steps to a reliable baseline
For every initiative, we follow four steps before the first line of code is written.
1. Choose a metric the business understands. Suitable options are capital tied up, cycle time, error rate or processing time. The metric should be expressible in euros, hours or cases so that management and the business unit speak the same language. Choose one primary metric and no more than two secondary metrics. Anything more dilutes the evaluation.
2. Measure the current state over a meaningful period. A single day or a particularly quiet week distorts the picture. The period should cover seasonal fluctuations, month end closing or typical peaks, insofar as they are relevant to the process.
3. Record the conditions. To keep the comparison fair later on, document what must not change and what is likely to change: service level, order volume, headcount, product range. If volume rises significantly after launch, the comparison has to take that into account.
4. Agree jointly on the improvement at which the project counts as a success. The business unit, IT and management agree on this threshold before the start. The counter threshold is just as important: below which value will the system be adjusted or shut down?
The effort involved is manageable. A workshop and a clean data extract from the ERP system, the ticketing system or the document repository are often enough.
Metrics by use case
| Use case | Metric | Condition for the comparison |
|---|---|---|
| Demand forecasting | Capital tied up in inventory | Same service level, comparable product range |
| Document processing, such as incoming invoices | Processing time per document, share of manual rework | Document volume, share of new suppliers |
| Internal knowledge assistant | Time to a reliable answer, share of test questions answered correctly | Fixed catalog of test questions, same document base |
| Customer service | Handling time per inquiry, share of cases resolved without follow up questions | Inquiry volume, channel, season |
| Quote preparation | Cycle time from inquiry to quote | Number and complexity of inquiries |
| Quality inspection in production | Defect escape rate, inspection time per part | Product mix, inspection specification |
Which metric fits depends on the use case. The following overview shows metrics that suit typical use cases, along with the condition you should record for a fair comparison.
Demand forecasting shows particularly clearly why the choice of metric is decisive. The reliable measure is not model accuracy but the difference in capital tied up at the same service level. Well configured demand forecasts typically reduce inventory by 10 to 50%, and Schneider Electric reduced its inventory by 10%. Such figures are only comparable if the service level stays constant. If you cut inventory at the cost of delivery capability, you have not saved capital, you have shifted a problem.
Common mistakes in measuring success
We encounter the same patterns in projects again and again:
From the baseline to the decision on operations
A baseline is not an end in itself. It serves three decisions in the lifecycle of an AI system.
After the pilot, the first measurement shows whether the use case holds up. An initial internal pilot in 72 hours does not yet deliver a reliable long term value, but it does give a first direction against the baseline. On the way to production use, for which we plan 4 to 10 weeks, measurement is extended to the real volume. During ongoing operations, the comparison shows whether quality remains stable or whether data and processes have changed so much that retraining is needed.
That is why, for us, the baseline belongs at the beginning and not at the end. We set it together with you before every pilot and carry it over into the reports of our operations service. This way, the business unit, IT and management are talking about the same number a year later. We describe how we bill for pilot, implementation and operations on our Pricing page.
Frequently asked questions
Sources
- Deloitte (2025) and Capgemini and Microsoft (2025), cited in nexpatch.ai/de/blog/orpheon-vs-plattformen (6% of AI initiatives pay for themselves in under a year, individual use cases in 3 to 12 months, large platforms in 2 to 4 years, 80% of initiatives require upfront investment in data architecture).
- MIT (2025): study on generative AI in enterprises (95% of pilots without measurable return). The methodology is limited; the figure should be read as a trend.
- NexPatch AI: "Bedarfsprognose" (Demand forecasting), nexpatch.ai/de/blog/bedarfsprognose (inventory reduction typically 10 to 50%, Schneider Electric inventory down 10%).