Model monitoring, changes and renewed review

Concept Drift in Production: When an AI Model Needs to Be Reviewed Again

Concept drift here means a change in the relationship between an AI model’s inputs and the outcome to be assessed using domain expertise. In the analysis of industrial measurements, this raises an operational question: Do the findings of the deployed model still fit the process for which it was evaluated? Different measurements alone do not answer this question. What matters is examining changes in inputs and the quality of the findings separately over time.[2][4]

System KnowledgeAuthor: Hannes SchubertPublished: Last updated:

A drift signal is a reason to investigate, not yet a diagnosis of the cause or the model’s performance.

What distinguishes concept drift from changes in input data

Gama and his co-authors refer to a change in the relationship between inputs and the target variable as “real concept drift”. It can occur with or without a change in the distribution of inputs. This article uses that narrower meaning.[2]

By contrast, Section 5.11.9.1 of ISO/IEC 22989 describes data drift and concept drift in terms of a decline in accuracy that has already occurred, distinguishing between its causes: changes in the statistical characteristics of operational data or a shift in the decision boundary. An observed input shift alone does not establish the performance problems described there.[1]

In the following hypothetical example, an unchanged model examines temperature windows from a defined holding phase. It was trained on cases assessed as normal using domain expertise and flags deviations from the learned pattern of normal behaviour. Such training data are not entirely unlabelled: Selecting cases as normal already constitutes an assessment.[3]

For subsequent evaluation, the model’s findings are compared with independent domain assessments of the windows. Applying this approach to a model of normal behaviour is a proposed use case; Gama’s discussion of supervised learning does not directly demonstrate its performance. The distinction between a fixed limit, a control chart and a trained model concerns the earlier choice of method.[2][3]

The subject is therefore an analytical model for measurements. Process learning, by contrast, concerns observed process behaviour and operational changes developed from it.

Defining the reference state and observation period

To identify a subsequent change, the comparison needs a recorded starting point. For this example, it would be useful to document the model version, preprocessing, holding phase and represented product variants together with the reference period. The same information should be recorded for the later observation period. This is a proposal for the process architecture, not a checklist prescribed by the sources. In the proposed setup, this also includes defining how often the comparison is repeated and which changes trigger it – such as a new variant or changed preprocessing.

The time reference must not disappear from the comparison. For evolving data streams, Gama and his co-authors point out that mixing the temporal order affects the evaluation. A model may perform well during one part of the sequence and poorly during another.[2]

In this example, it therefore remains clear which windows come from the reference period and which from later operation. Analyses can also be separated by product variant. If preprocessing has changed in the meantime, that also needs to be considered: Otherwise, inputs prepared in different ways may be compared as if they were equivalent observations.

When a different product mix changes the signal

Suppose the known variants have different typical temperature profiles, and one of them is manufactured more frequently in the later period. The combined distribution of inputs may then shift even though the relationship relevant to the domain assessment has not changed within the variants. This is an assumption of the example, not a measured operational finding.

A distribution comparison can investigate such a change without an outcome assessment already being available for every new window. Whether the model’s findings have become less reliable is a separate question: Rabanser and his co-authors explicitly distinguish detecting a distribution shift from its effect on predictive quality. Their study provides a general basis for this distinction, but not a validated method for this temperature time series.[4]

The situation is different if a new variant or range of values was not represented in the reference cases at all. A model of previous normal behaviour may flag such windows as deviations even though they are acceptable from a domain perspective and the relationship within the previous scope of use has not changed. Whether it actually flags them depends on the model and its settings. The lack of coverage is already a reason to review the reference data and scope of use – not proof of concept drift.[3]

A new composition of inputs therefore neither automatically indicates a performance problem nor provides an all-clear. In this example, the first question would be whether only the proportions of known variants have changed or whether the model is now being used outside the previously evaluated range.

What remains uncertain without domain assessments

Assessments may arrive late, contain errors or bias, and require considerable effort. Gama and his co-authors explicitly identify these limitations of feedback. A missing assessment must therefore not silently be treated as confirmation of a model finding in the analysis.[2]

For the temperature example, one possible review question is: Should this window have been investigated under the defined domain criteria? The answer is recorded separately from the original model flag. Where it remains unresolved, the case remains identifiable as unresolved. An outstanding review is thus not turned into an apparently correct result.

Reviewing only flagged windows is not enough: This allows unjustified flags to be investigated, but cannot identify anomalies the model did not flag at all. Unflagged windows therefore also need to be included in the domain review. How these cases were selected must remain visible in the analysis.

A targeted selection of difficult cases can reveal weaknesses, but does not by itself provide an overall error rate for the entire operation. The selection would need to support that conclusion. If assessments remain incomplete or restricted to certain variants, the judgement about model quality is subject to the same limitation.

The quality of human feedback also needs to be reviewed: A domain judgement is not automatically an error-free reference outcome. Nor does recording an assessment for evaluation automatically make it training data.[2]

Assessing inputs and outcomes together

Four situations can be distinguished for operational review. The following framework is a proposal for the architecture, not a diagnostic procedure taken from a standard.

Changed inputs, no deterioration demonstrated so far: The product mix, coverage of the scope of use and reliability of the outcome assessments are reviewed. If assessments are incomplete, “no deterioration demonstrated” does not confirm unchanged quality. A change in the model’s flagging rate is also an observation to consider: It shows that something has changed, not whether the flags are correct.

Findings assessed as worse, no input shift detected: An input comparison showing no issue does not rule out a change in the relationship relevant to the domain assessment. The assessment criteria or the data used for the review may also be unsuitable. This situation alone does not identify the cause.[2]

Changed inputs and findings assessed as worse: Together, these justify further investigation, but do not yet prove that the detected distribution shift caused the deterioration. Which variants and periods are affected is part of that investigation.

Too few reliable assessments: The input signal can still be described, but the quality of the outcomes remains uncertain. Continued use then requires a decision under that uncertainty; missing feedback is no substitute for evaluation. For this situation, it is possible to define in advance who decides and which interim options are available – such as restricted use or additional domain review until reliable assessments are available.

A single unusual window does not yet establish drift over time. It may nevertheless be an important reason for investigation from a domain perspective. Whether data preparation, the scope of use or the model version is subsequently changed depends on the findings.[2]

Observe inputs and domain-assessed outcomes separately and assess them together
Own schematic of the proposed review approach. Missing assessments remain unresolved; changed inputs do not automatically lead to retraining.

Comparing old and new model versions on the same cases

If a new version is to be evaluated, comparing two aggregate values from different periods is not enough to assess the difference between the versions: Different cases may also account for the result. For this example, a shared, held-out evaluation dataset on which both versions are applied would be suitable. This comparison setup is our own proposal.

The evaluation cases must remain separate from the data used to develop and tune the new version. ISO/IEC 22989 explains this separation as a way of avoiding an overestimate of model performance.[1]

Under this proposal, the shared evaluation dataset includes cases from current operation as well as earlier cases that remain relevant to the intended use. The analysis keeps product variants and periods separate. Both versions are assessed against the same domain criteria; uncertain reference assessments remain uncertain in the comparison.

Unjustified flags and missed anomalies are considered separately. Fewer flags are not an improvement if more relevant windows consequently go undetected. Similarly, a better overall result can conceal deterioration for a rare variant.

The result of this evaluation supports the decision on further use. When a change is made, its scope and the date from which the new version applies are recorded; if operation continues unchanged, the reasoning and remaining uncertainty are retained. This makes it possible to distinguish later whether the process, the assessment or the deployed model changed.

Primary sources and further reading

ISO/IEC 22989:2022 – Information technology — Artificial intelligence — Artificial intelligence concepts and terminology. Paid standard. Original source

Gama et al.: A Survey on Concept Drift Adaptation. ACM Computing Surveys 46(4), 2014, Article 44. References follow the accompanying 44-page repository manuscript, not the final publisher pagination. Original source

Chandola, Banerjee and Kumar: Anomaly Detection: A Survey. Technical report TR 07-017; references follow the accompanying manuscript version. Original source

Rabanser, Günnemann and Lipton: Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift. NeurIPS 2019. Original source