Measurements, analysis and expert review
AI Data Analysis in Production: Detecting Anomalies and Reviewing Findings
AI data analysis can help identify unusual patterns in production data. This article examines measurement time series for a process variable supplied by a connected device: What data and reference criteria does the analysis need, where does it run, and how does an unusual result become a finding that has been reviewed by someone with the relevant expertise?
An anomaly does not, by itself, identify its cause or the appropriate action.
What AI data analysis examines in measurement time series
In machine learning, model parameters are adjusted using data or experience. The trained model is the result of this training; it can then process new input data and calculate an output. Training and applying the model are different activities: Analysing a new measurement time series does not automatically mean that the model continues learning as it does so.[2]
Our fictional example is the temperature time series during the holding phase of a defined process step. For each batch, this specific phase is considered as a time window. The analysis is intended to flag unusual patterns for review; it is intended neither to diagnose the cause nor to release a batch. The example describes a possible architecture, not a product offering.
This requires knowing which device took the measurements and which batch and process phase the window belongs to. Context related to time and state allows heating and holding, for example, to be distinguished before a reference criterion is chosen. The signal can also be examined statistically without this association; its meaning in the process nevertheless remains unresolved.
The transfer of measurement data also needs to be considered: Timestamps and the status of a device message help put a time series in context before analysis. A gap in the time window is initially a gap in the available data, not yet a demonstrated drop in temperature.
A fixed limit, a control chart or a trained model?
Not every analysis needs a trained model. A specified temperature limit, a statistical control chart and a model for unusual patterns answer different questions. They do not form a progression from simpler to better analysis; machine learning and statistical methods are not sharply separated alternatives either.[3][4]
| Approach | Question in the example | Limit of the conclusion |
|---|---|---|
| Specified limit | Does the temperature exceed a limit specified for this step? | Compliance with the limit explains neither the entire time series nor the cause of an exceedance. |
| Statistical control chart | Does the selected process characteristic show a signal against the specified statistical reference? | A signal is a reason to investigate; statistical control limits must not be equated with product specifications. |
| Trained anomaly model | How unusual is the window under consideration compared with the learned reference pattern? | The result depends, among other things, on training data, inputs and the decision threshold. |
A control chart plots a selected quality characteristic, or a statistic calculated from it, against samples or time. NIST describes a centre line and upper and lower control limits. Points outside the limits are not the only possible reason to investigate: A non-random pattern within the limits can also indicate that the process is not in control.[3]
For the temperature example, the first task would therefore be to specify which characteristic the control chart tracks and how comparable samples are obtained. A shared holding phase narrows the comparison, but does not, on its own, establish its suitability. The choice of control chart and its statistical assumptions must fit the structure of the data.
Chandola, Banerjee and Kumar distinguish individual anomalous observations, contextual anomalies and anomalous groups of related observations. In the last category, a sequence can be anomalous even though individual points appear unremarkable when considered separately.[4] Applied to our example, a temperature rise might be expected during heating but require review during a holding phase; whether this actually holds depends on the specific process task.
What data training and model application require
For the remainder of the example, we assume a method that learns a reference pattern from holding phases assessed as normal by people with the relevant expertise. This is a chosen training assumption, not a description of all anomaly detection: The survey distinguishes methods using data labelled as normal and anomalous, methods using training data labelled as normal, and methods without these training requirements.[4]
Selecting these windows is a task that requires relevant expertise. The subsequent release of a batch alone does not establish whether its temperature time series is a suitable normal example for the intended analysis. For our model, the process phase, data completeness and known characteristics of the time series, for example, would need to be assessed before inclusion in the training dataset.
The model must also be tested using data that has not already determined its adjustment. ISO/IEC 22989 distinguishes training, validation and test data; the test dataset remains separate from training and model tuning. For system verification, the standard identifies test data that is representative of the expected inputs.[2]
In the running example, the specified model then receives a new holding-phase window or features calculated from it. The inputs used and how they are preprocessed are part of the model configuration. Training creates or changes the model; applying it initially only calculates the result for this window.[2]
AI on-premise or externally: Where do analysis and data processing take place?
Three operating arrangements can be compared for the temperature example. The following classification is our own architectural analysis; it is neither a product offering nor an adoption of NIST's cloud deployment models. Training and subsequent model application must be considered separately for each arrangement.
A – Processing on the company's own infrastructure: In the assumed arrangement, measurement storage, data preparation, training and model application take place on the company's own infrastructure at its premises. This is what on-premise means here. It does not establish anything about remote access or additional external connections; these would need to be clarified separately.
B – Dedicated operation by a service provider: The selected windows are transferred to an environment allocated to the company at the service provider for training and model application. The data is therefore outside the company's premises, even if the company's own employees administer the environment. Who can view or change data, models and logs remains a separate operational question.
C – External analysis through an interface: Storage and preparation remain within the company; the external service receives the agreed inputs. For a training run performed there, these could be selected historical windows with their context; for an individual application, the current window or its features. What data actually leaves the company follows from this specification – not from the word “API”.
These arrangements can be combined, for example by training externally and applying the model internally. “Hybrid” here refers to this combination, not automatically to a hybrid cloud as defined by NIST. Private cloud and on-premise are not synonymous either: NIST allows a private cloud to be on or off premises and to be operated by the organisation or a third party. A model interface alone is not a separate deployment model in that classification.[1]
Clarifying changes and operational workload
The operating decision involves more than the location of the computers. For our example, it would be necessary to specify who selects training data, who versions the model and preprocessing, and who introduces changes into operation. For an external service, it is necessary to clarify which changes the company controls itself and how it learns about changes made by the provider.
ISO/IEC 22989 describes retraining as updating a trained model with different training data and identifies data drift and concept drift, among other factors, as possible reasons. This does not establish a practical rule to retrain immediately whenever a change is observed.[2] In the temperature example, the first step would be to examine whether the process, measurement or data preparation has changed and what this means for the existing reference criterion.
A model that remains unchanged while it is being applied must be distinguished from continuous learning. Continuous learning involves ongoing model updates during operation; the standard associates this with additional continuous validation. This specific description must not be applied indiscriminately to every deployed model. The standard also addresses operation and monitoring independently of this.[2]
Three questions help with operational planning:
- Who maintains data preparation, the model and the technical environment during operation?
- What computing power, storage capacity and processing time does the planned training and model application run require?
- Who reviews changes and unusual results, and what working time is allocated for this?
These questions are a planning aid, not a cost calculation. Whether the company's own infrastructure or an external service is better suited to the task cannot be decided from the term AI alone. The work being considered also includes handling the findings after calculation.
Reviewing an unusual finding
An anomaly detection method can output a score, meaning an assessment of how unusual an observation is. If this is used to flag an observation, a threshold must be specified. Chandola and colleagues describe both such scores and direct normal/anomalous labels.[4] The score is therefore not automatically a probability of a particular fault cause.
With the calculation unchanged and the rule “higher score = more unusual”, a lower threshold causes more windows to be classified as anomalous. This can identify additional genuinely anomalous windows, but can also flag additional normal windows. A higher threshold reduces the number of flags, but can miss anomalous windows. Which cases are affected must be examined using data assessed by people with the relevant expertise; the survey explicitly notes the imprecise boundary between normal and anomalous behaviour.[4]
For our holding-phase example, a review record could bring together the following information. The list is our own design proposal, not a requirement of the cited sources.
- Device, batch, process step, holding phase and time window examined;
- measurement data used and known gaps or status information;
- versions of preprocessing and the model, and the threshold used;
- calculated result and the reference criterion used for the review;
- assessment by a person with the relevant expertise, the responsible person and the resulting action.
A flag is – like any learning signal – initially a starting point for investigation. In the example, it would be necessary to distinguish between a measurement problem, an unsuitable comparison and an actual process change; the flag alone does not decide this.
For human oversight to remain effective, the person conducting the review must have access to the relevant data and scope to act. A finding expressed in words could support this review; in the arrangement chosen here, its wording replaces neither the analysis nor the decision made using relevant expertise. What matters is which finding led to which reasoned action.
Primary sources and further reading
[1] Mell, Peter; Grance, Timothy: The NIST Definition of Cloud Computing. NIST Special Publication 800-145, 2011, Section 2, pp. 2–3. Original source.
[2] ISO/IEC 22989:2022: Information technology — Artificial intelligence — Artificial intelligence concepts and terminology. In particular 3.3.5–3.3.16, 5.11.7–5.11.9 and 6.2.4–6.2.7. Paid standard. Original source.
[3] NIST/SEMATECH: e-Handbook of Statistical Methods, Sections 6.3.1, What are Control Charts?, and 6.1.6, What is Process Capability? Accessed on 7 October 2026. Original source.
[4] Chandola, Varun; Banerjee, Arindam; Kumar, Vipin: Anomaly Detection: A Survey. Technical Report TR 07-017, University of Minnesota, 2007, in particular Sections 1.2 and 2.2–2.4. Original source.