Reference intervals, flags and detection quality
Anomaly Detection in Production: When Does an Anomalous Pattern Count as Detected?
Anomaly detection searches data for observations or patterns that deviate from expected behaviour. In production, this can concern a single measurement, but also a continuous pattern over time. Whether such an anomaly counts as detected depends on what the analysis is meant to flag: individual time points, an entire window or an interval with a beginning and an end. A detection somewhere within the pattern does not yet tell us how completely it has been captured.[1][2]
Before detections are counted, it must be clear what counts as one case and how the flags are assigned to that case.
What anomaly detection detects – and what does not follow from it
An anomaly is initially a statement about data in relation to an expectation. Which deviation matters depends on the task; the boundary between normal and anomalous behaviour is not always clear. A flagged temperature pattern therefore does not, on its own, establish either a product defect or its cause.[1]
In the following hypothetical example, temperature is examined during a defined holding phase of a production unit. The product variant and the phase under consideration remain the same for the comparison. What it is reasonable to expect from the pattern also depends on the state of the process: a temperature change during heating is assessed differently from one during a holding phase.
The example assumes a method whose reference has been derived from time series selected as normal for this task. This is an assumption of the example, not a property of every anomaly detection method. The available data and assessments differ between methods and applications.[1]
A single value, a measurement vector or a continuous pattern over time?
An analysis can assign a single judgement to each completed holding phase. The window is then either flagged or not flagged. This does not yet show which part of its pattern was anomalous. A different method may instead flag individual time points or intervals on the timeline. Only this more detailed output allows us to ask whether the beginning, duration and end of an anomalous interval have been captured.
The number of measured variables must also be distinguished from this. A data point can contain temperature alone or several variables considered at the same time. A vector of several measured variables is multivariate; that does not make it a time interval. The temporal unit and the number of features are two different properties of the analysis.[1]
Chandola, Banerjee and Kumar distinguish anomalous individual observations, contextual anomalies and anomalies in a group of related observations; context dependence can apply to individual values as well as groups.[1] Tatbul and co-authors classify anomalies that extend over a continuous time interval as a subset of contextual and collective anomalies.[2] For the temperature example, the decisive question is therefore whether the anomaly being sought is a single value or behaviour over a period of time.
When an event counts as detected
For the example, a continuous interval within the holding phase is assessed as a relevant anomaly by someone with the relevant expertise. It serves as the reference for the comparison. The rest of the pattern shown is also assumed to have been reviewed and assessed as unremarkable for this task. The example therefore contains no unassessed intervals; a real trial may well contain them.
One possible output flags only a short interval within the reference interval. It has therefore detected the anomaly as such. Whether this partial detection is sufficient for the task is a different question. If the duration of the anomaly is to be captured, some of the required information is missing.[2]
A second output covers almost the entire interval. A third flags several separate segments within it. If only the number of flagged measurement points is considered, the difference between a continuous and a fragmented output can be lost. For a task that is meant to capture an event as one coherent occurrence, that difference may be important.[2]
An output may additionally flag an interval outside the reference interval. Because this was assessed as unremarkable in the example, the flag is unwarranted for the defined task. The fact that the same output covers the relevant interval well does not turn the additional flag into a correct detection.
Tatbul and co-authors describe an evaluation approach for such time ranges. It distinguishes whether an anomaly is detected at all, how much of it is covered, where the overlap lies within the interval and whether several separate flags overlap the same interval. The weighting can be adapted to the task. This is a research approach, not a general requirement that a particular degree of partial coverage must always be sufficient.[2]
In a conventional count, precision and recall refer to two different sets: precision is the proportion of correct positive outputs among all positive outputs; recall is the proportion of detected positive reference cases among all positive reference cases. Here, the positive class is the anomaly defined for the comparison – not an established product defect. The unit must remain the same for this count: measurement points, windows or events matched according to a defined rule.[2]
For partially overlapping time ranges, simply counting flagged points is not enough to describe the quality of event detection. A metric must make clear how partial coverage and multiple flags were handled. If different units or matching rules are used, metrics with the same name are not directly comparable.[2]
The position of a flag remains separate from its availability. A method may flag an early interval only after the entire window has ended. That does not imply an early notification during the ongoing process. Whether the finding is available in time for domain review depends on the intended workflow and when it is actually made available.
What should be defined before the first comparison
The boundaries of the reference interval deserve the same attention as the flags. If the transition between normal and anomalous behaviour is unclear, a slightly different choice of beginning or end can change the overlap. In the example, the reference established through domain assessment therefore also includes the criteria used to draw those boundaries.[1]
Five points can be derived from this for such a comparison:
- The unit: What is being assessed: a measurement point, a window or a continuous interval?
- The matching rule: When does a flag belong to a reference interval, and how are overlaps handled?
- The coverage: Is finding the event sufficient, or do the captured duration and position also count?
- The fragmentation: How are multiple flags for the same interval and flags spanning several intervals taken into account?
- The assessment basis: Which intervals have been assessed using domain expertise, and which remain unresolved?
Partitioning a time series can also change the evaluation task. If a continuous event is cut at a partition boundary, the resulting segments have different boundaries. In their experimental setup, Tatbul and co-authors partitioned the data for training and testing so that anomaly ranges remained intact despite the segmentation. This is the procedure they describe, not a general rule for every data split. For a given trial, it should remain clear whether segments still belong to the same event.[2]
Which data are used for model development and which are held back for testing is part of defining the training and test data. The event matching described here adds to that separation: it records the unit to which the test result refers.
For intervals without a domain assessment, it is not possible to decide whether a flag is correct or an anomaly has been missed. Such intervals must not be silently treated as unremarkable reference data in the evaluation. The comparison result applies to the set that has actually been assessed.
Three things therefore remain distinguishable in the evaluation: the reference interval established through domain assessment, the method’s output and the rule used to match the two. If this rule changes, the metric may change even though the same data and the same flags are present.[2]
Successfully completing a detection task does not yet determine what action follows in the plant. A specific unusual finding requires domain review. The comparison first establishes what the method detected under the defined conditions – and what its metrics actually count.
Primary sources and further reading
[1] Chandola, V.; Banerjee, A.; Kumar, V. (2007): Anomaly Detection: A Survey. University of Minnesota, TR 07-017, particularly Sections 1.1–1.2 and 2.1–2.3. The repository manuscript was used, not the typeset ACM version. Original source.
[2] Tatbul, N.; Lee, T. J.; Zdonik, S.; Alam, M.; Gottschlich, J. (2018): Precision and Recall for Time Series. Advances in Neural Information Processing Systems 31, particularly Sections 1, 2, 4 and 5.1. Original source.