Model application, execution time and timely results

AI Inference in Production: How Measurement Data Becomes a Timely Result

AI inference means applying a trained model to input data to produce a result. During training, the model is built or adapted using data; during inference, it is applied. For the analysis of measurements in production, this raises a practical question: When is the result available for the next work step?[2]

System KnowledgeAuthor: Hannes SchubertPublished: Last updated:

What matters is when a usable finding becomes available in the workflow; the model’s computation time alone does not answer that question.

What AI inference means

Inference is not limited to language models. The MLPerf Tiny research benchmark also includes anomaly detection using sound data, for example. This demonstrates another application area, but does not establish whether a particular model is suitable for a temperature time series.[3]

Applying a model and selecting a suitable analysis method are different tasks. Whether a fixed limit, a control chart or a trained model suits the question must be clarified before planning runtime requirements. A shorter computation time does not make an unsuitable method more suitable.

When the process needs the result

The architectural example concerns a batch whose temperature is recorded during a defined holding phase. An already trained model assesses the completed measurement window and provides a finding for subsequent domain review. The example concerns neither automatic batch release nor temperature control.

In this example, the finding is intended to be available for domain review before the batch reaches the next process step; plant operations set this deadline, not the model.

If the model requires the complete window, this run can only be assessed once the necessary measurements are available. The duration of the holding phase must therefore be distinguished from the subsequent processing time. A model for completed windows does not imply an early-warning system capable of providing the same finding during the phase.

Domain review may require additional time. It is not completed simply because the model calculates its result quickly.

What takes time between measurement data and a finding

A simple processing chain can be divided into several time components: measurement data is made available and, where necessary, transmitted, prepared for the model, scheduled for analysis, processed and then made available as a result. Depending on the architecture, steps may be arranged differently or overlap. The chain is a measurement plan for the example, not a universally applicable formula.

A short model run may follow a longer wait. The start and end of any stated duration must therefore be defined: Does measurement begin when the complete measurement window is available, when the request is accepted or only when the model starts? Does it end with the model output or with the finding becoming visible in the application? Without these boundaries, two runtime figures cannot be meaningfully compared.

The research paper on the MLPerf Inference Benchmark explicitly states that preprocessing is not timed in its setup. MLPerf Tiny likewise excludes preprocessing and postprocessing from its measurement window. Such results can support a clearly defined measurement of model execution; they do not automatically describe the entire duration from a measurement value to an available finding.[1][3]

Schematic processing chain with the model step highlighted.
Schematic workflow shown in serial order, starting with the complete measurement window. Model execution time is only part of the total technical duration; domain review follows separately. Widths and spacing do not represent measured time components. Depending on the architecture, steps may be arranged differently or overlap. Plant operations define the required deadline.

The input side also needs a clear reference: timestamps and the status of a measurement message help distinguish when a value originated from when it arrived. When timestamps from different computers are compared, their time base must be suitable for that measurement.

Why throughput and response time need to be assessed together

Response time describes the duration of a defined operation. Throughput describes how many inputs or requests are processed within a period of time. A system can process many windows overall while still delivering individual results too late. Both measures are therefore needed to answer the process question.[1]

MLPerf Inference distinguishes several load scenarios and uses different metrics depending on the scenario. The server scenario considers throughput subject to a latency condition; an offline scenario concerns inputs that are already available. These are methodological differences, not timing requirements that can be taken from the benchmark and applied to a batch.[1]

For our example, an additional load case arises when several units of equipment finish their holding phases at almost the same time. Several windows then arrive for analysis together. The test must examine whether these arrivals cause waiting times and whether the findings still become available within the defined deadline. This case is derived from the production workflow; it is not attributed to the benchmark scenarios.

Batching requests can increase throughput while also increasing response time. The serving research on Clipper describes precisely this trade-off. Whether batching helps in a particular setup must therefore be measured under the intended load and against the required deadline.[2]

An average alone does not show how long slower operations take. The distribution of response times can also be examined, for example a percentile and the deadline overruns actually observed. A percentile from a test run is not a promise that a maximum deadline will be met in every future case.[1]

What capacity the specific run requires

These measurement principles can be used to derive a test plan for the intended environment. It starts with the application run and its conditions, rather than a blanket statement about the computing power required. The following questions are an architectural proposal for the example, not a mandatory benchmark specification.

  • Which model version is run, with what input format and window size?
  • Which preparation and result-processing steps belong to the measured workflow?
  • Which software and technical environment are used?
  • What load is expected, and which simultaneous window arrivals should be tested additionally?
  • Which domain-specific quality criteria and deadline must the run meet?

The defined total duration is then measured along with its individual components, where these can be captured in the setup. Throughput, response-time distribution, failed operations and deadline overruns are recorded separately. This makes it possible to investigate whether the model itself, a queue or another part of processing limits the workflow.

Comparing different environments requires the same domain task and a traceable record of the conditions. MLPerf Inference combines performance measurements with quality requirements. For our comparison, this means that a faster variant is unsuitable if it no longer meets the previously defined result quality.[1]

If the model version is also changed during the capacity test, the model versions should be compared on the same cases. Runtime measurement does not replace this domain comparison. Nor can a token rate for a language model establish how quickly a different model assesses our temperature windows.

What remains unresolved when a result is late

For the display in this example, three pieces of information remain separate: whether the window has already been assessed, whether the result was made available within the defined deadline and what the completed assessment found. A result can therefore be both late and unremarkable.

This distinction concerns the analysis result and complements the distinction between missing and delayed measurement messages on the input side. An unusual finding then requires a domain review. An assessment that is still pending must not appear as an unremarkable result in this display.

Primary sources and further reading

[1] MLPerf Inference Benchmark. Vijay Janapa Reddi et al., arXiv:1911.02549v2, 2020. The methods and measurement boundaries of this version are used, not current benchmark rankings. Original source

[2] Clipper: A Low-Latency Online Prediction Serving System. Daniel Crankshaw et al., arXiv:1612.03079v2, 2017. Repository manuscript. Original source

[3] MLPerf Tiny Benchmark. Colby Banbury et al., arXiv:2106.07597v4, 2021. Original source