From the scanned code to the correct assignment

Data Matrix code: structure, differences from QR codes and use in operations

A Data Matrix code stores data in a two-dimensional pattern of light and dark modules. It can, for example, be printed on a label or marked directly on a component. A suitable reader captures the pattern and decodes its content. Whether the data read matches the expected material and the current work step is a separate question for the application.[1][2]

System KnowledgeAuthor: Hannes SchubertPublished: Last updated:
Brown laboratory bottle with a blue cap and an added Data Matrix code on its label
Illustrative image/photomontage: laboratory bottle with a subsequently added Data Matrix code and a fictional example identifier.

What is a Data Matrix code?

Data Matrix is a code type that represents data in a square or rectangular area. Unlike a linear barcode, its arrangement uses both dimensions. The ECC200 variant uses Reed-Solomon error correction, which can compensate for certain errors in the captured data.[2]

Such a code can carry an internal operational identifier. In GS1 DataMatrix, by contrast, the data it contains follows the rules of the GS1 system. Its appearance alone therefore does not establish whether the content identifies a material number, a batch or an individual item. The application must evaluate the content and its structure to determine this.[1][2]

Structure and size

An L-shaped, continuously dark finder pattern helps the reader locate the symbol and determine its orientation. Light and dark modules alternate along the opposite sides. The encoded data area lies within this boundary. GS1 DataMatrix also requires a clear quiet zone one module wide on every side.[1][2]

Two aspects of size need to be distinguished: the number of modules and their physical size. How many modules are required depends on the data to be encoded and its encoding. In GS1 applications, the module size, also called the X-dimension, follows the requirements of the particular application. A universally applicable minimum size in millimetres would therefore be misleading.[1][2]

This has a practical consequence: additional data may require more modules. If these are squeezed into the same available area, the individual modules become smaller. Planning must therefore align the amount of data, the marking area, the production method and the reading conditions. The quiet zone forms part of the space required.[1][2]

Error correction does not guarantee recovery from arbitrary damage. The GS1 specification describes the correction capabilities for each symbol size in terms of codewords. This does not support a general assurance that a specified percentage of the visible code area may be missing.[2]

Comparing Data Matrix and QR codes

Both are two-dimensional codes. However, a QR code is not a subtype of Data Matrix. Differences include the finder pattern and error correction. In practice, the data structure and readers supported by the application also matter.[2][4]

FeatureData Matrix ECC200 / GS1 DataMatrixQR code in the standard form considered here
Finder patternL-shaped boundary supplemented by alternating modulesProminent finder patterns at three corners
Error correctionFixed capacity for each symbol sizeFour selectable error correction levels
GS1 applicationGS1 DataMatrix uses GS1 element strings; Data Matrix with Digital Link must be distinguished from itA QR code can carry GS1 Digital Link in URI form; GS1 QR Code also exists with its own GS1 data structure

[2][4]

For an organisation, selection therefore starts with the application: What information needs to be read? What data structure does the target system expect? What marking can be applied to the object and read under the actual conditions? A blanket ranking of the two code types does not answer these questions.

What GS1 DataMatrix adds

GS1 DataMatrix combines the code type with a standardised data structure. Application Identifiers, or AIs, specify the meaning of the data that follows. AI 01 denotes the GTIN, AI 10 a batch or lot number, and AI 21 a serial number.[1][2]

The FNC1 control character in the first position indicates GS1 use. Certain consecutive data fields also require separators. The parentheses used around AIs in the human-readable text are not part of the encoded data. These rules help software interpret the fields correctly; they do not establish the authenticity of the marked object.[1][2]

GS1 DataMatrix and Data Matrix with GS1 Digital Link must also be distinguished. The former uses GS1 element strings. Digital Link represents identification data in a web-compatible URI structure that can be contained in either a Data Matrix or a QR code. The application must support the variant being processed.[4]

Reading, interpreting and checking quality

In scanning, GS1 distinguishes capturing the symbol from decoding the captured image. The resulting data is then passed to an information system for further processing. GS1 DataMatrix uses image-based readers or suitable camera systems.[1][2]

Symbol verification serves a different purpose: it assesses the quality of the symbol against defined criteria. A successful scan and a verification result therefore express different things. In their guidance on interpreting verification results, the GS1 General Specifications explicitly state that the correctness of the data content cannot be confirmed without additional software linked to a database. Such a connection alone does not guarantee correctness either.[2]

It is also necessary to check whether the human-readable text and code content agree. A fictional example: a label states batch CH-50, while the code contains CH-56. Checking symbol quality alone does not resolve this contradiction. An overall system can provide additional checks for this purpose; these must be distinguished from symbol verification.[2]

Even a symbol that was flawless when produced can be damaged later. A sample check does not automatically confirm the quality of every symbol in a production batch either. For operations, this means that marking and capture should be tested under actual conditions of use, including cases in which no usable content can be read.[2]

From scanning to the process step

Before assignment, it is necessary to ask what the identifier actually denotes. For trade items, the GTIN distinguishes the item type; the additional batch or lot number assigns it to a corresponding group. Only the combination of GTIN and serial number identifies an individual instance in the model considered here. AI 21 is therefore not a globally unique object identifier on its own.[2]

Two containers of the same item type and batch can carry the same encoded content if no distinguishing data is added. Scanning that content does not reveal which of the two containers is in front of the reader. Distinguishing individual containers requires identification suitable for that purpose.[2]

For process architecture, a sequence of checks can be derived from this: first, the identifier is read and its data structure evaluated. The identified object or lot is then compared with the expected context. Only the business rule determines the action that follows and the result to be documented. This sequence of checks is an architectural inference, not a general GS1 requirement.[1][2]

The EPCIS event model illustrates the separation of identification and context: among other things, it distinguishes individual objects from quantities at class level and adds information about time, location and business step. Here, eventTime denotes the time at which the event occurred according to the capturing application; recordTime concerns recording in the repository. These times are not evidence of physical execution obtained from the code itself either. EPCIS serves here as an example of separating the layers, not as a prerequisite for using Data Matrix.[3]

When the assignment does not match

The following example uses only fictional internal identifiers, not GS1 keys. A work step expects material type MAT-17. MAT-19 is read. The code was decoded successfully, but the identified material does not match the requirement. For this example process, the rule is that the material assignment remains unresolved until the discrepancy has been clarified and the decision documented. This is a chosen process rule, not a function provided by the code.

A second case shows why the level of identification matters. Two containers of MAT-17 belong to batch CH-42. If their codes contain only this material and batch identifier, two identical scan results can mean either two containers or the same container being read twice. The code content alone cannot distinguish between them.

Even with a unique individual-object identifier, the process context remains decisive: the same container can legitimately be read first at goods receipt and later again during staging. A repeat at the same work step, by contrast, can trigger an unintended duplicate posting if every capture is treated as a new transaction without checking. The application therefore needs a rule that fits the action and the expected quantity; indiscriminately rejecting identical identifiers would be equally inadequate.

Before implementation, the following should be clear for every capture point: What level is being identified, what state is expected, and what happens if a scan does not match or is repeated? Only these decisions make it possible to determine what a scanned Data Matrix code may trigger in the specific workflow.

Primary sources and further reading

  1. GS1 DataMatrix Guideline, Release 2.5.1, January 2018
  2. GS1 General Specifications, Release 26.0, January 2026
  3. EPCIS Standard, Release 2.0, June 2022
  4. GS1 GO: What is the difference between the 2D barcode options …, updated 15 May 2025