Reliable data throughout the lifecycle

Data Integrity: When Data Is Complete, Consistent and Traceable

Measurement M-23 has been captured completely. What remains when only the numerical value and object label are exported?

System KnowledgeNESS Online GmbHPublished: Last updated:

Brief definition: Data integrity is the degree to which data is complete, consistent, accurate, trustworthy and reliable, and retains these properties from creation through retention or controlled deletion. [1]

What does data integrity mean?

The MHRA defines data integrity as the degree to which data is complete, consistent, accurate, trustworthy, and reliable, and retains these attributes throughout the data lifecycle.[1] PIC/S uses an almost identical definition.[2] The FDA summarizes the core requirements for CGMP data as completeness, consistency, and accuracy.[3]

Data integrity is therefore a property of how data is handled as a whole. It does not arise only during a later review. The design of a task, the recording of an observation, its attribution to a person, and its connection to the affected material already determine how reliable the resulting record can be.

A complete measurement loses its context during export

In a fictional case, measurement record M-23 contains a value with its unit, measured object, equipment identifier, time, and required change history. Its capture is traceable. During a later export, however, only the numerical value and an abbreviated object label are transferred.

The file opens, and its row count is correct. Nevertheless, the recipient lacks the basis for interpreting the value correctly. Without the unit, they cannot calculate reliably. Without the change history, they may not recognize whether they are looking at an original or corrected state.

An integrity review therefore follows the record across the transition: what information was available before, what was transferred, and which functionality must be preserved at the destination? Successful transport does not answer these questions. The export rule and its output must be corrected before the reduced representation is used as complete evidence.

Data must remain reliable throughout its lifecycle

The data lifecycle begins with generation or initial capture. This may be followed by processing, review, use in decisions, transfer, retention, archiving, and ultimately controlled deletion.[1][2][3]

A record may therefore be correct when created yet become unusable later. This can happen, for example, when:

  • required metadata is separated from the actual values,
  • a migration changes meaning or relationships,
  • file formats can no longer be read,
  • changes obscure the original entry,
  • permissions do not allow actions to be attributed unambiguously,
  • parts of the dataset are lost during export or archiving, or
  • an existing backup cannot be restored reliably.

Data integrity therefore requires an end-to-end perspective. What matters is not only the final state, but also the conditions under which data was generated, changed, reviewed, used, and retained.

Why context and metadata are part of the record

The numerical value 23 is of little use without context. Additional information is needed to understand whether it means 23 milligrams, 23 kilowatt-hours, a point in time, or an internal identifier.

The FDA describes metadata as contextual information needed to understand data. Examples include timestamps, user identifiers, equipment identifiers, material status, material identification numbers, and audit trails.[3] The MHRA likewise treats metadata as an integral part of an original record.[1]

In an operational process, metadata may answer questions such as:

  • Who recorded a value or caused it to be generated automatically?
  • When and within which task was it generated?
  • Which batch, material, or piece of equipment does it concern?
  • Which unit, method, and SOP version applied?
  • Which previous state was changed?
  • Which review, correction, or release followed?

Metadata is not the same as data provenance. It does, however, provide a substantial part of the context from which relationships describing origins and generation can be modeled. Data Provenance describes how data connects to operations, source data and actors.

Cannabis laboratory results illustrate why a reported number needs its analytical context.

Published laboratory comparison · cannabis

THCA: the quantification limit matters

THCA: the quantification limit mattersLC-UV/PDA laboratories with sufficient LOQ: Hemp Oil 1, 17%; 2, 22%; 2a, 12%.Laboratories with sufficient LOQ (%)Hemp Oil 117 %Hemp Oil 222 %Hemp Oil 2a12 %0255075100
NIST CannaQAP, p. 21. Published percentages for LC-UV/PDA laboratories at the respective consensus levels; no pass/fail assessment.[6]

For data transfer, preserve the method, unit, quantification limit (LOQ) and result qualifier. The chart does not establish data manipulation.

Data integrity in migrations and interfaces

Data frequently moves between systems, formats, or organizational areas of responsibility during its lifecycle. Each transition can change relationships, units, time information, status information, or metadata.

When data is transferred to another format or system, Annex 11 requires checks that its value and meaning are not altered.[4] In practice, this means that a migration is not successful merely because the same number of files arrived at the destination.

Checks may include:

  • completeness of the transferred records,
  • preservation of identifiers and relationships,
  • correct assignment of units and time information,
  • transfer of required metadata and histories,
  • readability and usability for analysis in the target system,
  • handling of errors, duplicates, and functions that cannot be transferred, and
  • traceable approval of the migration result.

Similar principles apply to interfaces. Built-in checks can prevent incomplete or structurally invalid data from being processed further without detection. The required controls depend on the significance and risk of the information being transferred. For the procurement of compliance software, this translates into concrete demonstration cases: an export must preserve the required context, and a failed transfer must be identifiable as an error.

Originals, true copies, retention, and backups

A record does not always consist solely of the visible content of a document. For dynamic electronic data, metadata, relationships, processing steps, and interactive capabilities may also be needed to preserve content and meaning.

A printout or PDF is therefore not automatically a complete replacement record. Whether a copy qualifies as a true copy depends on whether it preserves the required content, meaning, and associated metadata.[1][3]

Backups and archives also serve different purposes:

Term Main purpose
Backup Recovery after loss or failure
Archive Long-term, protected, readable retention of originals or suitable true copies

The MHRA makes clear that recovery backups do not replace long-term retention of data and metadata.[1] Annex 11 requires regular backups and checks of their integrity, accuracy, and restorability.[4]

A successfully created backup is therefore only the first step. Without tested restoration, there is no evidence that the data will actually be usable when needed.

Changes, corrections, and audit trails

An audit trail keeps relevant changes traceable. A correction is not automatically an integrity problem. Errors can occur and must be correctable. What matters is that the change does not silently replace the original context.

The MHRA describes an audit trail as metadata about actions associated with creating, modifying, or deleting records. It should allow the history of such events—including “who, what, when, and why”—to be reconstructed without obscuring or overwriting the original record.[1]

An audit trail alone does not establish data integrity, however. Its evidential value depends on whether:

  • relevant operations are actually captured,
  • time and identity are reliably attributed,
  • ordinary users cannot disable or modify the function,
  • reasons for relevant changes are documented,
  • entries remain available in an understandable form, and
  • relevant events are reviewed on a risk basis.

PIC/S explicitly notes that audit trail functionality must be appropriately configured and verified, and that relevant audit trails must be reviewed regularly according to risk.[2]

Roles, permissions, and separation of duties

Not everyone needs the same permissions. A person performing a work step has a different task from someone reviewing or approving its result or managing the technical configuration.

For computerized systems, PIC/S describes individual user identifiers, different roles, and the principle of least privilege. Administrator rights should be strictly controlled and separated from ordinary activities.[2] Annex 11 requires appropriate access levels, defined responsibilities, and documentation of the creation, modification, and cancellation of access authorizations.[4]

Two effects are important for data integrity:

  • Actions can be attributed to a specific identity and role.
  • Unauthorized entries, changes, deletions, or approvals are restricted by technical means.

Shared accounts weaken attribution. Conversely, an individual account does not prove that an action was correct in substance. Roles, permissions, process requirements, and subsequent review must work together.

Place controls at the transitions

The export case requires different controls from the original measurement. At capture, attribution, units, and the measurement method are central, for example. During transfer to the target system, selection, conversion, metadata, and error handling must also be checked.

For every transition, it should be clear who determines the required scope and who assesses the transfer. A technically responsible team can check that fields were transferred correctly; the substantive significance of an omitted item also needs to be assessed.

Later controls can detect deficiencies and enable targeted corrections. They cannot, however, reconstruct missing original observations at will. The FDA emphasizes system design and controls for detecting errors and omissions throughout the data lifecycle.[3]

Complete, consistent, and accurate

These three terms form the core of many regulatory definitions, but answer different questions.

Completeness

Complete data contains the information required for the intended purpose. For an executed production task, this may include more than the final measurement: the material batch used, the quantity actually used, the applicable SOP version, relevant times, the performing role, deviations, corrections, and the review result.

For automated compliance evidence, merely retaining this information is not enough. Its context must be preserved during execution so that it can support a specific statement about the operation.

For batch-specific documentation, an Electronic Batch Record brings such information together in the context of the batch that was actually executed.

Completeness does not mean storing everything indiscriminately. Scope and granularity must fit the process, the significance of the data, and the risk. A risk assessment does not make required records optional; retrospectively selecting only favorable results is not completeness. A dataset is not complete if it lacks precisely the information needed to reconstruct an activity or decision.

Consistency

Consistent data does not contradict itself without an explainable reason. Identifiers, chronological sequences, states, and relationships must fit together logically.

Clarification is needed, for example, if material is recorded as consumed only after a task is completed, a release precedes its associated review, or two systems assign different states to the same batch. A discrepancy does not automatically imply an integrity violation. It must, however, be visible and explainable in its operational context.

Accuracy

Accuracy concerns the correct representation of an observation, entry, or calculation. Technical controls can check input formats, value ranges, or calculations. They cannot always determine whether a real-world observation was factually correct.

Accuracy is therefore linked to operational execution: suitable measuring equipment, understandable requirements, qualified people, verifiable calculations, and appropriate checks. Data integrity is thus neither exclusively an IT task nor exclusively a documentation issue.

ALCOA+ makes the attributes of a record concrete

ALCOA+ brings together nine attributes of records, including attribution, contemporaneous capture, completeness, and availability.[1] The mnemonic supports the assessment of a specific record.

In the export case, attention shifts to the lifecycle: attributes that M-23 possessed at capture must also be appropriately preserved during transfer and retention. A one-time check at the beginning is not enough.

Preserving the record and its context across processing steps

In 420+, the material, SOP version, person or role, and result are linked within the task context. This relationship remains understandable after correction and transfer. The ledger architecture places the record in its history.

For process events, this includes the required time, object, and result relationships. Their transfer to process mining must preserve this meaning in the chosen analysis.

Regulatory finding

Finding: The FDA cited a drug manufacturer for the absence of a required dilution factor in its laboratory information management system for microbiological results. According to the letter, products were released solely on the basis of the raw plate counts entered into the system.[5]

Assessment: Data integrity also concerns the relationship between a measured value and the result derived from it. Correctly storing a raw value is insufficient if a required calculation step is missing.

Review question: Does the system check the basis of the calculation before a result becomes eligible for release?

Letter dated 18 August 2026 · Source checked on 18 September 2026.

This section presents the selected regulatory finding as stated at the time of the letter. Company responses and subsequent developments are not assessed here; this account does not describe the company’s current compliance status.

Reliability must survive a change of system

A reliable approach therefore follows the record and its required context through modification, transfer, and retention. The decisive test is whether it can still be understood and reviewed for its intended purpose at each destination.

The scope of a data transfer review follows from the intended use and applicable requirements. It cannot be replaced by a blanket description of a system as secure, immutable, or compliant.

Primary sources and further reading

  1. Medicines and Healthcare products Regulatory Agency (MHRA), GXP Data Integrity Guidance and Definitions, Revision 1, March 2018, particularly Sections 3, 5, and 6. Original source
  2. Pharmaceutical Inspection Co-operation Scheme (PIC/S), Good Practices for Data Management and Integrity in Regulated GMP/GDP Environments, PI 041-1 of July 1, 2021, particularly Sections 2.4–2.6, 5, and 9. Original document
  3. U.S. Food and Drug Administration (FDA), Data Integrity and Compliance With Drug CGMP: Questions and Answers, Guidance for Industry, December 2018, particularly Question 1. Original source
  4. European Commission, EudraLex Volume 4, Annex 11: Computerised Systems, January 2011 revision, particularly Sections 1–2, 4–7, 9–13, and 16–17. Original source
  5. U.S. Food and Drug Administration (FDA), Warning Letter: Reliance Life Sciences Private Limited, MARCS-CMS 730801, 18 August 2026. Item 1, subsection “Inaccurate Microbiology Laboratory Records”. Original source Accessed 18 September 2026.
  6. Abdur-Rahman, Phillips, Wilson (2021): Cannabis Quality Assurance Program: Exercise 1 Final Report. NISTIR 8385. Original. Section 2, THCA, p. 21; reporting recommendations p. 23. Checked 18 September 2026.