Data integrity used to be treated like a review-stage problem.
A record was created. A test was executed. A report was generated. A batch file was assembled. Then quality teams reviewed the evidence and looked for missing signatures, incomplete entries, transcription errors, unexplained changes, or gaps in documentation.
That model is no longer strong enough.
In modern life sciences environments, data integrity failures often begin much earlier than final review. They begin when data is created, configured, transferred, transformed, linked, approved, archived, or used to support a decision. By the time a reviewer finds the issue, the real failure may already have spread through requirements, testing, change control, reporting, deviations, release decisions, and supplier oversight.
That is why data integrity is moving upstream.
The point is not that final QA review no longer matters. It does. But review alone cannot carry the full weight of data integrity. Regulated organizations need to design processes where trustworthy data is produced by default, preserved in context, and controlled throughout its lifecycle.
Recent enforcement patterns reinforce this direction. In a 2026 warning letter to Tentamus India, FDA cited unofficial personal diaries used to record procedures, observations, results, method changes, and deviation descriptions. FDA also stated that reliability of data is compromised when a firm fails to maintain complete records of the conditions and data associated with all tests. The letter required a comprehensive investigation into inaccuracies, omissions, alterations, deletions, record destruction, non-contemporaneous record completion, and other data integrity deficiencies. (U.S. Food and Drug Administration)
That is not a paperwork problem at the end of the process. It is a control problem at the beginning.
Why Data Integrity Can No Longer Wait Until Review
Final review is important, but it is naturally limited.
A reviewer can identify missing information, questionable entries, unexplained changes, or incomplete records. But if the data was created outside a controlled process, stored in an unofficial location, manually transcribed several times, disconnected from the relevant test or requirement, or altered without a clear audit trail, the reviewer is already working downstream from the real failure.
The Tentamus warning letter illustrates this clearly. FDA noted that unofficial personal diaries were used to record CGMP-related data and that the firm lacked procedures to control their use or ensure proper retention. FDA also stated that the lack of complete data compromised the quality unit’s ability to ensure compliance. (U.S. Food and Drug Administration)
This is the upstream lesson: quality units cannot review data properly if the process does not first ensure that data is complete, retained, attributable, and created in the right system.
Data integrity therefore has to be built into workflow design, system validation, access control, audit trails, documentation practices, and change management. It cannot be rescued by review alone.
What “Moving Upstream” Actually Means
Moving data integrity upstream means shifting control earlier in the data lifecycle.
Instead of asking only “Was this record reviewed?”, teams also ask:
Was the data created in the right place?
Was the user authorized?
Was the entry contemporaneous?
Was the source record retained?
Were changes controlled and explained?
Was the data linked to the correct requirement, test, deviation, batch, or change?
Was the system configured to prevent invalid or incomplete entries?
Was the audit trail reviewable?
Was the record available to quality in context?
This is a different mindset.
In a downstream model, data integrity is checked after work is done.
In an upstream model, data integrity is designed into how work happens.
That distinction matters because modern digital operations move quickly. If flawed data enters a workflow early, it can travel through reports, dashboards, deviations, test records, approvals, and decisions before anyone notices. Small data integrity gaps can become large compliance problems because digital systems move mistakes efficiently too. Tiny gremlins, excellent logistics.
FDA’s Data Integrity Guidance Already Points Upstream
FDA’s data integrity guidance is not new, but its message remains very relevant. FDA states that the purpose of its guidance is to clarify the role of data integrity in drug CGMP. It also says the guidance was developed in response to increased findings of data integrity lapses in inspections, and that FDA expects all data to be reliable and accurate. (U.S. Food and Drug Administration)
The same FDA page states that CGMP regulations and guidance allow flexible and risk-based strategies to prevent and detect data integrity issues. It also says firms should implement meaningful and effective strategies to manage data integrity risks based on process understanding and knowledge management of technologies and business models. (U.S. Food and Drug Administration)
The key word is “prevent.”
Prevention requires more than inspection readiness. It requires process design. It requires understanding where critical data originates, where it moves, how it is modified, how it is reviewed, and how it supports regulated decisions.
That is why upstream data integrity is not a new regulatory invention. It is a practical response to what the guidance already expects: reliable, accurate data supported by meaningful, risk-based controls.
Why Modern Digital Systems Make Upstream Control More Important
Life sciences organizations now operate across increasingly connected systems. Data may begin in a laboratory instrument, move through middleware, enter a LIMS, support a deviation investigation, appear in a dashboard, feed a report, and later become evidence in a validation or quality review.
Each handoff creates an integrity risk.
Data can be lost, duplicated, transformed incorrectly, manually re-entered, disconnected from its context, or stored without enough metadata. Access rights may be misaligned. Audit trails may be difficult to review. Interfaces may not be validated deeply enough. Changes may affect data flows without being fully assessed.
The European Commission’s 2025 consultation on revised Chapter 4, Annex 11, and new Annex 22 shows that regulators are also moving toward more explicit expectations for digital data control. The Commission states that the revised Annex 11 establishes enhanced requirements for lifecycle management of computerized systems, comprehensive application of Quality Risk Management, ongoing maintenance of system requirements, supplier oversight, and stronger controls for data integrity, audit trails, electronic signatures, and system security. (Public Health)
That direction matters.
Data integrity is no longer just a recordkeeping topic. It is a computerized system lifecycle topic.
The Problem With “Data Integrity at the End”
When data integrity is treated as a final review exercise, teams often discover problems too late.
A test record may be complete, but the evidence may not be attributable to the right execution step.
A deviation may be closed, but the original observation may have been captured outside the controlled system.
A report may be approved, but the underlying data source may not be clearly traceable.
A requirement may be verified, but the test result may rely on a screenshot stored without proper context.
A batch or validation package may look complete, but the audit trail may reveal unexplained edits, deleted entries, or changes without documented rationale.
At that point, review becomes investigation. Investigation becomes remediation. Remediation becomes delay.
This is why upstream data integrity is more efficient as well as more compliant. It is cheaper to prevent bad data from entering the workflow than to repair a broken evidence chain later.
Data Integrity Begins With System Design
One of the most powerful ways to move data integrity upstream is to build controls into system design.
A well-designed system should reduce the chance that users create incomplete, ambiguous, or untraceable records. Required fields, controlled values, role-based permissions, electronic signatures, validation rules, automated timestamps, version control, audit trails, and linked workflows all support integrity before review begins.
This matters especially in validation and quality systems.
If a test step requires evidence, the system should make evidence capture part of execution, not an optional afterthought. If a deviation is created from a failed test, it should remain linked to the test run and step where it occurred. If a change affects validated functionality, impacted requirements and tests should be linked directly to the change request. If an approved record is revised, previous versions should remain preserved and reviewable.
Upstream data integrity means the workflow itself helps preserve the story.
Data Integrity Also Begins With Requirements
Data integrity is often discussed in terms of records, but it should also begin in requirements.
If requirements do not define how critical data should be created, processed, transferred, reviewed, retained, and archived, validation may miss the controls that matter most. A system may be tested for visible functionality while the underlying data lifecycle remains weak.
For example, requirements should address questions such as:
Which records are GxP critical?
Which fields are mandatory?
Which data changes require reason capture?
Which actions require audit trail entries?
Which records require electronic signature?
Which interfaces transfer regulated data?
Which reports support quality decisions?
Which data must remain available for inspection?
Which users can create, modify, approve, or deactivate records?
The draft revised Annex 11 consultation materials also reinforce the importance of system requirements. The European Commission states that the revised Annex 11 strengthens expectations around the definition and ongoing maintenance of system requirements. (Public Health)
That is an upstream control. If requirements are weak, data integrity controls are likely to be weak too.
Audit Trails Need to Move Upstream Too
Many organizations think about audit trails during review. They ask whether the audit trail can be reviewed when needed.
That is important, but audit trail control starts earlier.
Teams need to define which events are auditable, which changes require old and new values, which actions require reason capture, which users can perform critical actions, how audit trails are protected from modification or deletion, and how reviewers can search, filter, export, and interpret entries.
If these controls are not defined during system design and validation, audit trail review becomes more difficult later. A system may technically have an audit trail, but the audit trail may not support meaningful review.
The European Commission’s Annex 11 consultation page specifically highlights strengthened controls related to data integrity, audit trails, electronic signatures, and system security. (Public Health) That grouping is important because these controls do not work alone. Access control, audit trails, electronic signatures, and security all help protect data integrity from the start.
An audit trail should not be a black box that opens only during an inspection. It should be part of the daily control model.
Supplier Data Integrity Is Still Your Data Integrity
Modern life sciences organizations rely heavily on suppliers, service providers, contract laboratories, SaaS vendors, cloud providers, and external testing partners.
That does not move data integrity responsibility out of the organization.
FDA’s CGMP page states that CGMP regulations contain minimum requirements for methods, facilities, and controls used in manufacturing, processing, and packing of drug products, and that they help ensure a product is safe for use and has the ingredients and strength it claims to have. (U.S. Food and Drug Administration) In practice, that responsibility extends into how organizations select, qualify, oversee, and rely on suppliers whose data supports regulated decisions.
The Tentamus warning letter makes the supplier point especially visible. FDA stated that customers rely on the integrity of laboratory data to make decisions regarding drug quality, and that strict control is needed to ensure laboratory data is retained and additions or modifications are authorized and appropriately documented. (U.S. Food and Drug Administration)
That means supplier oversight is an upstream data integrity control.
Before relying on a supplier’s data, organizations should understand the supplier’s documentation practices, audit trail controls, data retention, deviation handling, method validation, access governance, and ability to provide complete original records when needed.
A certificate, report, or exported result is only as trustworthy as the system that produced it.
Why Change Control Is a Data Integrity Control
Change control is often discussed as a validation activity, but it is also a data integrity activity.
A configuration change can affect required fields.
A software update can affect audit trail behavior.
An interface change can alter data transfer.
A role change can expand or restrict access.
A report change can affect quality decisions.
A test automation update can change evidence capture.
A vendor release can introduce new data handling logic.
If change impact assessment does not evaluate data integrity, the organization may miss risks until after records have already been generated.
This is another reason data integrity is moving upstream. Every change should be assessed for potential impact on data creation, processing, transfer, retention, audit trail behavior, electronic signatures, reporting, backup, archival, access, and review.
Change control should not simply ask, “Does the system still work?”
It should ask, “Can we still trust the data the system creates and manages?”
AI Makes Upstream Data Integrity Even More Important
Artificial intelligence adds another layer to the upstream data integrity conversation.
AI-assisted systems may help draft requirements, propose risk assessments, generate test cases, summarize records, classify deviations, or support decision-making. These activities can improve efficiency, but they also introduce new questions about source context, input quality, output review, traceability, and human approval.
If AI-generated content enters a regulated workflow without strong review and traceability, the data integrity risk begins at generation, not approval.
The European Commission consultation notes that the new Annex 22 on Artificial Intelligence establishes requirements for AI and machine learning in manufacturing of active substances and medicinal products, including intended use, performance metrics, training data quality, test data management, continuous oversight, change control, model performance monitoring, and human review procedures where needed. (Public Health)
That is upstream governance language.
AI does not make data integrity less important. It makes early control more important because AI can generate or transform information quickly. The faster the engine, the better the brakes need to be.
Where AI-Native Validation Infrastructure Fits
This is where AI-Native Validation Infrastructure, or ANVI, becomes relevant.
ANVI is not just about using AI in validation. It describes a validation foundation where systems, requirements, risks, tests, evidence, deviations, changes, approvals, audit trails, and periodic reviews remain connected and controlled throughout the lifecycle.
That matters for data integrity because upstream control depends on connected context.
A mature ANVI approach can help ensure that data is not only stored, but linked to the process that created it, the requirement it supports, the test that generated it, the deviation it triggered, the change that affected it, and the approval that accepted it.
This is the direction modern validation is moving: away from isolated record review and toward connected lifecycle control.
Data integrity becomes stronger when the system knows the story behind the data.
What Life Sciences Teams Should Do Now
Teams can begin moving data integrity upstream with practical steps.
Start by mapping critical data flows. Identify where GxP data is created, modified, transferred, reviewed, retained, and archived.
Review system requirements. Confirm they define data integrity controls clearly enough to support validation and testing.
Strengthen access controls. Ensure users have appropriate roles, unique accounts, and limited privileges based on responsibility.
Evaluate audit trail design and reviewability. Confirm that critical events are captured, protected, searchable, and reviewed according to risk.
Connect data integrity to change control. Assess whether changes affect data flows, audit trails, reporting, electronic signatures, interfaces, or access.
Improve supplier oversight. Confirm that external data sources and service providers maintain records that are complete, attributable, retained, and reviewable.
Make evidence contextual. Link evidence directly to test steps, deviations, requirements, changes, and approvals wherever possible.
Bring AI use under governance. Define intended use, review procedures, traceability expectations, and approval requirements for AI-assisted outputs.
These steps are not cosmetic. They move control closer to the point where data is born.
The Strategic Value of Upstream Data Integrity
Organizations that move data integrity upstream gain more than inspection readiness.
They reduce rework because data issues are prevented earlier.
They improve audit confidence because records are more complete and easier to defend.
They strengthen quality oversight because data is linked to context.
They reduce investigation burden because original evidence is easier to reconstruct.
They improve validation efficiency because requirements, tests, evidence, deviations, and changes are connected.
They support better decisions because teams trust the information in front of them.
Over time, upstream data integrity becomes a strategic advantage. It allows regulated organizations to move faster without letting control fray at the edges.
Conclusion
Data integrity is moving upstream because modern regulated operations leave no safe place for late discovery.
By the time a final review finds incomplete, inconsistent, or poorly controlled data, the real issue may already have affected testing, validation, change control, deviations, reports, supplier decisions, or quality oversight. The better approach is to design workflows, systems, requirements, audit trails, access controls, and change processes so that reliable data is created and preserved from the beginning.
FDA’s data integrity guidance states that FDA expects data to be reliable and accurate, and that firms should use meaningful, effective, risk-based strategies to prevent and detect data integrity issues. (U.S. Food and Drug Administration) Recent warning letters continue to show what happens when records are incomplete, unofficial, destroyed, altered, backdated, or disconnected from quality oversight. (U.S. Food and Drug Administration)
For life sciences teams, the lesson is clear.
Do not wait until review to protect data integrity.
Build it upstream.
