Pharma AI Compliance 2026: Annex 22 and the Validation File

Life sciences spent two years asking whether regulators would write AI rules. In 2026 they did, and the rules are more specific than most quality organizations expected.

By Mohamed Ali|September 10th, 2026|11 Min Read

On September 8th, 2026, HHS confirmed four senior FDA appointments, and one of them was new: Jared Seehafer as the agency's first Deputy Commissioner for Technology and Artificial Intelligence. The role sits above the individual centers, and its stated remit includes both AI in regulatory submissions and inappropriate AI use in GMP settings. When a regulator creates a deputy commissioner for something, that something has stopped being a pilot.

Pharma AI compliance in 2026 is no longer a reading exercise about discussion papers. There is a draft GMP annex that says exactly which model types may touch a critical decision, an FDA framework that asks for a credibility file before it asks about your accuracy number, a set of joint FDA and EMA principles, and a horizontal European law whose deadlines just moved in a way most quality teams have not yet absorbed. This guide walks through all four, and then through the thing nobody put in the budget, which is the paperwork.

1. The Year Pharma AI Got Its Own Rulebook

Two years ago a life sciences company asking how to validate a model got a shrug and a reference to general computerized system guidance. That gap closed in roughly eighteen months, and the sequence matters because each document assumes the previous one.

WhenWhat happenedWhy it matters
January 2025FDA draft guidance on AI used in regulatory decision making for drugs and biologics, docket FDA-2024-D-4689Introduced the risk based credibility assessment, comments closed April 2025
July to October 2025Public consultation on Annex 22, drafted by EMA with PIC/SFirst dedicated AI framework for GMP manufacturing
January 14th, 2026FDA and EMA publish ten joint guiding principles of good AI practice in drug developmentFirst transatlantic alignment, covering research through manufacturing and safety monitoring
June 30th to July 1st, 2026EMA multistakeholder workshop on generative AI in manufacturingThe exclusion in the draft annex is being reconsidered right now
July 2026The EU AI Act omnibus moves the high risk deadlinesRelief on the horizontal law, no relief on the sectoral one
September 8th, 2026FDA names its first deputy commissioner for technology and AIAI governance becomes a cross agency function, not a center project

The joint principles published on January 14th are short and worth reading in full, because they are the vocabulary both agencies now use. They cover human centric design, a risk based approach, adherence to standards, a clear context of use, multidisciplinary expertise, data governance and documentation, model design and development practices, risk based performance assessment, life cycle management, and clear essential information. Nothing there is exotic. What is new is that the same ten words now appear in questions from two regulators on two continents.

For scale, the FDA had already reviewed more than 300 submissions containing AI components by 2023, well before any of this existed. The documents did not create the practice. They caught up with it, and the teams feeling the most pain right now are the ones whose models went live during the gap.

2. Annex 22 and the Line Through Generative AI

Annex 22 to EudraLex Volume 4 is the first regulatory text anywhere that tells a manufacturer where a model may and may not sit. Its consultation ran from July to October 2025 and finalisation is expected by the end of 2026, so the version everyone is working from is still a draft. The draft draws one hard line and it is worth quoting the logic rather than paraphrasing it loosely.

The annex applies to static models, meaning models whose parameters are fixed and that do not adapt during use. Dynamic models that keep learning in live operation are excluded from critical GMP applications, and as European Pharmaceutical Review summarised in its analysis of the draft, generative AI and large language models are excluded from any GMP critical decision. The door stays open for the non critical lane: summarising deviation reports, searching standard operating procedures, drafting maintenance notes, provided a qualified person stays in the loop.

Model typeGMP critical decisionsNon critical support work
Static model, fixed parametersIn scope, with the full validation burdenAllowed
Dynamic model that adapts in useExcluded as draftedCase by case, under the quality system
Generative AI and LLMsExcluded as draftedAllowed with a qualified person in the loop

Five requirements carry most of the weight. Intended use has to be documented before testing starts, including input ranges, edge cases, likely error modes and sources of bias. The model has to perform at least as well as the process it replaces, with the metrics fixed before testing rather than chosen afterward. Test data must be separated from training and validation data through technical and procedural controls, not by good intentions. The system has to record which features drove each classification or decision, using tools such as SHAP values, LIME or heat maps. And every change to the model, the system or the input sources triggers a revalidation assessment, which is the same discipline we described for prompt versioning in production, applied to a far less forgiving audience.

The clause that reshapes vendor conversations: Annex 22 places full responsibility for validation evidence on the regulated company, whether the model was built in house or bought from a supplier. A vendor certificate is not evidence. If your supplier cannot hand over test data separation proof, performance baselines and feature attribution records, you are the one who will be explaining their absence to an inspector.

One more thing, and it is the reason this section says “as drafted” so often. On June 30th and July 1st, 2026 the EMA ran a two day expert workshop in Amsterdam specifically to revisit the generative AI exclusion, after the 2025 consultation showed support for enabling it under conditions. The agency asked experts for control measures, guardrails, validation paradigms, human oversight models and cybersecurity limits that could support a risk based approach instead of a flat prohibition. Plan for the draft as written, and watch for a revision that trades the ban for a set of conditions you will have to meet.

3. The FDA Framework Starts With Context of Use

The FDA approaches the same problem from the submission side, and its draft guidance from January 2025 is built on a seven step credibility assessment. The order is the point. Teams that start at step five, which is where the interesting engineering lives, end up rewriting the first four steps under time pressure.

StepWhat it asks
1. Question of interestWhich regulatory decision is this model actually informing
2. Context of useHow and where the model operates, its inputs, outputs, and the human roles around it
3. Model riskModel influence multiplied by decision consequence, classified low, medium or high
4. Credibility planThe validation strategy, written before the testing, sized to the risk
5. ExecuteTesting, validation and analysis as planned
6. Document results and deviationsThe credibility report, including everything that did not go to plan
7. Determine adequacyIs credibility established for this context of use, or are mitigations required

Two ideas in that table do real work. Context of use means credibility is never a property of a model on its own. The agency defines it as trust, established through credibility evidence, in a model's performance for a particular context of use, which means the same model validated for one decision is unvalidated for the next one. Model risk is the product of how strongly the output drives the decision and how much harm a wrong output causes. High autonomy raises risk, genuine human review lowers it, and patient safety consequences raise it further. That is the formal version of the argument we made in our guide to human-in-the-loop AI, now with a regulator attaching validation stringency to the answer.

Documentation expectations follow from the risk tier: the question and context of use, the risk classification, a credibility plan written before execution, results with performance metrics covering accuracy, sensitivity, calibration and uncertainty, every deviation from the plan, and life cycle maintenance procedures. The report can travel with the application or be retained for inspection. The guidance comment period closed in April 2025 and a final version has been expected through 2026, so the sensible posture is to build to the draft now rather than wait for a document whose structure is unlikely to change.

4. What Actually Changed on the EU AI Act Calendar

Most life sciences AI roadmaps written in 2025 had August 2nd, 2026 circled as the date European high risk obligations landed. That date came and went without those obligations arriving, and a surprising number of quality plans still assume otherwise.

An omnibus amendment to the AI Act, agreed politically in early May 2026 and adopted in July ahead of the original deadline, moved the dates. As Gibson Dunn set out in its analysis of the agreement, obligations for standalone high risk systems listed in Annex III moved from August 2nd, 2026 to December 2nd, 2027, and obligations for high risk AI embedded in regulated products under Annex I, the category that includes medical devices, moved from August 2nd, 2027 to August 2nd, 2028. The deferral is unconditional, with fixed dates rather than the conditional trigger the Commission had originally floated.

What did arrive on schedule is narrower and easy to miss. Article 50 transparency duties apply, so people have to be told when they are interacting with an AI system, with a grace period running to December 2nd, 2026 for watermarking on systems already on the market. General purpose AI model obligations have been in force since August 2025. So the honest summary for a life sciences organization is that the horizontal law gave you time, and the sectoral rulebook did not. Annex 22 and the FDA credibility framework are the documents setting your next two years, not the AI Act. Our broader walkthrough of the EU AI Act covers the horizontal picture for teams outside the regulated product perimeter.

There is a trap in the extra time. Sixteen extra months on Annex III and twelve on Annex I are enough for a program to lose its sponsor, and quality organizations know how that ends. The work that the AI Act would have forced, which is an inventory, a risk classification, technical documentation and a human oversight description, is the same work Annex 22 and the FDA framework demand on their own schedule. Treat the delay as budget relief, not as permission to stop.

5. What Goes Into the Validation File

Strip away the acronyms and every one of these frameworks asks for the same folder. An inspector or reviewer opens it and expects to find, in order: the intended use statement and context of use, the risk classification with its justification, the data lineage including provenance and usage rights, evidence that the test set was held separate, performance against the baseline process with the metrics that were fixed in advance, feature attribution records showing what drove each decision, a description of the human oversight that actually happens rather than the one in the policy, the change control log with revalidation triggers, and the monitoring plan for the life of the system.

Two items in that list are where real programs fail. The first is the performance baseline, because it requires knowing how well the humans or the prior method performed, and that number often does not exist. Teams discover in week three of a validation effort that they are being asked to beat a process that was never measured. Start measuring the incumbent before you start building the replacement.

The second is monitoring, and there is a subtlety that trips up people who read the static model requirement too literally. A static model does not drift, but the world feeding it does: raw material suppliers change, a sensor is recalibrated, an upstream system starts rounding differently. The parameters hold still while the input distribution walks away, which is exactly the failure mode covered in our guide to drift detection. Pair that with structured robustness work along the lines of adversarial AI testing and you can answer the one question inspectors always ask, which is how you would know if the model had quietly stopped working.

6. The Documentation Load Nobody Budgeted For

Add it up and the cost of this regime is not compute, licences or data science headcount. It is writing. Every model needs an intended use statement, a context of use description, a risk justification, a credibility plan, a credibility report, a change control entry per modification, a monitoring plan and a briefing for the quality council. Companies running twenty models in the regulated estate are looking at a documentation program, not a paperwork task, and it usually lands on the two or three people who already own validation for everything else.

Worse, the source material keeps moving. A team that wrote its gap assessment against the July 2025 draft of Annex 22 now has to reconcile it against the January 2026 joint principles, the outcome of the June 2026 workshop and a set of AI Act dates that changed in July. Keeping a current internal view of the rules is itself a recurring job.

TheBar Perspective

This is the work TheBar is shaped for. You give it a prompt, and a master agent plans the job, runs live web research across the regulatory sources, reads the files you already have, and returns the artifact: a gap assessment against the current draft, a first pass context of use statement for each model in the inventory, the skeleton of a credibility plan, the deck for the quality review board, or a small internal page where the team keeps the state of the rules in one place instead of in six inboxes.

Now the boundary, which in this industry has to be explicit. TheBar is not a validated GxP system and does not belong inside a GMP critical decision, a batch release or a submission pipeline. It is also a cloud backed desktop app: prompts and responses travel to linesNcircles servers, so patient identifiable data and unpublished trial data stay out of it. Use it on published regulatory texts, your own draft procedures and the material your quality system already allows outside validated systems.

What it produces is a draft that a qualified human reads, corrects and signs. That is not a limitation bolted on for compliance reasons, it is the same division of labour Annex 22 describes for the non critical lane: the machine drafts and searches, a qualified person decides. The signature is the deliverable, and the signature is human.

It is worth noting how close this sits to the pattern in regulated healthcare generally. The constraint in our guide to HIPAA compliant AI tools was never the drafting ability of the model, it was the contractual and architectural question of where data goes. Life sciences adds a second question on top, which is whether the output can be defended in an inspection two years later.

7. A 90 Day Readiness Plan

Ninety days is enough to know where you stand, which is a different goal from being compliant. Aim for an inventory you trust, a risk classification you can defend and three complete files rather than twenty half written ones.

WindowWhat you doArtifact at the end
Days 1 to 30Inventory every model and AI assisted step in the regulated estate, including the spreadsheet macros nobody calls AI. Split critical from non critical. Name one owner per entry.An inventory with a draft context of use per model
Days 31 to 60Classify risk as influence times consequence. Write credibility plans for the three highest. Verify test data separation and pull the supplier evidence you are missing.Three credibility plans and a supplier evidence gap list
Days 61 to 90Run a mock inspection on one file. Close the explainability and change control gaps it exposes. Brief the quality council with a decision on what gets retired.One complete validation file, a gap log and a board level deck

The mock inspection in the last window earns its keep. Ask someone outside the project to open a file and request the performance baseline, the test set provenance and the feature attribution for a specific decision made six months ago. The gaps it surfaces are the ones a real inspection would find, and finding them yourself costs a week instead of a response letter. Discovery programs running ahead of this curve, the kind we described in agentic R and D, tend to have the strongest models and the thinnest files, so start the inventory there.

One sequencing note: do not wait for final texts. Annex 22 is expected to be finalised around the end of 2026 and the FDA guidance has been pending finalisation through the year, but both set out a structure that is stable even where the details are not. Nobody has ever been cited for having documented their intended use too early.

The Rules Arrived Before the Files Did

The gap in life sciences right now is not between companies that use AI and companies that do not. It is between models that are running and files that would survive being opened. Two regulators published their expectations in 2026, a third document is being rewritten as we speak, and the European delay bought time rather than forgiveness. The organizations that come out of this well will be the ones that treated the validation file as a deliverable from the start, not as something assembled the month an inspector books a visit.

To be precise about the boundary: TheBar is a free desktop app for chat, documents, slides, websites and web research. It is not a validated GxP system, does not belong in a GMP critical decision or a submission pipeline, and does not act inside your quality systems. Prompts and responses travel to linesNcircles servers, so keep patient identifiable and unpublished trial data out of it. What it does is take a regulatory question, plan the work, search the live web and hand back a brief, a deck or an internal page that a qualified person reviews, edits and signs.

Turn Annex 22 Into a Gap Assessment

Try TheBar, the free AI desktop app for chat, documents, slides, websites and web research. Point it at the published regulatory texts and your own procedures, and get a gap assessment and a quality council deck instead of a folder of PDFs.

Download TheBar Now