Knowledge Topic
Annex 22 and AI Agents: Human-in-the-Loop Oversight for Regulated Manufacturing
Annex 22 AI agents guidance points regulated manufacturers toward bounded intended use, human-in-the-loop oversight, auditability, and documented accountability.
Industry context
Industry context
Annex 22 should get the attention of every regulated manufacturer exploring AI agents. It is not a broad AI hype document. It is proposed GMP guidance for artificial intelligence and machine-learning systems used in the manufacture of medicinal products and active substances.
At a glance
At a glance
Annex 22 matters for AI agents because it makes one principle clear: autonomy is not a compliance strategy. Human accountability is. The European Commission’s draft Annex 22 provides proposed GMP guidance for AI and machine-learning systems used in the manufacture of medicinal products and active substances.
Most enterprise AI agents today rely on large language models. They may retrieve documents, summarize records, call tools, draft workflow updates, and recommend actions. That makes them powerful, but harder to fit into critical GMP use cases where validation, predictability, explainability, and data integrity are essential.
The practical answer is not to ban every agent. It is to design agents around narrow intended use, clear risk classification, human checkpoints, source evidence, confidence flags, monitoring, audit logs, training, and change control.
For regulated manufacturers, “human in the loop” cannot mean a casual review button. It must mean a trained person has the authority, evidence, time, and responsibility to accept, reject, correct, or override the agent’s output.
On this page
Overview
Overview
The practical message is clear: automate assistance, not accountability.
The European Commission’s draft Annex 22: Artificial Intelligence focuses on AI use in GMP contexts. The related European Commission stakeholder consultation on EudraLex Volume 4 Chapter 4, Annex 11, and Annex 22 is also important because it places AI guidance alongside documentation and computerized-system guidance.
That matters because many AI agents today are built on large language models. They can retrieve records, summarize documents, call tools, draft responses, and recommend workflow actions. In regulated settings, those abilities create a serious question: how do you keep the human accountable when the AI system is doing more of the work?
The answer is not to ban every AI agent. The answer is to design AI agents around intended use, risk classification, source evidence, meaningful human checkpoints, audit trails, and monitoring. The more important the workflow, the stronger the control model must be.
Annex 22 should also be read alongside Europe’s broader AI governance direction. The official Regulation (EU) 2024/1689 Artificial Intelligence Act establishes a risk-based framework for AI in the European Union.
Three pains Annex 22 forces teams to face
Three pains Annex 22 forces teams to face
AI agents can cross system boundaries
A chatbot may produce an answer. An agent may query QMS, retrieve batch records, compare deviations, summarize CAPAs, draft a response, or recommend escalation. The more systems an agent can touch, the more explicit the control model must be.
For example, an agent that summarizes deviation history for human review is very different from an agent that recommends batch release. The first may be a bounded assistance workflow. The second touches critical GMP decision-making and should face a much higher control threshold.
For human oversight specifically, the EU AI Act Service Desk summary of Article 14 is useful because it explains oversight in terms of monitoring, interpretation, intervention, and awareness of over-reliance. The accessible Article 14 human oversight text reinforces the same point: human oversight should reduce risk, not simply appear as a checkbox.
LLM outputs are hard to validate for critical decisions
Many agents rely on probabilistic models. Their answers may vary based on prompt, context, retrieved records, or model behavior. This creates challenges for critical GMP use cases where validation, repeatability, explainability, and data integrity matter.
That does not mean AI agents have no role. It means their role should be bounded, documented, and reviewed.
For life sciences manufacturers, Annex 22 should be connected to quality risk management. ICH Q9(R1) Quality Risk Management is relevant because AI-agent design should be proportionate to potential impact on patient safety, product quality, and data integrity.
Human-in-the-loop can become theater
A review button is not enough. A human who lacks training, time, evidence, or authority cannot provide meaningful oversight. In regulated workflows, the reviewer must be able to detect error, challenge the output, override the system, and own the final decision.
AI agents in GMP settings also need to be considered alongside established computerized-system lifecycle practices. ISPE’s GAMP 5 Guide, 2nd Edition overview is relevant because it frames computerized system validation and lifecycle control in regulated environments.
Ask this → Get that
Ask this → Get that walkthrough
Ask what the agent is intended to do
“What is the agent’s intended use?”
Get
A narrow scope, such as “summarize open deviation history for human review” or “flag missing documentation before QA review.” Avoid broad mandates like “support batch release” or “manage deviations.”
Ask whether the workflow is critical GMP
“Could this directly affect patient safety, product quality, or data integrity?”
Get
A risk classification. If the workflow is critical, the organization should ask whether the AI model type, validation approach, monitoring, and control model are appropriate.
Ask whether the model is appropriate
“Is the underlying model suitable for this intended use?”
Get
A documented decision. For LLM-based agents, the safer pattern is to keep the agent in non-critical, bounded, human-reviewed support tasks unless the system can meet the required control expectations.
Define the human checkpoint
“Where must a qualified person intervene?”
Get
Required review before regulated decisions, record changes, batch actions, deviation approvals, CAPA closure, or escalations. The checkpoint should be designed into the workflow, not added as an afterthought.
Make uncertainty visible
“Can the human see the evidence and uncertainty?”
Get
Source records, confidence notes, missing evidence, out-of-scope flags, and conflicting information. The reviewer should not have to guess whether the AI answer is complete.
The NIST AI Risk Management Framework gives teams a broader AI governance structure for mapping, measuring, managing, and governing AI risks. That framework pairs well with Annex 22 because both emphasize context, risk, monitoring, and accountability.
Document the oversight
“How is review recorded?”
Get
Reviewer identity, training status, timestamp, records reviewed, final decision, edits made, rejected recommendations, and escalation path.
The FDA’s draft guidance on Computer Software Assurance for Production and Quality System Software is relevant because AI agents in production and quality workflows should be assured according to intended use and risk.
Monitor both the model and the human process
“How do we know oversight is effective?”
Get
Output sampling, reviewer consistency checks, exception trend review, access review, model change control, periodic reassessment, and training updates.
This is why Annex 22 belongs directly inside governed AI in manufacturing. Governance is not paperwork after the fact. It is the design of the workflow.
Proof metric or mini case study
Proof metric or mini case study
A useful metric is human review effectiveness. Track the percentage of agent outputs where the reviewer opened source records, resolved missing evidence, documented a decision, and either accepted, edited, rejected, or escalated the output.
A second metric is out-of-scope stop rate. If the agent encounters a critical GMP decision, missing source evidence, or a task outside intended use, it should stop and route to a qualified person.
A third metric is reviewer consistency. If different reviewers handle the same type of agent output in very different ways, the process may need better training, clearer criteria, or a redesigned interface.
This is the difference between decorative oversight and real oversight. Annex 22 treats the human as part of the controlled process. The human is not a rubber stamp. The reviewer must have a defined role, competence, evidence, authority, and responsibility.
Common questions
Frequently asked questions
Does Annex 22 ban AI agents?
No. The practical reading is more precise. AI agents should be kept away from direct critical GMP decisions unless the system can meet the required controls. Agents can still support bounded, non-critical, human-reviewed work.
What does human-in-the-loop mean for AI agents?
It means a qualified person has a defined role, evidence, authority, time, training, and responsibility to review, correct, reject, or override the agent.
Next step
Annex 22 AI Agent Governance Checklist
Use the Annex 22 AI Agent Governance Checklist to map intended use, risk level, human checkpoints, source evidence, audit logs, reviewer training, and escalation rules before deploying AI agents in regulated workflows.
Sources
References
- European Commission. (2025). Annex 22: Artificial Intelligence draft guideline.
This is the central reference for the article. It is relevant because it proposes GMP expectations for AI and machine-learning systems used in medicinal product and active substance manufacturing.
- European Commission. (2025). Stakeholders’ Consultation on EudraLex Volume 4 Good Manufacturing Practice Guidelines: Chapter 4, Annex 11 and Annex 22.
This page is relevant because it explains the consultation context and positions Annex 22 alongside documentation and computerized systems guidance.
- European Union. (2024). Regulation (EU) 2024/1689, Artificial Intelligence Act.
The AI Act is relevant because it establishes a broader European risk-based AI framework. It reinforces the importance of human oversight for high-risk systems.
- EU AI Act Service Desk. (n.d.). Article 14: Human oversight.
This reference is useful because it summarizes human oversight in practical terms, including monitoring, interpretation, override, and awareness of over-reliance.
- Artificial Intelligence Act. (n.d.). Article 14: Human oversight.
This source provides an accessible rendering of the human oversight provision and reinforces that oversight must reduce risk, not simply appear in the workflow.
- International Council for Harmonisation. (2023). ICH Q9(R1) Quality Risk Management.
ICH Q9(R1) is relevant because Annex 22 agent design should be risk-based, especially where patient safety, product quality, or data integrity could be affected.
- International Society for Pharmaceutical Engineering. (2022). GAMP 5 Guide, 2nd Edition overview.
GAMP 5 is relevant because AI agents in GMP settings must be considered alongside established computerized system validation and lifecycle practices.
- National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework 1.0.
NIST AI RMF gives teams a general risk-management structure for AI governance, including mapping, measuring, managing, and governing AI risk.
- U.S. Food and Drug Administration. (2022). Computer Software Assurance for Production and Quality System Software.
This FDA draft guidance is relevant because AI agents in production and quality workflows should be assured according to intended use and risk.