
I keep having the same conversation with quality and IT leaders at pharma manufacturers. Most of them are sharp. They know the risks of sending GMP data to third-party AI platforms. They’re asking the right questions. But a few have started saying something that worries me: “The productivity gains are real, nobody’s had a breach yet, and we can’t afford to fall behind.”
That last part is the real driver. The fear of falling behind is pushing people to strap something together, to plug their data into whatever model is available, just to keep pace. And when that happens, “nobody’s had a breach yet” becomes the justification for skipping the data governance work they’d normally never skip.
What’s Happening
AI adoption in pharma manufacturing is accelerating fast. McKinsey estimates generative AI could drive up to 30% efficiency gains in biopharma operations through reduced deviations, better equipment effectiveness, and faster investigations. The use cases are real and growing: deviation management, batch record review, root cause analysis, real-time release testing. Organizations like Sanofi, Takeda, and Amgen are already deploying these tools across manufacturing and quality.
At the same time, the regulatory walls are going up. The EU AI Act classifies AI used in pharmaceutical quality and process control as “high-risk,” requiring documented risk assessments, human oversight, and transparency. The FDA and EMA jointly released guiding principles for AI in drug development in early 2026. ISPE’s GAMP Community of Practice published best practices for AI in regulated environments in mid-2025. The message is consistent: innovate, but you’d better be able to explain what your AI is doing and where your data lives.
The problem is that the speed of adoption is outpacing the rigor of implementation, particularly around data.
What Most People Are Missing
Here’s the part that keeps me up at night. The Varonis 2025 State of Data Security Report found that 99% of organizations have sensitive data exposed to AI tools. Not “at theoretical risk.” Exposed. Kiteworks found that 27% of life sciences organizations acknowledge more than 30% of their AI-processed data contains sensitive or proprietary information.
Think about what that means in a GMP context. Someone on the quality team uploads a deviation history to get a root cause suggestion. A process engineer pastes batch parameters into an AI platform for optimization insights. A scientist feeds a proprietary formulation into a cloud-based tool for structural analysis. None of these people are being reckless. They’re under pressure to move faster, and these tools are right there. But every one of those actions, well-intentioned and productivity-enhancing, creates a data exposure that cannot be undone. Unlike a traditional breach where you change passwords and revoke access, data that enters a third-party AI training pipeline is permanently outside your control.
And the “nobody’s had a breach yet” comfort blanket? It’s already threadbare, and the incidents that have happened are directly tied to AI tools. In 2023, Samsung employees pasted proprietary source code and internal meeting notes into ChatGPT, exposing trade secrets to a third-party training pipeline. A bug in OpenAI’s platform that same year leaked chat history titles and payment data from ChatGPT Plus subscribers. In 2024, over 225,000 OpenAI credentials were found on dark web markets, stolen through malware, giving attackers full access to users’ conversation histories and whatever sensitive information they’d shared. Italy’s data protection authority fined OpenAI €15.6 million that December. LayerX’s 2025 security report found that 18% of enterprise employees paste data into GenAI tools, and more than half of those paste events include corporate information.
These aren’t hypothetical risks. They’re documented exposures that happened because people used AI tools the way most people use them, by pasting in real data to get real answers.
Implications
The FDA has already flagged that cloud-based AI applications can impact oversight of pharmaceutical production data and records. Existing quality agreements between manufacturers and third-party cloud providers often contain gaps in risk management, gaps that widen when AI systems are monitoring and controlling manufacturing processes. BioPhorum’s DISCO team found that digital maturity mismatches between sponsors and CMOs create real obstacles to secure data exchange, even when one partner is technically advanced.
There’s also a competitive dimension. Your manufacturing data, your process parameters, deviation patterns, yield optimization history, is hard-won institutional knowledge. When that data flows into a vendor’s cloud-hosted model that trains on aggregated customer inputs, you’re handing over process intelligence that took years to develop.
And if you can’t document where your data went, who accessed it, and how an AI-assisted decision was made, you have a compliance gap that will eventually become a finding.
My POV
Here’s what bugs me about this conversation: it’s framed as a tradeoff. Benefits versus risk. Productivity versus security. As if you have to choose. I get why people feel that way. When every conference talk is about AI, when other sites in your network are already piloting tools, the pressure to just get something running is real. But you don’t actually have to choose.
The answer is self-hosted AI models deployed in your own cloud environment. Your data stays on your infrastructure. The model comes to your data, not the other way around. No batch records leaving your network. No proprietary process data entering a third-party training pipeline. No quality agreements with gaps you can’t see.
This isn’t a niche technical approach anymore. Deloitte’s 2025 Technology Industry Outlook noted that private, on-premise AI deployments are expanding across enterprises worldwide, driven by data protection requirements and IP security. Open-weight models and improved hardware have made self-hosted deployment practical and production-ready, and providers like AWS and HPE have launched air-gapped deployment options specifically for regulated industries.
The organizations I see making real progress with AI in GMP are the ones that resolved the data question first. They started with “how do we ensure zero data leaves our perimeter?” and worked backward to capability from there. When quality and IT are confident in the data architecture, they say yes to pilots. When they say yes, the organization learns and scales. The data governance decision unlocks everything downstream.
Actionable Takeaways
- •Stop treating data exposure as an acceptable tradeoff. The pressure to keep pace is real, but “benefits outweigh the risk” is a pre-incident rationalization. In any other area of GMP, we wouldn’t accept “it hasn’t gone wrong yet” as a risk justification. Apply the same standard to your AI data flows.
- •Evaluate self-hosted deployment before defaulting to SaaS. Self-hosted models running in your own cloud or on-premise infrastructure eliminate third-party data exposure entirely. The technology is mature, the costs are reasonable, and the compliance posture is dramatically stronger.
- •Map your data flows today, not after the pilot. Where does your GMP data go when it enters an AI system? What does the vendor retain? Is it used for model training? What jurisdiction are the servers in? If you can’t answer these, you’re flying blind.
- •Qualify your AI vendor the way you’d qualify any GMP supplier. Ask for their data handling SOPs. Understand their retention and deletion policies. Confirm whether your data is segregated or pooled. If they can’t support a quality agreement that covers AI-specific risks, that’s your answer.
- •Use maturity frameworks to find your gaps. BioPhorum’s Digital Plant Maturity Model and ISPE’s Pharma 4.0™ assessment are specifically designed for this. Data integrity is consistently identified as a top inhibitor, so know where you stand before you layer AI on top.
- •Demand audit trails and explainability from day one. Under the EU AI Act, this is legally required for high-risk pharma QC applications. Even outside Europe, regulators expect you to explain AI-assisted decisions. Build the documentation infrastructure before an inspector asks for it.
- •Start contained, then scale deliberately. Pick a use case where the data scope is narrow and the value is clear. Prove it works within your own perimeter, then expand.
