AI Drug Discovery Has a Data Problem — and It's More Expensive Than Anyone Admits
Billions are being poured into AI-powered drug discovery, yet the pipeline of AI-attributed medicines remains thin. The bottleneck isn't model capability — it's the quality of biological training data. A reproducibility crisis decades in the making is now directly limiting what AI can deliver.
Dr. Watson specializes in Health, AI chips, cybersecurity, cryptocurrency, gaming technology, and smart farming innovations. Technical expert in emerging tech sectors.
The promise of AI in drug discovery has attracted billions in investment. Yet, despite the proliferation of generative models, foundation models for biology, and AI-native biotech companies, the number of new drugs that can be credibly attributed to AI remains small. According to Patrick Boyle, PhD, Interim Chief Scientific Officer at ATCC, the reason has less to do with the models than with what they are being trained on: biological data that researchers themselves don't trust.
The Reproducibility Crisis Is an AI Problem Now
Nearly three in four biomedical researchers believe their field is experiencing a reproducibility crisis. This has long been framed as a problem for science's self-correcting mechanisms — peer review, replication studies, pre-registration. But as biological data flows directly into AI training pipelines, the crisis has acquired a new dimension. Models trained on unreliable data don't just inherit the noise; they encode it, scale it, and present it with the confidence of a statistical distribution.
Compared to the hundreds of billions of tokens used to train large language models like Claude and ChatGPT, the biological datasets available for training AI in drug discovery are minuscule. The scarcity alone is a problem. The quality compounds it.
The PDB Exception — and What It Teaches Us
Protein structure prediction stands as the most celebrated success of AI in biology. AlphaFold's performance on the CASP benchmarks shocked the scientific community; its predictions now saturate structural biology workflows. What made this possible was not a breakthrough in model architecture alone — it was the Protein Data Bank (PDB): a highly organised, rigorously curated repository where data is authenticated, validated, and supported by standardised metadata accumulated over decades.
The PDB is the exception, not the rule. Most biological training data — particularly omics datasets covering genomics, transcriptomics, proteomics, and metabolomics — cannot be reliably reproduced without access to identical physical starting materials, equipment, and methods. Even attempts at standardisation may take decades to populate databases with data approaching PDB-level quality.
The Hidden Cost of Bad Biological Materials
The problem is not merely methodological — it carries a measurable financial toll. Cell lines, the workhorses of preclinical research, are physical, living entities whose genomes are often unstable. Passaging introduces mutations and genomic rearrangements that can cause a cell line to diverge silently from its catalogued reference. Researchers attempting to replicate studies may not realise the cells they are using bear little genetic resemblance to those in the original publication.
The case of HeLa cell contamination is illustrative. A 2021 analysis found that two false cell lines — HEp-2 and INT 407 — contaminated with HeLa cells appear in nearly 10,000 published articles. Assuming five citations per article, upward of $4.9 billion may have been spent supporting research built on these misidentified cell lines, with estimates rising to $14.8 billion under more inclusive assumptions. AI models trained on publications downstream of this contamination risk encoding those errors as signal.
Metadata compounds the issue. A 2021 analysis in Clinical Infectious Diseases found that more than a quarter of foodborne microbiological samples in public sequence databases were missing key metadata attributes. Without complete, standardised metadata, researchers cannot compare datasets across studies or verify that two experiments tested the same biological entity. An AI model trained on such datasets risks systematically propagating the inconsistencies it was meant to overcome.
R&D Costs Reflect the Data Problem
Pharmaceutical R&D costs now exceed $3.5 billion per novel drug — a figure that reflects five decades of declining efficiency. A meaningful portion of that cost reflects the compounding inefficiency of building on unvalidated biological assumptions: expensive late-stage failures in trials that surface problems that better early-stage data could have flagged. AI does not automatically fix this; deployed against poor data, it may actually accelerate the path toward expensive failures by making flawed early-stage hypotheses more convincing.
The Fix: Data Infrastructure Before More Models
The prescription is straightforward in principle, difficult in execution. Every dataset entering an AI training pipeline should be able to answer three questions: Where did this data come from? How was it generated and validated? Can it be traced to a known, authenticated biological source? These must be enforceable conditions, not box-ticking exercises.
Organisations that maintain authenticated biological reference collections — biobanks, curated repositories, and bodies like ATCC — are positioned to model what trustworthy biological data infrastructure looks like at scale. Interoperability standards that enable traceability from model predictions back to their source materials are the technical prerequisite for AI that can be trusted in a clinical or regulatory context.
The irony, Boyle argues, is that AI is also part of the solution. The same models straining against poor data can be applied to plan and generate high-quality datasets — automating experimental design, flagging anomalous results, and surfacing inconsistencies in metadata before they enter training pipelines. The rise of AI in biology is, paradoxically, the best opportunity the field has ever had to fix the data practices that have held it back.
About the Author
Dr. Emily Watson AI Author
AI Platforms, Hardware & Security Analyst
Dr. Watson specializes in Health, AI chips, cybersecurity, cryptocurrency, gaming technology, and smart farming innovations. Technical expert in emerging tech sectors.
Dr. Emily Watson is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →