R&D Is Investing in AI, Just Not Where It Matters Most
Key Takeaways
- Adoption rates remain low for generative design, biomarker analysis, and ADME prediction, indicating ROI risk when investment outpaces scientists’ willingness to act on outputs.
- Scientific accountability demands traceability, reproducibility, and defensibility; confident but unverifiable answers fail operational standards for program decisions and are rationally rejected at the bench.
AI adoption in pharma drops sharply in the domains that matter most for discovery.
The pharmaceutical industry is on track to spend $25 billion on AI by 2030, and the case for that investment—with faster discovery, lower attrition, and sharper decision-making at every stage of the pipeline—is compelling. However, there are other numbers that should give every biopharma executive pause. Drug Discovery News reports that AI adoption in pharma drops sharply in the domains that matter most for discovery: 42% for generative design, 40% for biomarker analysis, and 29% for ADME prediction, with the bottleneck reportedly not the models themselves, but the fragmented data environment underneath them.
The gap between what leaders are funding and what scientists are willing to act on is the risk underlying one of the largest technology bets our industry has ever made. Investment in AI does not guarantee adoption, and if the scientists at the bench don't trust the tools, the promised return on that investment will never materialize.
The trust gap is real, and it has a clear cause: AI systems built on incomplete, unvalidated, untraceable data that scientists cannot defend, reproduce, or stake a program decision on. Closing the gap requires a shift in where pharma is putting its attention. The next phase of AI investment shouldn’t just be about bigger models and broader data. It has to be about data readiness from aggregating, normalizing, and harmonizing multiple sources of internal and external data into an authoritative, AI-ready database. Without those fundamentals, AI investments to drive drug discovery will continue to fall short of expectations.
Why scientists are skeptical of generalist AI, and what could change that
It may be tempting for executives to interpret slow AI adoption inside R&D as a cultural problem, with a workforce that needs to be brought along, trained up, or persuaded to embrace new tools. While those factors may also play a role, the bigger issue may be scientists who are hesitating to rely on AI outputs for valid reasons aligned with their roles, responsibilities and training.
Scientific work has a built-in accountability standard: every conclusion has to be defensible, every result has to be reproducible, and every claim has to be traceable to a source. That standard isn’t a preference; it’s how science works and how programs that depend on it move forward without backtracking later. AI outputs don’t meet that standard by default. A confident-sounding answer that turns out to be wrong, or one whose source can’t be located, fails the professional bar scientists are trained—and accountable—to apply.
A scientist hesitating to act on an AI-generated answer isn’t resisting the technology. They’re doing what their training requires: refusing to stake a program decision on something they can’t defend, reproduce, or trace. Skepticism, in that context, isn’t the obstacle to AI adoption. It’s the standard the technology has to clear.
The trust problem compounds in drug discovery
The trust issues playing out across science and healthcare become more pronounced when it comes to drug discovery, and the reason has less to do with AI itself than with what AI is being asked to learn from. Broad models trained on the open web were not created with the complexity of drug discovery data in mind. When these models are pointed at specialized scientific questions, errors are introduced before the first answer is generated.
Drug discovery depends on domain-specific data structures, controlled vocabularies, and annotation standards that general-purpose datasets don’t reflect. A single chemical compound may appear under thousands of different names across databases and publications, and AI models trained on that inconsistency inherit the confusion and carry it forward. Outputs that look authoritative on the surface can mask a foundational mismatch between the question being asked and the data the model has to work with.
In my role at CAS, I hear a consistent theme from pharma leaders facing challenges in achieving adoption and accelerating discovery with AI initiatives: data readiness matters more than data volume. Pouring more general scientific content into a model doesn’t make it better at drug discovery. It often makes the inconsistency problem worse.
The downstream effect is what makes the stakes so high. Scientists aren’t trained to catch the failure modes of a language model, and they shouldn’t have to be. When a subtly wrong AI output enters a drug discovery workflow, it can move through a program before anyone notices, leading to wasted time and budget on results that can’t be reproduced, or to delays when compromised work has to be repeated. One widely cited case traces a single AI-introduced error through a chain of downstream impacts, showing how quickly one bad output can propagate when the underlying data isn’t built for the work being done.
This is what makes the trust gap so consequential in our industry. The cost of misplaced trust in drug discovery is far higher than in almost any other domain where AI is being deployed.
The solution: Human-curated, traceable data foundations
Addressing the trust gap in AI comes down to three things: a knowledge base scientists can trust, an AI system that leverages content appropriately, and traceability that lets scientists see exactly where an answer came from before they act on it.
Trust starts with the underlying content. When AI draws from a rigorously human-curated knowledge base—scientific information aggregated across literature, patents, and structured databases, normalized under consistent vocabularies and identifiers, and harmonized by scientific experts into an authoritative, AI-ready data asset—the foundation can support the work scientists need it to do. The integrity of an output cannot exceed the integrity of the data feeding it, and in drug discovery, the difference between curated and uncurated content is the difference between an answer worth acting on and one that has to be re-verified.
However, a trusted foundation only matters if the system built on top of it can use it well. Drug discovery questions are rarely answered in a single retrieval step. They require breaking a complex problem into targeted, multi-step search paths and applying domain-specific reasoning at each one. AI designed by scientists for R&D that approaches scientific questions this way produces more nuanced, defensible answers than a general-purpose model because the workflow is designed to mirror how a scientist would actually approach the problem.
Finally, none of this is useful if the scientist can’t see where a given answer originated. A recommendation a scientist can’t trace is a recommendation they can’t defend. When an output can’t be verified, scientists must manually re-verify the finding before acting on it, adding hours and days to workflows that AI was supposed to accelerate. The verification burden erodes the ROI case for the tool entirely.
Trust will decide whether AI delivers
The conversation around AI in pharma has been dominated by questions about what models can do, how large they can get, and how much data they can be trained on. The more important question—and the one that will determine whether any of this investment pays off—is whether scientists trust the outputs enough to act on them.
That isn’t a question scientists can answer on their own; it’s a question for the executives funding these initiatives. Biopharma companies that want value from AI in drug discovery should focus on the fundamentals, including curated data foundations, validated scientific content, and systems that make every output traceable. When we get those right, AI becomes something scientists can move faster with, not something they have to slow down to verify. That’s when we’ll start to see the return on our industry’s investment.





