PubMed added more than 1.5 million citations in fiscal year 2023, roughly 4,300 a day, according to the National Library of Medicine. On top of that literature, medical affairs teams must absorb congress data, real-world evidence, and field insights from healthcare professionals (HCPs), and the volume makes it harder for pharma organizations to act on clinical evidence as quickly as it arrives.
AI can shorten synthesis time, but it brings its own failure modes. Researchers testing large language models (LLMs) on medical evidence summarization found the models could produce factually inconsistent and overly convincing summaries, with more errors on longer texts, according to a 2023 study in npj Digital Medicine. The U.S. Food and Drug Administration has since proposed a risk-based credibility framework for AI models used to support drug regulatory decisions.
Together, these findings point executives toward a narrower question: which decision should AI improve, and how will leaders know it did?
Emerj Senior Editor Yolandi de Weerdt explores that question with Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer, on the AI in Business podcast, in a conversation about what practical AI adoption looks like in evidence-heavy, regulated pharma workflows.
This article examines three core insights that matter most for leaders as pharma teams use AI to turn clinical evidence into action:
- Measure AI at the decision it improves: Work backward from the clinical or business decision AI should improve, and judge success by a downstream result the organization already tracks, such as site activation or data quality.
- Verify AI synthesis against primary sources: Check AI-generated summaries against the original material before they reach field reports or strategy decks, where a single conflated finding can compound.
- Fund the learning curve before counting gains: Build a training quarter and a supervised-use quarter into AI business cases for regulated, evidence-heavy work.
Listen to the full episode below:
Episode: Accelerating Evidence to Action in Pharma with Practical AI Adoption – with Nabil Khan of Pfizer
Guest: Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer
Expertise: Medical Affairs, Clinical Development, Clinical Trial Operations, Scientific Engagement
Brief Recognition: Nabil Khan is Medical Director for Internal Medicine Antivirals at Pfizer, where he gathers field insights from HCPs and carries them back to clinical development teams. Before joining Pfizer in 2024, he monitored clinical trial sites and supported studies in vaccines, cardiometabolic disease, ophthalmology, and neurology in senior clinical research roles at PPD, ICON, and IQVIA. He served as Director of Clinical Operations at Impact Physician Group. A physician with more than a decade in clinical research and operations, he holds a medical degree from Xavier University School of Medicine.
Measure AI at the Decision It Improves
In pharma, the evidence itself is rarely the problem; the delay comes in turning it into decisions. Nabil Khan sees that gap as one of the industry’s most consequential bottlenecks.
Medical affairs teams continuously review publications, congress presentations, real-world evidence, and field insights, and access to that material is rarely the constraint. The pressure builds, he says, because the ability to synthesize, analyze, and act on clinical and scientific information has not scaled at the same pace as the volume being generated.
In Khan’s experience, the bottleneck varies by seat. In the field, medical teams have to judge which HCP conversations yield clinically relevant insight. Inside medical affairs or clinical development, those insights arrive secondhand and have to be turned into strategy. On both sides, he says, the main struggle is telling noise apart from actual signal and evidence.
Khan identifies three ways AI changes that workflow. It speeds up synthesis: instead of spending weeks reviewing large volumes of publications, congress abstracts, field insights, and real-world data, teams can process that material much faster and give their time to interpretation and decisions. It detects emerging patterns and signals that are difficult to spot in manual review. And it creates a more consistent evidence base, improving alignment between medical affairs, clinical development, and commercial teams, which he calls a stabilizing feature. The goal, in his words, is to “allow the experts to spend less time gathering the information and more time acting on it.”
Asked what a senior leader should press on when a team says AI is helping it move faster, Khan compares AI adoption without a defined problem to owning a toolbox full of tools with no sense of which one fits the job. His starting point is the decision itself:
“Successful adoption of AI has to start with a clearly defined business problem, not necessarily the technology itself. So leaders should ask themselves and should ask their teams: What decision are we trying to improve on? What decision are we trying to accelerate? What decision are we trying to scale?”
— Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer
His answer points to four questions executives can put to any proposed AI evidence workflow:
- What decision are we trying to improve? With the business objective defined up front, Khan says, a team can see whether AI is pointing it in that direction or not.
- What evidence is the system working from? AI is only as reliable as the information it is fed, trained on, and analyzing, which makes data quality critical.
- Where do governance and expert oversight enter? Khan urges establishing governance early in a regulated industry to ensure transparency, compliance, and oversight, and keeping physicians and scientific experts from outside the company closely involved so outputs stay clinically relevant and actionable.
- Which downstream outcome will show it worked? In Khan’s view, success shows up only once the decision AI supported plays out in practice.
Trial site selection shows how far downstream that last answer can sit. AI can help identify potentially relevant HCPs in a region, but the test comes only after field medical teams vet the investigators and clinical development takes the sites through protocol training, regulatory review, compliance, and activation: are those sites producing clean data? In Khan’s account, only then can a team see whether the work upstream has succeeded.
Verify AI Synthesis Against Primary Sources
The speed and consistency that make AI attractive for evidence synthesis are also what make its errors costly. Khan puts the risk in terms of consistency:
“Sometimes consistency in the wrong direction is not a good thing. You think of consistency as being a positive, as being a good thing, but if you’re consistent in the wrong direction, obviously that’s not a good thing. The risks that I have seen are that if there is some sort of mistake or some sort of inconsistency, it will amplify that negativity.”
— Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer
His team uses AI heavily at medical congresses, both to plan limited time on site and to summarize the huge volume of material presented. What they receive, he says, has to be double-checked, because separate presentations sometimes end up merged in the output:
“Sometimes posters and presentations can somehow get amalgamated because their titles are very similar. And not only are their titles similar, but their topics can be similar. Maybe they even have the same speaker, but the same speaker is speaking at two different presentations. AI is taking that information and sometimes putting it together when it should not be put together. And then once it’s put together, when you are expanding on that information, it’s expanding upon incorrect information.”
— Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer
His operating principle is to treat AI as “a tool rather than a source,” and he returns to it at the close of the conversation: if AI is given poor or inaccurate information, it amplifies the misinformation, and in healthcare as on social media, amplified misinformation is very difficult to contain once it spreads.
Applied to his Congress example, the principle comes down to three checks before an AI summary moves to other teams:
- Traceability to primary material: Congress summaries, literature reviews, and evidence digests keep a clear route back to the publication, abstract, poster, or presentation behind each claim.
- Closer review for look-alike inputs: Similar titles, overlapping subject matter, and shared presenters trigger extra scrutiny of the synthesis.
- Named accountability: A designated human expert validates each AI-assisted synthesis before other teams build on it.
Fund the Learning Curve Before Counting Gains
In a regulated industry, access to an AI tool and readiness to use it on regulated work are two different things. Asked which step organizations most often skip when adopting AI, Khan names training, though he is not sure whether companies skip it outright or simply do not take it as seriously as they should.
In pharma, employees must learn where AI use intersects with regulatory requirements, compliance obligations, copyright, and external communications, in addition to learning the tool itself. He explains why a general-purpose technology creates that exposure:
“AI wasn’t created for one specific area. It wasn’t made just for pharma; it wasn’t made just for accounting or healthcare. It’s a general product, and so there are a lot of regulations, a lot of compliance issues, and a lot of copyright issues that people who are not used to dealing with those issues are not going to be good at. If they simply use AI to give them answers and to expand on topics for them, they could do that at the risk of being non-compliant and going against federal regulations.”
— Nabil Khan, Medical Director for Internal Medicine Antivirals at Pfizer
The exposure extends to the company, Khan adds: when an employee puts AI-generated content out that proves non-compliant, the organization carries the risk as well as the individual. The tools themselves differ too. ChatGPT, Claude, Copilot, and the many other systems available each behave a little differently, so teams need to learn how to use them in their own context, and using AI on a personal phone, he notes, is a different activity from using it on a work computer for regulated work.
Khan frames any timeline with a caveat: large organizations are still in the trial-and-error phase of AI implementation, working out where it works best. With that in mind, his sequence for the learning curve runs in three stages:
- Training, about one quarter (three to four months): Teams learn how the tools work, how to prompt them, and where regulatory and compliance limits apply.
- Supervised use, about another quarter: Employees apply AI to real work, see what succeeds and fails, and learn to judge whether an output is reliable or needs to be verified.
- Routine integration: AI moves into day-to-day workflows once teams can interpret its output with confidence.
The pace also depends on the use case. Summarizing or drafting emails, he notes, requires less readiness than summarizing a clinical or journal article or building a presentation from gathered medical evidence. For that evidence-heavy work, Khan’s minimum estimate is five to six months, and he expects meaningful adoption to take longer in practice.
