Real-world evidence and AI: How EHR data is reshaping drug development decisions
16:59 · July 22, 2026 · RSS APP - AI Primary Research

Real-world evidence AI drug development is moving from pilot projects into routine use as EHR AI tools mature. Explore how sponsors and regulators are adapting.
Summary
Real-world evidence derived from electronic health records is shifting from regulatory aspiration to operational practice in drug development as natural language processing and machine learning tools mature enough to handle the volume and unstructured nature of clinical text. Sponsors now apply these methods to extract diagnoses, medication changes, and adverse events from physician notes and reports that structured fields alone cannot capture, turning raw real-world data into analysis-ready conclusions that inform decisions across the development lifecycle.
The distinction between real-world data and real-world evidence remains central: the former consists of messy, often free-text records collected outside randomized trials, while the latter requires validated study designs and extraction pipelines to produce regulatory-grade findings. Transformer-based models adapted for clinical text have become common for this extraction work because they handle long, jargon-heavy passages more reliably than earlier keyword or rule-based systems. Validation of these pipelines stays a critical and frequently scrutinized step, since unvalidated outputs can distort downstream trial planning or safety assessments.
Applications now extend to early trial design, where real-world data helps estimate incidence rates, refine eligibility criteria, and construct external control arms, particularly in oncology and rare-disease settings where traditional comparators are impractical. Machine learning models also support patient stratification by identifying response patterns across genomic, clinical, and treatment variables drawn from large EHR cohorts, though demographic or geographic biases in the source data require sensitivity checks before results influence protocols. Post-approval, the same NLP techniques scan narratives for adverse-event signals that claims codes often miss, feeding pharmacovigilance systems with earlier detection.
Regulatory agencies have responded with structured pathways rather than treating AI-derived evidence as a wholesale replacement for randomized trials. The FDA’s March 2024 draft guidance on non-interventional studies outlines expectations for observational designs intended to support effectiveness or safety claims, building on earlier mandates from the 21st Century Cures Act. The EMA operates DARWIN EU, a federated network spanning roughly 250 million patients, to standardize data quality and generate studies that inform both pharmacovigilance and broader regulatory review. Both agencies emphasize documentation and validation of extraction methods without imposing categorically different standards for AI-assisted versus manual curation.
Data quality, representativeness, and acceptance criteria continue to limit how far real-world evidence can travel toward labeling decisions. Sponsors that invest early in validated pipelines and documented models are better positioned to integrate these sources routinely rather than as one-off supplements.
Why it matters
This article is highly relevant for Dutch AI researchers in healthcare and pharma, as it details the European Medicines Agency's (EMA) DARWIN EU network and regulatory stances on AI-extracted data. It provides actionable insights into deploying transformer-based NLP models for EHR mining within the EU regulatory context.











