Connecting RWD to Better Understand the Patient Journey
Key Takeaways
- Tokenization substitutes PHI with non-meaningful unique identifiers, enabling cross-database matching while protecting identity and enriching context when combining claims with EHR-derived clinical nuance.
- Multi-provider diseases such as NMIBC require linkage across primary care, specialty, pathology, oncology, labs, claims, and genomics to capture staging, treatment sequences, and outcomes longitudinally.
Why effective RWD linkage starts with a specific research objective — and how tokenization, clean rooms, and AI make it possible.
Real-world data (RWD) is an essential resource for finding insights and generating evidence in life sciences R&D, from clinical development to commercialization and clinical practice. In order to realize the potential of RWD, it’s becoming increasingly necessary to link data from disparate sources.
Yet linking health data isn’t a simple matter. In addition to patient privacy concerns, the fragmentation and lack of interoperability between health data systems present complex technical challenges.
In order to effectively connect and apply a diverse range of RWD, organizations need a clear understanding of fundamental data linkage issues and strategic approach tailored to specific research questions.
Unlocking a range of sources with tokenization
Individual sources of RWD typically lack the depth required to illuminate the complete patient journey. Understanding a patient’s history, comorbidities, treatments and outcomes typically requires integration across datasets, including claims, electronic health records (EHRs), images and genomic data.
Yet linking data must be done securely — without compromising patient privacy. One foundational tool for securely unlocking the value across a range of real-world data sources is tokenization.
Tokenization replaces protected health information (PHI), including patient name, social security number, address, date of birth, with unique identifiers. These tokens have no meaningful value aside from de-identifying the PHI. Once the token is generated, researchers can match tokens from different databases without exposing patient identities.
This linking of databases offers a richer view with more contextual information. For example, while claims data and physician notes in an electronic health record (EHR) both include the ICD code for a diagnosis, the claims data can confirm the official code and precisely what was billed for. EHR notes, on the other hand, complement this information with details about related conditions, symptoms that led to the diagnosis, and factors that provide a fuller picture of the patient experience.
With many medical conditions, the patient journey includes a number of stages. Gaining a clear picture of the case requires connecting data from a variety of sources. For example, with non-muscle invasive bladder cancer (NMIBC), a patient may have records with a primary care physician, urologist, pathologist, and oncologist. For a comprehensive view, it’s necessary to link data from all the providers as well as claims, labs and genomics.
Opportunities and challenges with data linking
Linking real-world data sources opens a range of opportunities to researchers, many which are obvious but some are counter-intuitive or just being realized.
Mike D’Ambrosio, a SVP who leads real-world research for Parexel, a global CRO and biopharmaceutical services company, cites the example of social determinants of health (SDOH) Z codes, used to document non-medical socioeconomic and environmental factors that profoundly shape health outcomes.
“I think it’s underestimated how valuable that can be within the linkage environment, particularly in disease indications that disproportionately affect either certain genders, races or patients who are on the lower end of health equity,” says D’Ambrosio.
Another opportunity D’AMbrosio cites is the potential to bridge the gap between the clinical trial findings and how to practically implement a new therapy into a care setting.
“One example of what we’ve been looking at is tokenizing clinical trial failure patients, which is a bit of a novel concept for most people,” says D’Ambrosio. “We tend as an industry to kind of almost discard that cohort of patients because they're not the perfect match to gain the best kind of outcome. So it's an interesting concept to follow them as a kind of surrogate, a kind of quasi real world population, particularly around how in a more dynamic fashion an indication is treated.”
While data linkage offers unending potential possibilities to explore, there are significant challenges. The fragmentation of data in various formats and systems, without interoperability, makes connecting information effectively a puzzle.
Dr. Sarah Matt, a physician and health tech strategist, points out that EHR data is something researchers want connected to everything, but the reality is it’s a challenging hurdle.
“We still have great difficulty connecting electronic medical record data with everything else,” says Dr. Matt. "There are great standards in the world… but we're not all using the same versions. Oftentimes we’ll have hospitals within a system that are not in the same pipeline as the others. We need to be more interoperable at baseline with more standards.”
A broader view: data ecosystems
At core, researchers utilizing RWD must ensure data is fit-for-purpose, representative, and complete enough to answer specific research objectives — while preserving patient privacy.
While tokenization has become a valuable tool, it alone is not a solution to the fragmentation problem.
Praveen Haridas leads the real-world data and solutions efforts for Amazon Web Services, and has been talking with life sciences customers and healthcare providers to understand the challenge from the technology side.
“How do we drive connectivity in a privacy preserving way?” asks Haridas. “Tokenization
is a component to it, linkage is a component to it, but it's more about collaboration – this is a data collaboration problem, if you will. Our priority in solving it is building connectivity in the ecosystem.”
Haridas points to the creation of virtual clean rooms or clean room frameworks, engaging multiple health systems or data producers, and allowing for data collaboration and analysis — while preventing unnecessary movement and ensuring data remains under the control of the producer.
“We've been working to expand the clean room technology to meet the scale,” says Haridas. “We are potentially looking at 50 terabytes of data. Pharma companies are looking at how they can link their trial data into the RWD sources without revealing the identity, and then do the analysis without any party moving the datasets.”
Employing AI
In recent years, as pharmaceutical and biotech companies have generated enormous amounts of clinical trial and launch data and gained expanded access to RWD, they’re increasingly turning to AI workflows to search for actionable insights to improve decision-making.
AI has significant potential in data standardization, automating the conversion of different data formats, and supporting evidence generation.
“From a life science perspective, when we think about AI and the data we already have, we're seeing more and more synthetic data and digital twins being utilized as a foundation for what experiments we're going to do next, how we're going to make the next clinical trial,” says Dr. Matt. “This is an area where AI can absolutely be a good spot to utilize tokenized data and deidentification.”
Yet it is important that AI solutions are safely and responsibly implemented, as patients and providers are often skeptical of "black box" solutions. Success requires a "human-in-the-loop" approach, clear guardrails, and focusing on solving pervasive, urgent problems rather than just applying technology for the sake of marketing.
Strategic data linking
What’s the best approach to the open-ended, infinite realm of data linkage? Just because data can be linked, doesn’t mean it’s worthwhile.
The best practice for data linkage is to start with a clear, specific research objective.
According to D’Ambrosio, we’ve moved past the initial hype period of tokenization, when the thought was to tokenize everything and hope for the best.
The key, he says, is “a more thoughtful approach in terms of integrated evidence planning” that convenes disparate stakeholders and examines a multiyear timeline, asking “What evidence are we going to need at different points throughout that kind of journey?”
By working backward from specific goals, organizations can avoid unnecessary tokenization efforts, and ensure that the chosen data sources and linkage strategies directly support required insights.
The potential rewards of effective data linkage are enormous. Indeed, RWD already plays a key role in gaining more comprehensive views of the patient journey and evidence generation in life sciences research. With a strategic approach that focuses effort on salient research tasks and requirements, organizations can efficiently open pathways to connect RWD in ways that uncover critical insights and improve patient lives.
Related to this article








