Feature|Videos|September 3, 2026

The Mechanics of Analyzing Linked Data Sets

Carelon Research's Mark Cziraky on why trusted research environments don't change how analysis is done — just how fast and richly it can begin.

The pharma and biotech industries run on data, especially on the R&D side. Data is one of the most valuable assets, but it can also cause some of the biggest headaches.

Due to the wide array of sources and collection methods, data sets may not always be easily compatible with others. Adding to the complications, big companies generally view data as proprietary and aren’t comfortable sharing it.

Things are only getting more complex with the rising acceptance of real-world data (RWD). In April of this year, David Lazerson and Fabio Lievano wrote for Pharmaceutical Executive about the rising importance of RWD. In their piece, they said, “FDA is changing the rules. Marking a fundamental shift in how we define acceptable clinical proof, it has indicated that one well-controlled trial, supported by confirmatory evidence––in particular, real-world evidence––will become the default for drug approvals. Many are applauding this as a sign of more regulatory flexibility.”

Pharmaceutical Executive recently spoke with Mark Cziraky, president of Carelon Research. The company just signed a deal with Manifold that allows life sciences researchers and data analysts to analyze Carelon Real World Data (RWD) alongside third-party datasets.

The deal is part of an ever evolving landscape in which data and how its controlled is becoming one of the most important tasks of any company. During his conversation, Cziraky discussed shifts in how RWD is perceived and how companies are viewing it as a different kind of asset.

Pharmaceutical Executive: What are the mechanics of analyzing linked data sets?
Mark Cziraky: The analytical process itself is largely the same whether you're working with data you've licensed and brought in-house or data you're accessing within a trusted research environment. What changes significantly is everything that happens before the analysis begins — the process of acquiring the data, bringing it in, and the level of quality and detail you can actually access.

There are meaningful restrictions on what data can be included when it's transferred in de-identified flat files, compared to what's available when data sits within a more governed environment. A trusted research environment can expand both the volume and the richness of data available to researchers, which ultimately leads to better real-world evidence.

Once you bring in the compute and begin the actual research, the work follows standard scientific practice. You identify whether the data can support the evidence you're trying to generate, run preliminary queries to inform study design, and then execute with the same rigor you'd apply in any environment — developing protocols, applying appropriate analytical methodologies, generating outputs, conducting quality assurance, and developing the evidence package.

At the end of the day, whether you're acquiring data in flat files and loading them into a data lake or accessing data within a governed environment, it's all about the evidence you're trying to generate and the insights that come from it. The model is secondary to the outcome.