VerifiedEvidence: highv1.0.0

Real-World Data

Routinely collected information about patient health or healthcare delivery that may support research and decision making when fit for a defined purpose.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Real-World Data

Real-world data (RWD) are information about health status or healthcare delivery collected in the course of routine practice or other nontraditional research settings. Electronic health records, claims, registries and some digital-health sources can contain RWD. An analysis of suitable data can produce real-world evidence (RWE), but a large dataset is not evidence of a treatment effect merely because it exists.

Where these data come from

Different sources record different parts of the care pathway for different purposes. A clinical record may contain laboratory results and symptoms, while a claim documents billed services and enrollment. The original reason for collection shapes what is missing and how a variable should be interpreted.

SourceOften useful forCommon blind spot
Electronic health recordDiagnoses, test results, prescribing, encounters and some clinical outcomes.Care outside the network, unstructured notes, incomplete treatment taking and changing documentation.
Insurance or administrative claimsBilled services, dispensing, covered costs and longitudinal enrollment within a payer.Clinical severity, services not billed or covered, and medications prescribed but not filled.
Disease or product registryDefined populations, treatments and condition-specific outcomes.Selective site participation, incomplete capture and variable follow-up.
Patient-reported or digital sourceSymptoms, functioning, experience and measurements outside appointments.Device access, engagement, missingness and differences in measurement.
Linked sourcesCombining clinical detail, utilization and outcomes across systems.Linkage error, incomplete coverage, incompatible dates and privacy constraints.

The label “real-world” describes the data-generating setting, not a guarantee that the sample represents everyone receiving care. Some randomized or pragmatic studies can use routinely collected data for outcomes; “RWD” does not automatically mean “observational,” and “trial data” and “RWD” need not be mutually exclusive in every design. The study design and the origin of each variable should be reported separately.

From a decision question to fit-for-purpose data

A dataset is useful only in relation to a specified question. A study of treatment uptake needs reliable dates and eligibility; a comparative-effectiveness study also needs exposure, outcome, confounder and follow-up information; a costing study needs a perspective and relevant resource or payment fields. Data can be excellent for one purpose and unsuitable for another.

  1. Specify the target question. Define population, treatment strategies, comparator, outcome, time origin, follow-up, setting and target measure before inspecting apparent effects.
  2. Map the care pathway. Identify how people enter the source, where treatment occurs, which encounters are captured and where records may leave the system.
  3. Assess relevance. Check whether the source covers the population, geography, time period, interventions, outcomes and resource use needed for the decision.
  4. Assess reliability. Validate definitions, coding, dates, completeness, consistency, linkage and changes in capture or practice over time.
  5. Design the analysis. Align eligibility, treatment assignment and start of follow-up; address confounding, selection, censoring and missing data as the question requires.
  6. Report limits. Show the cohort flow, operational definitions, data-quality findings, assumptions and remaining uncertainty, including what the source could not observe.

Relevance asks whether the right information exists for the decision; reliability asks whether the recorded information is sufficiently accurate and complete. Both matter. Privacy permissions, lawful data access and ethics review must also be resolved for the intended use; availability to an analyst is not itself authorization.

Worked example: a crude association changes with case mix

Imagine 1,000 patients on treatment A and 1,000 on treatment B in a routine-care dataset. There are “higher-risk” and “lower-risk” patients, and both treatments have the same observed outcome rate within each risk group: 10% in higher-risk and 2% in lower-risk patients. Treatment A is used much more often in the higher-risk group. These figures are constructed to show how the overall association can differ from the within-group comparisons.

Treatment and risk groupPatientsEventsObserved event rate
A, higher risk6006010%
A, lower risk40082%
A, all patients1,000686.8%
B, higher risk2002010%
B, lower risk800162%
B, all patients1,000363.6%

The crude A-minus-B risk difference is $6.8%-3.6%=3.2$ percentage points. If each treatment's observed risk is standardized to a common population that is 50% higher risk and 50% lower risk, both become $0.5(10%)+0.5(2%)=6%$, yielding a standardized difference of zero. The different crude rates in this example reflect the specified risk-group composition, not a within-group rate difference.

Standardization does not prove the treatments are equally effective. There could be unmeasured severity within the two groups, different follow-up, outcome recording or treatment adherence. A causal interpretation requires defensible design and assumptions about comparability, measurement and selection. The worked data are illustrative, not a study or recommendation to switch treatment.

Biases that routine records can introduce

Treatment choice in practice is often influenced by prognosis, clinician judgment, preferences and access. Confounding by indication arises when the reason a patient receives a treatment also predicts their outcome. Missing a key severity variable cannot be cured by adding more rows of otherwise similar data. Risk adjustment, matching, weighting or other methods rely on measured covariates and their own assumptions; none automatically recreates randomization.

Time alignment is equally important. If a patient is labeled “treated” from diagnosis but only starts therapy weeks later, the period during which they had to survive to receive it can create immortal-time bias. Define eligibility and time zero, assign strategies coherently, and handle treatment changes and grace periods with an appropriate design. A target-trial emulation can make these design choices explicit, but calling a study an emulation is not a substitute for checking identification assumptions.

Other hazards include outcome misclassification, missing data related to health status, incomplete capture when patients change insurers or hospitals, duplicate records, coding incentives, linkage errors and changes in technology or clinical practice. Patient-reported outcomes and health-state utilities may be absent from claims even when costs and utilization are well observed. Censoring at loss of enrollment is not automatically noninformative. Analyze how each limitation could change the direction and magnitude of the result, rather than listing it as a generic caveat.

Uses in health economics and technology assessment

RWD can inform epidemiology, baseline event risk, treatment patterns, adherence, resource use, costs, patient characteristics and longer-term outcomes. These inputs can make an economic model more relevant to a local population, but their definitions must align with the model's states, cycle timing, perspective and price year. A utilization count is not itself a cost; the analyst must apply appropriate unit costs and avoid double counting.

Comparative-effectiveness estimates from RWD may complement trial estimates when the question concerns patients or practice patterns underrepresented in trials. Transporting trial effects to a real-world population also requires care: a change in baseline risk can change absolute benefit even with a stable relative effect, while a change in actual treatment effect needs separate evidence. For a decision beyond observed follow-up, RWD can support survival extrapolation, but unobserved late outcomes and selection into follow-up still limit certainty.

Data and evidence occupy distinct steps. RWD are the recorded observations; RWE is an inference derived through an explicit design and analysis. A real-world cost estimate, safety signal and causal treatment-effect estimate have different evidentiary demands. Neither the FDA's nor NICE's frameworks imply that every observational estimate is sufficient for a particular regulatory or reimbursement decision.

Reproducibility, governance and reporting

Describe the provenance of each source, its collection purpose, database versions, covered population and dates, coding systems, linkage, cleaning and analysis. Publish operational definitions and a cohort flow where lawful, including counts lost at each eligibility step. Pre-specify or transparently distinguish planned from exploratory analyses, then report outcome validation, missingness and sensitivity analyses relevant to the claim.

  • Check time zero. Eligibility, treatment classification and follow-up must align so that no group receives guaranteed event-free time by definition.
  • Check ascertainment. Confirm whether care outside the system and outcomes after loss of coverage can be seen.
  • Check denominator and selection. Explain who appears in the data and whose results the analysis can represent.
  • Check measurement. Validate critical codes and dates, especially treatment initiation, outcomes, covariates and costs.
  • Check causal claims. State which confounders were measured, what assumptions remain and how alternative plausible choices affect the estimate.
  • Check rights and access. Follow the applicable privacy, consent, data-use and ethics requirements; de-identification labels alone do not settle every obligation.

Sources and further reading

The U.S. FDA's real-world evidence page defines RWD and distinguishes it from evidence derived through analysis. The NICE real-world evidence framework gives a structure for assessing data suitability and conducting quantitative real-world studies. The European Medicines Agency's real-world evidence page places routinely collected data in regulatory context. The two-treatment counts are an original teaching example, not evidence about actual treatments.

Library

Publications

1
  • Journal article

    Good Practices for Real-World Data Studies of Treatment and/or Comparative Effectiveness: Recommendations from the Joint ISPOR-ISPE Special Task Force on Real-World Evidence in Health Care Decision Making — Berger, Sox, Willke, Brixner, Eichler, Goettsch, Madigan, Makady, Schneeweiss, Tarricone, Wang, Watkins & Mullins, Vol. 20, No. 8 ed., 2017 (Value in Health)

    The joint ISPOR-ISPE recommendations on good procedural practice for real-world data studies (observational studies and registries) used to inform healthcare decisions — study registration, replicability and stakeholder involvement — the reference for RWE credibility in HTA.

Frequently Asked Questions (6)

  • What is real-world data?

    Data on patient health status or healthcare delivery routinely collected outside a traditional clinical trial, such as records, claims, and registries.

    Source: Makady et al. 2017

  • Where does real-world data come from?

    Real-world data is health information collected in the course of ordinary care rather than within a controlled trial, drawn from sources such as electronic health records, insurance claims, disease registries, and increasingly wearable devices. It comes from patients as they actually live and are treated, capturing the broad, unselected population that trials leave out. This breadth is its value, though the data were recorded for care or billing rather than research, so they are often messy and incomplete. Care as it happens is its origin. Sherman and colleagues (2016) describe this.

    Source: Sherman et al. 2016

  • What are the main sources of real-world data?

    The main sources of real-world data include electronic health records from clinical care; administrative and insurance claims data generated by billing and reimbursement; disease and product registries that collect standardised data on defined patient groups; and, increasingly, patient-generated data from apps, wearables, and surveys. Each source captures different information and has different strengths and limitations. So real-world data comes from a range of routinely collected sources, chiefly records, claims, and registries, together with newer digital sources, which between them cover different aspects of healthcare and patient experience in ordinary practice, and are often combined through linkage to give a fuller picture.

    Source: Makady et al. 2017

  • How is real-world data used?

    Real-world data is used to study how treatments are used and how they perform in routine practice, addressing questions such as long-term effectiveness and safety, use in populations underrepresented in trials, treatment patterns, and healthcare use and costs. Analysed appropriately, it generates real-world evidence to inform regulation, health technology assessment, and clinical decisions. So real-world data is used to complement trial evidence by providing information from ordinary care, supporting decisions where trials are lacking or where real-world performance matters, though its use requires methods to address the biases inherent in data not collected for research under controlled conditions.

    Source: Makady et al. 2017

  • What are the strengths and limitations of real-world data?

    The strengths of real-world data include large, broad populations reflecting routine practice, long follow-up, and relevance to real-world use, at lower cost than dedicated trials; its limitations include that it is not collected for research, so it may have missing or inconsistent information, measurement error, and, crucially, confounding, since treatment is not randomised. So real-world data offers scale and relevance but requires careful methods to address its quality issues and confounding, and its findings are interpreted with attention to these limitations, since the absence of randomisation and the nature of routinely collected data mean bias is a central concern in drawing conclusions from it.

    Source: Makady et al. 2017

  • How does real-world data relate to real-world evidence?

    Real-world data relates to real-world evidence as the raw material to the conclusion: real-world data is the routinely collected information, and real-world evidence is the clinical evidence about a treatment's use, benefits, or risks derived from analysing that data. In other words, real-world data is the input, and real-world evidence is the output produced by studying it. So real-world data and real-world evidence are distinct but linked, with the quality of the evidence depending on both the quality of the underlying data and the rigour of the analysis, since sound real-world evidence requires appropriate data and appropriate methods applied to it.

    Source: Sherman et al. 2016

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 24 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-ES-RWE-011

Stable URI · Machine-readable · Resolvable · CC BY 4.0