Concept Architecture
Electronic Health Record
An electronic health record (EHR) is designed to support care over time, but its entries also become a source of information for service planning, quality measurement and research. This page follows the path from a clinical encounter to a recorded data item, then explains what has to be checked before those data can support health-economic estimates. A record's presence in an EHR is evidence of documentation in that system, not automatically evidence that an event occurred exactly as coded or that an unrecorded event did not occur.
What an EHR records during care
Clinicians and other authorised staff enter information as part of care delivery and associated workflows. An EHR can hold structured fields, narrative notes, laboratory results, medication orders, observations and dates, while the exact content and access vary by system and setting. The distinction between an order, dispensing, administration and actual use is especially important when studying medicines.
| EHR element | What it may show | What a researcher must check |
|---|---|---|
| Diagnosis code | A condition documented for an encounter or problem list. | Whether the code denotes a confirmed diagnosis, suspected condition, historical problem or billing convention. |
| Medication order | A clinician's prescribing intent. | Whether the medicine was dispensed, administered and taken; each is a distinct event. |
| Procedure or visit | A documented service at a recorded time and location. | Whether services delivered elsewhere are visible and whether duplicate records exist. |
| Laboratory result | A measured value and possibly units, method and reference range. | Whether dates, units and measurement context are comparable across sites. |
| Narrative note | Clinical detail that may not fit structured fields. | Whether extraction preserves negation, timing and uncertainty. |
An electronic medical record is sometimes used for a record within a provider's practice; an EHR often implies information intended to travel across care settings. Terminology varies in practice, so the exchange capability and actual data coverage should be described rather than inferred from the label. A patient-controlled personal health record is a different arrangement of stewardship and access.
How information moves between systems
An EHR cannot provide a complete longitudinal history simply because a person has received care somewhere. Sending data, matching the right patient, understanding a code and incorporating the information into a usable workflow are separate tasks. Interoperability therefore involves technical exchange, shared meaning and practical use, with permissions and governance at each handoff.
For a multi-site dataset, investigators should document source systems, patient linkage, care settings, coding systems, mapping versions and observation windows. A standard data model can help harmonise tables and terminology, but mapping does not restore events never captured by the source. Missing external care can make a person appear untreated, event-free or lost to follow-up when the data merely cease to observe them.
Turning care records into research evidence
The EHR is collected chiefly for operational and clinical purposes, so a research question needs a separate design. Specify an eligible population, an index date, exposure, comparator, follow-up, outcomes, covariates and the observation period before querying records. Document how each clinical concept becomes an operational rule, such as whether one diagnosis code suffices or confirmation from another encounter is required.
| Research choice | Why it changes the estimate | Practical check |
|---|---|---|
| Who enters the dataset | A system covers only people and care it can observe. | Describe enrolment or encounter requirements and likely care outside the system. |
| When follow-up begins | Future information used to define exposure can create immortal-time bias. | Align eligibility, treatment assignment and time zero. |
| What counts as an outcome | Documentation and coding vary across settings and over time. | Validate outcome definitions where feasible and inspect code changes. |
| What is missing | Absence of a value may depend on clinical need or access. | Report completeness and assess plausible missingness mechanisms. |
| How treatment is selected | Clinical severity and access influence treatment and outcomes. | Measure important confounders and assess residual confounding. |
An EHR-derived comparison of treated and untreated people is observational. Even after adjustment, it does not inherit randomisation; causal claims require assumptions about exchangeability, consistent measurement, time alignment and adequate overlap that should be tested where possible. A study should report sensitivity analyses for important alternative definitions and unobserved data limitations.
Using EHR data in economic evaluation
EHRs may help estimate event rates, resource use, treatment pathways and patient characteristics for a model. A recorded admission count can be an input to costing, but the EHR typically does not supply the full unit cost, patient payment, productivity loss or preference-based quality-of-life weight needed for every evaluation. The analyst must specify perspective, population, time horizon and any linkage to claims, tariffs, surveys or registries.
Consider a fictional hospital cohort of 200 people followed for one year. If 36 admissions are observed, the recorded admission rate is $36/200=0.18$ per person-year only if every person contributes a full year of observable follow-up; more generally, divide 36 by the sum of observed person-years. At an illustrative £2,400 per admission, the observed admissions contribute $36\times £2{,}400=£86{,}400$ in assigned admission costs, or £432 per enrolled person; these are not total care costs or proof that an intervention saved £432 per person.
The calculation assumes all 36 entries are distinct admissions and that the chosen unit cost applies to each. If another hospital's admissions are absent, the rate and assigned cost understate activity in the intended population; if 20 people are observed for only six months, the observed denominator is 190 person-years rather than 200, and the rate is $36/190\approx0.1895$ admissions per observed person-year. That adjustment fixes unequal observed time, but does not recover admissions outside the data source.
| Spreadsheet item | Illustrative input or formula | Result or interpretation |
|---|---|---|
| People enrolled | 200 | Defines the cohort, not the exposure time. |
| Observed admissions | 36 | Check for repeat or duplicate encounters. |
| Observed person-years | =180+20*0.5 | 190 when 180 contribute one year and 20 contribute half a year. |
| Admissions per person-year | =36/190 | Approximately 0.1895; label the observed denominator. |
| Assigned admission cost | =36*2400 | 86400 pounds for these admissions only. |
| Assigned cost per enrolled person | =86400/200 | 432 pounds; uses people, not person-years, as denominator. |
Checking fitness for a particular decision
Fitness depends on the decision and cannot be certified by a generic data-quality score alone. Compare the EHR's population, setting, dates and measured variables with the target population and required model inputs. Then check completeness, plausibility, consistency, duplication and changes in coding or workflow across sites and time.
- Provenance: Record where each variable came from, how it was transformed and which terminology or mapping version was used.
- Coverage: Identify care settings and people excluded from the system, including changes in capture over time.
- Measurement: Distinguish a clinical event from an order, a coded proxy and a manually abstracted outcome.
- Linkage: Evaluate patient-matching error, data permissions and the effects of unlinked records.
- Governance: Use the applicable legal basis, access controls, minimisation and disclosure safeguards for the jurisdiction and study.
- Reproducibility: Preserve cohort logic, code lists, extraction dates and checks so that another analyst can audit the estimate.
De-identification or pseudonymisation can reduce disclosure risk but does not itself establish permission for every secondary use or eliminate re-identification risk. Privacy, security, patient trust and representativeness are part of responsible data use, not an afterthought to a large sample. A large EHR cohort can yield a precise estimate of a biased quantity if data capture or study design is wrong.
Sources and further reading
The US Office of the National Coordinator's EHR and health information exchange FAQ explains the practical EHR, medical record and personal record distinction. The WHO classifications interoperability hub describes terminology support for clinical and secondary uses. OHDSI's description of observational research and its OMOP common data model documentation explain why care data require standardisation and quality checks. The cohort and cost numbers above are original illustrative calculations, not findings from these sources.
Related Concepts (2)
Frequently Asked Questions (6)
What is an electronic health record?
A digital record of a patient's medical history maintained by providers, capturing diagnoses, laboratory results, and treatment history in detail.
Source: Strom, Kimmel & Hennessy 2019
What clinical detail does an electronic health record hold?
An electronic health record is a digital account of a patient's care kept by providers, holding the clinical detail generated during treatment: diagnoses, laboratory results, medications, imaging, and clinicians' notes. This depth is what distinguishes it from billing sources, since it records not just that a service occurred but what was found and done, including values that reveal disease severity and response. For research it offers a rich, patient-level view of real care, though its records are shaped for treatment rather than analysis. Clinical richness is what it holds. Casey and colleagues (2016) describe this source.
Source: Casey et al. 2016
What does an electronic health record contain?
An electronic health record contains detailed clinical information generated in patient care, including diagnoses, symptoms, and problem lists; laboratory and test results; medications prescribed; procedures and treatments; clinical notes and observations; and measurements such as blood pressure and other readings, recorded over time. It may link to imaging and other data. This detail exceeds that of billing-based data. So an electronic health record contains rich, patient-level clinical data reflecting the care provided, offering laboratory values, clinical findings, and other detail that administrative and claims data lack, which makes it valuable but complex as a source of research data.
Source: Strom, Kimmel & Hennessy 2019
Why are electronic health records used in research?
Electronic health records are used in research because they contain detailed clinical information, such as laboratory results, diagnoses, and clinical measurements, that administrative and claims data lack, enabling richer studies of real-world care, effectiveness, and safety. They cover large populations over time and reflect routine practice. This clinical depth allows more precise definition of conditions and outcomes and study of clinical detail. So electronic health records are used for their combination of scale and clinical richness, providing real-world data with detail suited to studying the effects and use of treatments, complementing the breadth of claims data and the rigour of trials.
Source: Rothman, Greenland & Lash 2008
What are the challenges of using electronic health records for research?
The challenges of using electronic health records for research include variability and inconsistency in how data are recorded, since they are entered for care not research; missing data, as not all information is captured for all patients; the difficulty of extracting and standardising data from free text and different systems; incomplete capture of care occurring elsewhere; and confounding in observational analyses. Privacy and data governance also pose issues. These challenges mean electronic health record data require careful cleaning, validation, and methods to address missingness and confounding, and their findings interpreted with awareness of the limitations of data generated in routine care.
Source: Strom, Kimmel & Hennessy 2019
How do electronic health records differ from claims data?
Electronic health records are created for patient care and contain detailed clinical information, such as laboratory results, notes, and measurements, while claims data are generated for billing and capture coded diagnoses, procedures, and prescriptions but lack clinical detail. Electronic health records offer clinical depth, whereas claims data offer breadth across insured populations and capture billed care and costs. Electronic health records may miss care elsewhere, while claims may miss clinical detail. So the two differ in purpose, content, and coverage, with electronic health records providing clinical richness and claims data providing coverage of billed care, and they are sometimes linked for research.
Source: Strom, Kimmel & Hennessy 2019
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 24 Sep 2026
Content version: 1.0.0
Canonical Identity
- Term code
- HE-ES-ES-021
Stable URI · Machine-readable · Resolvable · CC BY 4.0