VerifiedEvidence: highv1.0.0

State Reward

The cost or health outcome value assigned to a health state within a Markov model, accrued for each cycle a patient occupies it.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Concept

Theoretically, State Reward is the numerical value assigned to occupying a particular health state during a model cycle within a state-transition model. A state reward represents the expected quantity accumulated while remaining in that state, such as healthcare costs, health utilities, life years or quality-adjusted life years. In health economics, state rewards enable model outcomes to be calculated by combining state occupancy with state-specific values over time.

Mathematically, state rewards are represented as a reward vector associated with the health states of a Markov or state-transition model. At each cycle, the expected reward is obtained by multiplying the state occupancy vector by the reward vector. Cumulative rewards are calculated by summing cycle-specific rewards across the model time horizon, with discounting applied where appropriate.

In practice, state rewards are estimated from economic evaluations, clinical studies, health utility instruments, administrative databases or published literature. Separate reward vectors are commonly specified for costs and health outcomes. During model execution, rewards are accumulated at each cycle to estimate discounted lifetime costs, life years, quality-adjusted life years and other measures required for economic evaluation.


Purpose

Used to assign costs, utilities or other outcomes to health states, enabling calculation of cumulative health and economic outcomes throughout a state-transition model.


Mathematical Formulae

Primary Formula

Expected reward during cycle t:

R? = ?????

where:

  • ??? = state occupancy vector at cycle t
  • ?? = vector of state rewards
  • R? = expected reward during cycle t

Supporting Formulae

Cumulative reward:

R = ????? R?

Discounted cumulative reward:

R = ????? R? / (1 + d)?

where:

  • d = discount rate
  • T = total number of model cycles

Related Mathematical Methods

  • Markov modelling
  • State-transition modelling
  • Matrix algebra
  • Cohort simulation
  • Discounting
  • Probabilistic sensitivity analysis

Example

A Markov model contains three health states with annual utility rewards:

Health StateUtility Reward
Stable Disease0.88
Progressive Disease0.60
Death0.00

If the cohort occupancy vector after one cycle is:

??? = [0.80, 0.15, 0.05]

the expected utility reward is:

R? = (0.80 ? 0.88) + (0.15 ? 0.60) + (0.05 ? 0.00) = 0.794

Thus, the cohort accumulates 0.794 quality-adjusted life years during that cycle before discounting.


Excel Implementation

FunctionExample FormulaHealth Economics Application
SUMPRODUCT=SUMPRODUCT(StateVector,RewardVector)Calculate expected state reward for each cycle
SUM=SUM(CycleRewards)Calculate cumulative rewards across the model horizon
PV=PV(DiscountRate,Cycle,0,-Reward)Apply discounting to future rewards where appropriate
MMULT=MMULT(StateVector,TransitionMatrix)Update state occupancies before calculating rewards

VBA (Optional)

Automate calculation and accumulation of discounted state rewards for multiple model strategies and simulation runs.


Sources

  • Puterman ML. Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley; 1994.
  • Sonnenberg FA, Beck JR. Markov models in medical decision making: a practical guide. Medical Decision Making. 1993;13(4):322?338.
  • Siebert U, Alagoz O, Bayoumi AM, et al. State-transition modeling: a report of the ISPOR-SMDM Modeling Good Research Practices Task Force-3. Medical Decision Making. 2012;32(5):690?700.
  • Briggs A, Claxton K, Sculpher M. Decision Modelling for Health Economic Evaluation. Oxford University Press; 2006.
  • NICE. Health Technology Evaluation Manual.
  • Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes. 4th ed. Oxford University Press.

Library

Publications

1
  • Journal article

    State-Transition Modeling: A Report of the ISPOR-SMDM Modeling Good Research Practices Task Force-3 — Siebert, Alagoz, Bayoumi, Jahn, Owens, Cohen & Kuntz, Task Force Report 3 ed., 2012 (Value in Health / Medical Decision Making)

    Best-practice guidance for cohort and individual-based state-transition (Markov) models, covering development, analysis, validation and reporting.

Frequently Asked Questions (6)

  • What is a state reward?

    The cost or health outcome value assigned to a health state within a Markov model, accrued for each cycle a patient occupies it.

    Source: Sonnenberg FA, Beck JR. Markov models in medical decision making: a practical guide. Medical Decision Making. 1993;13(4):322-338. doi:10.1177/0272989X9301300409.

  • What is an example of a state reward?

    A state reward is the value earned for spending time in a health state, applied for each cycle a patient remains there. Examples are the yearly cost of managing a chronic condition and the quality-of-life utility attached to living in that state, both accrued cycle by cycle for as long as the patient occupies it. Summing these rewards over the time spent in each state, across all states, builds the model's total cost and health. They are the running values of the model. Briggs and colleagues (2006) describe state rewards.

    Source: Briggs et al. 2006

  • How are state rewards accrued?

    State rewards are accrued by applying, each cycle, the reward value of the state a patient occupies, so that occupying a state for a cycle adds its per-cycle cost and health value. In a cohort model, the reward for a state each cycle is its value times the proportion of the cohort in it, summed across states and cycles. Over the horizon, these accruals accumulate into the total costs and effects. Rewards are thus added cycle by cycle according to where patients are.

    Source: Sonnenberg & Beck 1993

  • What do state rewards represent?

    State rewards represent the per-cycle consequences of occupying a health state: the cost reward is the cost of care and resources associated with the state per cycle, and the health reward is the quality-of-life value, or utility, of the state, which contributes to quality-adjusted life years. Some states, such as death, may carry zero or minimal rewards. By attaching these values to states, the model quantifies the costs and health effects of the conditions the states represent as patients experience them over time.

    Source: Sonnenberg & Beck 1993

  • How do state rewards differ from transition rewards?

    State rewards are accrued for each cycle a patient occupies a state, representing the ongoing per-cycle cost or health value of being in that state, whereas transition rewards are applied once at the moment a patient moves between states, representing a one-off cost or effect of the transition itself, such as the cost of an event. State rewards accumulate with time in a state; transition rewards occur at the point of movement. A model may use both to capture ongoing and one-off consequences.

    Source: Sonnenberg & Beck 1993

  • How do state rewards contribute to model outcomes?

    State rewards contribute to model outcomes by accumulating over the cycles patients occupy states: the total cost from a state is its cost reward times its occupancy, and the total health effect is its health reward times its occupancy, summed across states. These accumulated rewards give the model's total costs and quality-adjusted life years, which are compared between options. Because outcomes are built from the rewards accrued over state occupancy, the state rewards are among the key inputs determining the results.

    Source: Sonnenberg & Beck 1993

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 7 Oct 2025

Content version: 1.0.0

Canonical Identity

Term code
HE-EM-MM-019

Stable URI · Machine-readable · Resolvable · CC BY 4.0