Concept Architecture
Concept
Theoretically, Best-Worst Scaling (BWS) is a stated preference method in which respondents identify the most and least preferred, most and least important, or most and least desirable items from a set of alternatives. It is based on random utility theory and discrete choice theory, providing more information per choice task than conventional rating or ranking methods while reducing many forms of response bias. In health economics, BWS is widely used to elicit preferences for health states, treatment attributes and healthcare resource allocation.
Mathematically, Best-Worst Scaling is represented using random utility models in which the probability of selecting the best and worst items depends on the difference between their underlying utility values. Preference weights are typically estimated using conditional logit, mixed logit or hierarchical Bayesian models.
In practice, Best-Worst Scaling is implemented by presenting respondents with repeated subsets of items generated using an efficient experimental design. Responses are analysed to estimate the relative utility or importance of each attribute or item, supporting health technology assessment, patient preference studies and health policy research.
Purpose
Used to quantify preferences, estimate relative attribute importance, derive utility weights, inform health technology assessment, support patient preference research, and evaluate healthcare decision-making.
Mathematical Formulae
Primary Formula
P(i best, j worst) = exp(U? ? U?) / ?? ???? exp(U? ? U?)
Supporting Formulae
U? = X?? + �?
BW Score = Best Count ? Worst Count
Related Mathematical Methods
- Random utility modelling
- Conditional logit model
- Mixed logit model
- Hierarchical Bayesian estimation
- Experimental design
Example
A respondent evaluates four treatment attributes:
- Effectiveness
- Side effects
- Travel time
- Cost
The respondent selects Effectiveness as the best attribute and Travel time as the worst.
Across 500 respondents:
- Effectiveness: Best = 360, Worst = 28 ? BW Score = 332
- Travel time: Best = 35, Worst = 290 ? BW Score = ?255
Model estimation converts these responses into preference weights that quantify the relative importance of each attribute.
Excel Implementation
| Function | Example Formula | Health Economics Application |
|---|---|---|
| COUNTIF | =COUNTIF(B2:B501,"Best")-COUNTIF(C2:C501,"Worst") | Calculates the Best-Worst score for an attribute. |
| SUMIFS | =SUMIFS(D:D,A:A,G2) | Aggregates best or worst selections for each attribute. |
| PivotTable | Attribute ? Count of Best/Worst | Summarises respondent choices prior to statistical modelling. |
VBA (Optional)
Automate the calculation of Best-Worst scores and preparation of respondent-level datasets for discrete choice model estimation.
Sources
- Louviere JJ, Flynn TN, Marley AAJ. Best-Worst Scaling: Theory, Methods and Applications. Cambridge University Press.
- Flynn TN, Louviere JJ, Peters TJ, Coast J. Best-Worst Scaling: What It Can Do for Health Care Research and How to Do It. Journal of Health Economics.
- Lancsar E, Louviere JJ. Conducting Discrete Choice Experiments to Inform Healthcare Decision Making. Pharmacoeconomics.
- ISPOR Conjoint Analysis Good Research Practices Task Force Reports.
Related Concepts (2)
Library
Publications
1
NICE DSU Technical Support Document 11: Alternatives to EQ-5D for Generating Health State Utility Values — Brazier, Rowen, TSD 11 ed., 2011 (NICE Decision Support Unit (University of Sheffield))
Guidance on alternatives to EQ-5D — including SF-6D, HUI, condition-specific preference-based measures, direct valuation and vignette methods — for generating health-state utility values.
Frequently Asked Questions (6)
What is best-worst scaling?
A stated preference method in which respondents identify both the most and least preferred items from a set, rather than ranking every item.
Source: Finn & Louviere 1992
How does best-worst scaling work?
Respondents are shown a set of items and asked to identify the most and the least preferred, rather than ranking everything or rating each separately. Each such answer identifies the pair with the greatest difference in the set, which yields more information than a single choice while asking less than a full ranking. Repeating the task across sets constructed so that every item appears with every other allows the relative position of all items to be estimated on a common scale.
Source: Finn & Louviere 1992
What are the three forms of best-worst scaling?
The first presents a list of objects and asks which is most and least important, which suits establishing priorities among many items. The second presents a single profile described by attribute levels and asks which level is best and worst, which yields values comparable across attributes on one scale. The third presents several complete profiles and asks which is best and which worst, which is closest to a conventional choice experiment and supports the same kind of analysis.
Source: Flynn et al. 2007
Why use best-worst scaling rather than rating or ranking?
Rating scales allow respondents to score everything highly, which reveals nothing about trade-offs, and they are used differently by different people, so scores are not comparable across respondents. Full ranking is cognitively demanding and produces unreliable answers in the middle of long lists. Identifying the extremes is the easiest judgement to make and the one respondents make most consistently, and repeating it generates enough information to position everything without requiring the difficult intermediate comparisons.
Source: Marley & Louviere 2005
Where is best-worst scaling used in health?
It is used to establish which aspects of care patients prioritise when many candidate aspects exist, which the object form handles better than a choice experiment could. It is used to value the levels of attributes within a health state description on a common scale, which supports the construction of preference-based measures. And it is used for priority setting exercises where the task is to order a long list of possible interventions or service features rather than to value them in money.
Source: Flynn et al. 2007
What are the limitations of best-worst scaling?
The object form yields relative importance and no monetary or health value, so it establishes an ordering without saying how much better one item is in any external unit. Results depend on which items appear in the list, since importance is measured relative to the set presented. Analysis rests on assumptions about how the best and worst choices relate to a single underlying scale, and those assumptions are not always tested. Evidence that the results predict real behaviour remains limited, as it does across stated preference methods generally.
Source: healtheconomics.wiki
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 31 Jul 2025
Content version: 1.0.0
Canonical Identity
- Persistent URI
- https://healtheconomics.wiki/concept/best-worst-scaling
- Term code
- HE-EE-CBA-005
Stable URI · Machine-readable · Resolvable · CC BY 4.0