Concept Architecture
How a choice set creates an observable decision
A choice set is the group of alternatives presented together in one stated-choice task. Respondents compare the alternatives as complete profiles and select, rank, or otherwise evaluate one option. This page explains how alternatives, attributes, levels, opt-out options, and experimental design determine what preferences and trade-offs can be inferred from the observed choice.
Choice set, choice task, and experiment are different
A choice set is the alternatives available at one decision point. A choice task includes that set plus the question, instructions, framing, and response format. A discrete choice experiment usually contains several tasks so that preference parameters can be estimated from repeated choices across systematically varied profiles.
| Element | Meaning | Example |
|---|---|---|
| Alternative | One option within the set | Treatment A |
| Attribute | A characteristic on which options differ | Effectiveness |
| Level | The value assigned to an attribute | 70% response |
| Choice set | All alternatives shown in one task | Treatment A, Treatment B, and no treatment |
| Choice task | The choice set plus the decision question and presentation | “Which option would you choose?” |
| Experimental design | The planned allocation of profiles across tasks and respondents | A blocked efficient design |
Alternatives can be labelled or unlabelled
Labelled alternatives have meaningful names such as surgery, medicine, or watchful waiting. Unlabelled alternatives use generic labels such as Option A and Option B so that choices are driven mainly by shown attributes. Labels can improve realism but can also carry unmeasured beliefs, familiarity, or stigma.
The choice should match the real decision. A labelled design should measure or acknowledge important information respondents attach to the label, while an unlabelled design should not remove distinctions that make options clinically different.
Attributes define what respondents can trade
Attributes should represent features that matter to respondents, vary across realistic options, and can be described independently enough for meaningful choice. They are usually developed from literature, qualitative research, expert input, and direct engagement with the target population. Important omitted attributes can distort the values placed on those included.
Good attributes are:
- Relevant to the decision and understandable to respondents.
- Capable of varying across alternatives or policies.
- Distinct enough to avoid double counting the same concept.
- Plausibly influenced by the decision maker or intervention.
- Described without embedding value-laden or unequal wording.
- Supported by levels that cover meaningful trade-offs.
Levels must be realistic and sufficiently different
Levels define the values an attribute takes across alternatives. They must be plausible jointly as well as individually and broad enough to reveal preference gradients. Narrow levels can make trade-offs invisible, while extreme levels can create obvious choices or unrealistic scenarios.
Quantitative levels should use consistent units and time horizons. Risk should normally include the denominator and period, and cost should specify who pays, how often, and in which currency and price context.
The status quo can anchor the task in reality
A status-quo or current-care alternative represents what happens if no new option is selected. It can improve realism and allow estimation of willingness to switch. It can also carry familiarity, inertia, entitlement, and unmeasured characteristics that must be represented through an alternative-specific constant or richer specification.
The status quo should describe the respondent's actual or credible baseline. A generic “current care” option is weak when current care differs materially across people or settings.
Opt-out and no-choice options affect welfare interpretation
An opt-out allows respondents to select none of the presented active alternatives. It can reduce forced choices and better represent voluntary decisions. In settings where treatment or policy selection is unavoidable, an opt-out can be unrealistic and may alter the trade-offs being studied.
Researchers should distinguish:
- No treatment or no programme.
- Current care or status quo.
- Defer the decision.
- Choose another unlisted option.
- Cannot decide or task non-response.
These responses have different behavioural meanings and should not be combined automatically.
Random utility connects choice sets to preference estimates
Discrete choice models commonly assume that respondent (n) chooses the alternative with the greatest utility in choice set (C_{nt}). Utility has an observed systematic component and an unobserved component. The chosen alternative reveals relative preference under the model rather than directly measuring utility on an absolute scale.
$$ U_{njt}=V_{njt}+\varepsilon_{njt} $$
A linear systematic utility function can be written as:
$$ V_{njt}=\boldsymbol{\beta}^{\top}\mathbf{x}_{njt} $$
where (\mathbf{x}_{njt}) contains the attribute levels for alternative (j). Coding, interactions, nonlinearities, and alternative-specific constants determine the interpretation of (\boldsymbol{\beta}).
Conditional logit expresses choice probability
Under the multinomial logit assumptions, the probability that respondent (n) chooses alternative (j) from set (C_{nt}) is:
$$ P_{njt}=\frac{\exp(V_{njt})}{\sum_{k\in C_{nt}}\exp(V_{nkt})} $$
This model implies independence of irrelevant alternatives and homogeneous preferences unless extended. Mixed logit, latent-class, nested, or scale-adjusted models may be more appropriate when preferences or error variance differ.
Attribute-level coding changes interpretation
Categorical levels can use dummy coding, effects coding, or another planned scheme. Effects coding commonly sets the omitted level to the negative sum of estimated levels, allowing coefficients to be interpreted relative to the grand mean. Continuous coding assumes a functional form that should be tested.
Codebooks should preserve the displayed level, analytical code, reference, and expected direction. Changing coding after seeing results can create confusing or selective interpretation.
Experimental design creates identification
The experimental design selects which profiles appear together and across tasks. It should permit estimation of the intended main effects and interactions while controlling respondent burden. Full factorial designs are usually too large, so fractional factorial or statistically efficient designs are common.
Design evaluation includes:
- Level balance and overlap.
- Correlation among attribute levels.
- Statistical identification and expected standard errors.
- Number of alternatives, tasks, blocks, and respondents.
- Prior parameter assumptions used for efficient designs.
- Restrictions needed to prevent impossible combinations.
Efficiency is not the only design criterion
A statistically efficient design can still be cognitively difficult, unrealistic, or dominated. Respondents may use shortcuts when tasks are too complex, weakening data quality. Design optimisation should therefore combine statistical properties with qualitative testing and piloting.
D-efficiency is often based on the determinant of the parameter covariance matrix. It is meaningful only for the specified model and priors; it does not establish content validity or respondent comprehension.
Dominant alternatives can teach or distort
An alternative dominates another when it is at least as good on every valued attribute and better on at least one. Dominant tasks can test attention or help respondents learn, but repeated obvious choices provide little information about marginal trade-offs. Apparent dominance can disappear when omitted features or preference heterogeneity are considered.
Researchers should not automatically remove respondents who choose a dominated option. The response may reveal misunderstanding, non-compensatory preference, attribute non-attendance, mistrust, or a design error and should be investigated.
Implausible combinations undermine validity
Independent variation can produce profiles that cannot exist clinically or operationally. Restrictions can remove these combinations, but excessive restrictions can create correlation and reduce identification. The design should distinguish impossible, merely uncommon, and unfamiliar profiles.
Qualitative interviews and expert review should test whether respondents can imagine each complete alternative. Explanatory text should not repair a fundamentally incoherent profile.
Cognitive burden increases with complexity
Burden depends on the number of alternatives, attributes, levels, tasks, numerical concepts, and unfamiliar trade-offs. More information can improve realism while reducing attention and consistency. Accessibility, literacy, language, numeracy, disability, and fatigue should shape the design.
Possible safeguards include training tasks, plain language, visual aids, progressive familiarisation, blocking, hover definitions, interviewer support, and fewer attributes. Visual presentation must remain neutral and consistent across alternatives.
Order and position can affect choices
Respondents may favour the first, leftmost, default, or visually prominent alternative. Attribute order can also change attention. Randomisation or systematic rotation can reduce confounding, while fixed clinically meaningful order can improve comprehension when justified.
The survey system should record the order actually displayed. Analysis can test position effects and learning or fatigue across tasks.
Repeated tasks create panel data
Each respondent usually completes several choice tasks, so observations from the same person are correlated. Treating every choice as independent can understate uncertainty. Models or standard errors should account for clustering and repeated choice.
The number of tasks per respondent balances information against fatigue. More tasks do not always yield more useful information if later responses become inattentive or heuristic.
Preference heterogeneity matters
Average coefficients can conceal groups with different priorities. Heterogeneity may reflect observed characteristics, latent classes, random coefficients, experience, or scale differences. Subgroup claims should be planned, adequately powered, and interpreted cautiously.
Heterogeneity can be explored through interactions, mixed logit, latent-class models, or individual-level estimates where data support them. Scale heterogeneity should not be mistaken automatically for stronger preferences.
Attribute non-attendance changes interpretation
Respondents may ignore one or more attributes because they are unimportant, difficult, or used only after another attribute passes a threshold. Stated non-attendance and inferred non-attendance measure different things. Forcing compensatory utility can bias trade-off estimates when decision rules are non-compensatory.
Debrief questions, response time, eye tracking, and alternative models can provide evidence. The analysis should not exclude inconvenient behaviour merely to preserve the planned model.
Willingness to pay depends on the cost coefficient
When cost enters utility linearly, marginal willingness to pay for attribute (k) can be estimated as the negative ratio of its coefficient to the cost coefficient:
$$ WTP_k=-\frac{\beta_k}{\beta_{cost}} $$
The ratio can be unstable when the cost coefficient is small or heterogeneous. Cost sensitivity also depends on income, payment vehicle, frequency, and whether respondents believe payment is consequential.
Marginal rates of substitution quantify trade-offs
The same coefficient-ratio logic can express how much of one attribute respondents would trade for another. If time has coefficient (\beta_{time}), the time equivalent of attribute (k) is:
$$ MRS_{k,time}=-\frac{\beta_k}{\beta_{time}} $$
The result inherits the functional-form and scale assumptions. Nonlinear attributes require comparison over specified level changes rather than a single constant ratio.
Choice probabilities can support scenario prediction
Estimated utility can predict expected shares under hypothetical choice sets. Predictions are relative to the alternatives included and assume that estimated preferences and choice behaviour transport to the new setting. They are not market forecasts unless availability, awareness, access, constraints, and implementation are also represented.
Scenario reporting should show the exact profiles, uncertainty, and effect of adding or removing alternatives. Under multinomial logit, the independence assumption can make substitution patterns unrealistic.
Internal-validity tasks need careful use
Dominance tests, repeated tasks, fixed choices, and comprehension questions can identify potential problems. They can also create learning, signal a preferred answer, or unfairly exclude respondents. Exclusion rules should be pre-specified and sensitivity analyses should retain versus remove flagged responses.
Consistency is not the only sign of validity. Real preferences can be incomplete, constructed, or sensitive to information, especially for unfamiliar health decisions.
External validity is not guaranteed
Stated choices may differ from real behaviour because consequences are hypothetical, budgets are not binding, options are simplified, and actual access is constrained. Consequential framing, realistic tasks, representative samples, and comparison with revealed behaviour can strengthen external validity. None fully eliminates hypothetical bias.
The report should distinguish preference estimates from predicted uptake and observed utilisation. Implementation models should add eligibility, supply, affordability, and other barriers.
Equity affects whose choices are represented
Survey access, language, literacy, digital inclusion, disability, health status, and recruitment can exclude populations affected by the decision. Cost coefficients also reflect ability to pay as well as preference. Aggregate willingness to pay can therefore privilege higher-income respondents.
Equity review should examine sample representation, accessible administration, preference heterogeneity, income effects, and whether a choice set assumes options unavailable to some groups. Distributional results should accompany averages when they affect policy.
Worked choice-set example
Suppose a respondent chooses between two unlabelled treatments described by response probability, adverse-event risk, administration time, and monthly cost. Treatment A has systematic utility 1.2 and Treatment B has utility 0.7. Under a two-option logit model, the predicted probability of choosing A is:
$$ P(A)=\frac{e^{1.2}}{e^{1.2}+e^{0.7}}\approx0.622 $$
The 62.2% prediction applies only to that model, respondent representation, and choice set. Adding an opt-out or another treatment changes the denominator and potentially the predicted shares.
Common mistakes
Choice sets can look intuitive while embedding design choices that determine the result. The following errors commonly weaken inference. Each should be checked before preference estimates support economic or policy conclusions.
- Using attributes selected only by researchers without target-population input.
- Presenting levels that are individually plausible but jointly impossible.
- Treating current care, no treatment, and cannot decide as the same option.
- Omitting an important attribute and interpreting included coefficients causally.
- Maximising statistical efficiency while ignoring comprehension and burden.
- Treating repeated choices from one respondent as independent.
- Assuming all respondents attend to every attribute and trade continuously.
- Reporting willingness to pay without testing cost-coefficient stability and income effects.
- Using stated choice probabilities as direct uptake forecasts.
- Excluding respondents after seeing that their choices do not fit the preferred model.
Reporting a choice set and experiment
Transparent reporting should allow every displayed task to be reconstructed. It should link qualitative development, experimental design, administration, model, and interpretation. Screenshots or full task appendices are often necessary because wording and layout are part of the intervention.
- Report alternatives, labels, attributes, levels, units, frames, and opt-out meaning.
- Explain attribute development and target-population involvement.
- Report design method, priors, restrictions, efficiency, blocking, and randomisation.
- Report tasks per respondent, training, mode, layout, and accessibility.
- Report coding, utility specification, clustering, heterogeneity, and scale assumptions.
- Report validity tests, exclusions, response time, missingness, and sensitivity analyses.
- Separate preference, willingness-to-pay, predicted choice, and actual behaviour claims.
The decision standard
A credible choice set presents realistic alternatives that contain the information respondents need to make the intended trade-off without unnecessary burden or hidden framing. Its design supports identification, but statistical efficiency does not replace content validity and comprehension. The resulting choice reveals preference only within the stated alternatives, attributes, context, and model, so policy use must preserve those boundaries.
Related Concepts (2)
Frequently Asked Questions (6)
What is a choice set?
The group of alternatives, typically hypothetical options and sometimes a status quo, presented to a respondent in a single choice task.
Source: Louviere, Hensher & Swait 2000
How many alternatives should a choice set contain?
Most health studies present two or three hypothetical alternatives, sometimes with an additional option to decline. Two is the easiest task and yields the least information per question. Adding a third increases the information without much increasing difficulty. Beyond that, respondents begin to simplify, commonly by eliminating options on a single attribute rather than weighing everything, which produces answers the analytical model does not describe. The number is therefore a trade between information per task and the fidelity of the response.
Source: Louviere, Hensher & Swait 2000
Why does a choice set often include an opt-out or status quo option?
Because forcing a choice between hypothetical alternatives assumes the respondent would take one of them, and in reality many would take neither. Including an option to decline, or to continue with current arrangements, makes the task resemble the real decision and allows the value of the service as a whole to be estimated rather than only the relative value of its features. Omitting it produces estimates conditional on participation, which overstate uptake if applied to a population that could refuse.
Source: Louviere, Hensher & Swait 2000
How is a choice set constructed?
The combinations of attribute levels forming each alternative are selected by an experimental design rather than assembled at random, so that the effect of each attribute can be separated from the others. Alternatives are then paired into sets in a way that avoids one option dominating another on every attribute, since a dominated pair yields no information about trade-offs. Designs are built to be efficient given assumptions about the likely parameter values, and those assumptions are usually refined after a pilot.
Source: Lancsar & Louviere 2008
How many choice sets can one respondent complete?
More than intuition suggests and fewer than designers would like. Published health studies commonly present eight to sixteen sets per respondent, and evidence indicates that response quality holds over that range for tasks of moderate complexity before fatigue produces simplification. The sustainable number falls as the number of attributes rises, so a design with many attributes must use fewer sets. Testing for a change in response pattern across the sequence is the practical check, since fatigue shows as increasing reliance on a single attribute.
Source: de Bekker-Grob, Ryan & Gerard 2012
What makes a choice set fail?
A set containing an implausible combination undermines the credibility of the whole exercise, so designs must exclude combinations that could not occur, such as the cheapest option also being the fastest and most effective. A set in which one alternative is better on every attribute yields no trade-off information. A set whose alternatives differ on only one attribute reduces the task to a single comparison. Each of these wastes a question, and in a design with limited questions that loss is not recoverable.
Source: healtheconomics.wiki
Trust Record
Verified by Dr Darrin Baines
British health economist
Professional identity: darrinbaines.org
Verification date: 21 Sep 2026
Content version: 1.0.0
Canonical Identity
- Persistent URI
- https://healtheconomics.wiki/concept/choice-set
- Term code
- HE-EE-CBA-008
Stable URI · Machine-readable · Resolvable · CC BY 4.0