Verifiedv1.0.0

Welfare Economics

Welfare economics is the normative branch of economics that judges whether changes in the allocation of resources and health make society better off.

Last reviewedDarrin Baines IP Ltd

Concept Architecture

Welfare Economics: Pareto, Compensation Tests and Social Welfare Functions in Health

Every economic evaluation in health care rests on a prior judgement about what makes one state of the world better than another, and welfare economics is the part of economics that sets out those judgements and tests their logic. It supplies the criteria that lie behind a positive net present value in cost-benefit analysis, and it explains why a cost per QALY rests on a different set of value judgements. For health economists the subject matters because the choice of welfare framework decides what counts as a benefit, whose valuation is used and how gains to one person are set against losses to another. This page is a hub: it explains the difference between positive and normative economics, the Pareto criterion and its limits, the two fundamental theorems and why health care markets fail their conditions, the Kaldor and Hicks compensation tests and the monetary measures built on them, social welfare functions and Arrow's impossibility theorem, the debate between welfarism and extra-welfarism, and the trade-off between equity and efficiency. An illustrative worked example shows one policy failing the Pareto test, passing the compensation test and then being ranked differently by utilitarian and Rawlsian welfare functions. Detailed treatment of each linked topic sits on its own page.

Positive and normative questions in health economics

Positive economics describes and predicts: how demand for a service responds to a charge, or how a payment reform changes hospital activity. Normative economics judges: whether the change is desirable, and for whom. Welfare economics is the normative branch, and Arrow opened his 1963 paper on medical care by describing it as a study of the specific features of medical care as the object of normative economics.

The distinction matters in practice because a single analysis usually contains both kinds of statement. The estimate that a screening programme prevents a given number of deaths is positive and can be tested against evidence. The conclusion that the programme should be funded requires a criterion for comparing the gains of those who benefit with the losses of those who pay or whose care is displaced, and that criterion is a value judgement. Welfare economics leaves those value judgements in place but makes them explicit and checks that the conclusions follow from them.

The field developed in two phases. The older welfare economics associated with Pigou's The Economics of Welfare treated welfare as a sum that could be increased by raising the national dividend or by transferring income from rich to poor, which assumed that utility could be measured and compared across people. The "new" welfare economics of the late 1930s tried to reach conclusions without interpersonal comparisons of utility, and relied instead on the Pareto criterion and on hypothetical compensation. Pigou's analysis of divergences between private and social net product also founded the modern treatment of externality, which is covered on its own page.

Pareto efficiency and where the criterion runs out

The Pareto criterion is the weakest value judgement in welfare economics and therefore the least contested. A change is a Pareto improvement if it makes at least one person better off and nobody worse off, judged by each person's own preferences. An allocation is Pareto efficient when no such improvement remains, so any further gain to one person requires a loss to someone else.

The criterion runs out quickly in health policy for three reasons. First, almost every real policy creates losers: funding a new medicine displaces other care, and reconfiguring a service lengthens some journeys while shortening others, so a strict Pareto test would block nearly every decision. Second, there are many Pareto efficient allocations, which differ in how resources and health are distributed, and the Pareto criterion alone cannot choose between them. Third, the criterion says nothing about fairness. An allocation in which one group holds almost all the resources can be Pareto efficient if any transfer would make that group worse off.

For health economists the practical consequence is that Pareto efficiency is a necessary condition for a good allocation in most welfare frameworks but never a sufficient one. The broader efficiency concepts used in health care, including allocative efficiency and technical efficiency, are compared on the efficiency page.

The two fundamental theorems and why health care breaks their conditions

The fundamental theorems of welfare economics link competitive markets to Pareto efficiency. Arrow stated both in his 1963 paper and then asked how far the medical care market meets their conditions. His answer shaped the case for public intervention in health care that most health economists now take as their starting point.

  • The first theorem. If a competitive equilibrium exists, and all commodities relevant to costs or utilities are priced in the market, the equilibrium is Pareto optimal.
  • The second theorem. If there are no increasing returns in production, and certain other conditions hold, every Pareto optimal state is a competitive equilibrium corresponding to some initial distribution of purchasing power.

Together the theorems suggest a division of labour: markets handle allocation, and redistribution of income through taxes and transfers handles distribution. Arrow identified three preconditions for this separation to work (the existence of a competitive equilibrium, the marketability of all goods and services relevant to costs and utilities, and non-increasing returns), and argued that when the actual market departs significantly from the competitive model the separation of allocative and distributional procedures becomes, in most cases, impossible. He also noted that in practice almost no set of taxes and subsidies avoids some adverse effect on the achievement of an optimal state.

Health care fails these conditions in several connected ways, each of which is treated on its own page.

  • Externalities. Arrow's own example was communicable disease: a person who is not immunised raises the risk to others, and no market price compensates them. This is a failure of marketability and a classic externality.
  • Uncertainty and incomplete insurance. Illness is largely unpredictable, and Arrow argued that the special economic problems of medical care can be explained as adaptations to uncertainty in the incidence of disease and in the efficacy of treatment. Insurance is incomplete, partly because insurers cannot distinguish avoidable from unavoidable losses, so insurance dilutes the incentive to avoid them. This links to moral hazard and adverse selection.
  • Information asymmetry. The physician usually knows far more than the patient about the likely consequences of treatment, so the patient cannot act as the informed buyer the competitive model assumes. The consequences for the doctor-patient relationship are covered under information asymmetry and the principal-agent problem.

Arrow's central interpretive claim was that when the market fails to achieve an optimal state, society will to some extent recognise the gap and non-market institutions will arise to bridge it. Professional ethics, trust in the physician, licensing and non-profit hospitals are read in that light. The general framework for these departures from the competitive model is set out on the market failure page.

Compensation tests: Kaldor, Hicks and the potential Pareto improvement

The compensation tests were designed to rank policies that create losers without making interpersonal comparisons of utility. Kaldor proposed in 1939 that a change is an improvement if the gainers could compensate the losers and still be better off, whether or not compensation is actually paid; whether to pay it was, in his view, a political question separate from the efficiency judgement. Hicks proposed the same criterion in the same year, describing the reforms of interest as those that would allow compensation for the losses and still show a net advantage. A reverse form of the test, often labelled the Hicks test, asks whether the losers could profitably bribe the gainers to return to the original position; a change passes if they could not.

A change that passes the compensation test is called a potential Pareto improvement, and the test is usually called the Kaldor-Hicks criterion. It is the welfare-economic basis of cost-benefit analysis: if the sum of the monetary gains to everyone affected exceeds the sum of the monetary losses, including the opportunity cost of resources, the gainers could in principle compensate the losers.

The tests have well-known weaknesses. Scitovsky showed in 1941 that a change and its reversal can both pass the Kaldor test when the change alters the quantities of goods available as well as their distribution, so the criterion can fail to give a consistent ranking. His remedy, often called the Scitovsky double criterion, requires a change to pass the Kaldor test while its reversal fails it. More fundamentally, hypothetical compensation leaves the losers worse off in fact. Adding monetary gains and losses also gives each pound the same weight regardless of who gains or loses it, and because willingness to pay is bounded by income, the test gives more weight to the preferences of people with more money. Brouwer, Culyer, van Exel and Rutten point out that applied cost-benefit analysis therefore makes an interpersonal comparison after all, by assigning an equal unit weight to each monetary gain or loss.

Measuring gains and losses in money: compensating variation, equivalent variation and consumer surplus

The compensation tests need a monetary measure of each person's gain or loss, and welfare economics supplies two exact measures based on the utility levels a person reaches. The compensating variation is the sum that, taken from or given to a person after a change, returns them to their original utility level. The equivalent variation is the sum that, given to or taken from them instead of the change, would bring them to the utility level the change would have produced.

For an improvement in health, the compensating variation is the maximum willingness to pay for the gain, and the equivalent variation is the minimum willingness to accept to forgo it. Which measure is appropriate depends on whether people are taken to be entitled to the status quo or to the changed state. The two measures differ because of income effects, and they differ most for goods with few substitutes, of which a person's own health is an example.

Consumer surplus, the area under a demand curve above the price paid, is the older Marshallian measure. It is easy to estimate from market data but is in general only an approximation to the exact measures. For a change in a single price it lies between the compensating and equivalent variations, and Willig derived bounds showing that the approximation error is usually small when the surplus is a small share of income and income effects are modest. In health care, where prices are often absent or regulated, stated-preference methods that estimate willingness to pay directly are more common, as described on the willingness to pay page. Producer surplus measures the corresponding gain to suppliers, and the resources a programme uses are valued at their opportunity cost.

Social welfare functions: from Bergson and Samuelson to Rawls

A social welfare function makes the distributional judgement that the compensation tests avoid. Bergson introduced the idea in 1938 and Samuelson developed it in his 1947 Foundations of Economic Analysis: social welfare is written as a function of the welfare of each individual, and its form states explicitly how gains to one person are set against losses to another. In its general form:

W = W(U1, U2, ..., Un)

where W is social welfare, U1 to Un are the utility levels (or, in extra-welfarist versions, the health levels) of the n individuals, and W increases in each argument so that the function respects the Pareto criterion. The Bergson-Samuelson function can select a single best point among the many Pareto efficient allocations, but only once a normative choice about distribution has been made.

Three forms recur in health economics, and each answers the distributional question differently.

  • Utilitarian. W = U1 + U2 + ... + Un, so a unit of utility or health counts the same whoever receives it. Maximising total QALYs from a fixed budget is the health analogue.
  • Rawlsian (maximin). W = min(U1, U2, ..., Un), so social welfare is judged by the position of the worst-off person. The form is associated with Rawls's A Theory of Justice, although Rawls applied his difference principle to primary goods rather than to utility.
  • Weighted or inequality-averse. Gains to worse-off people receive a weight greater than one, or the function is made concave so that equal increments add less welfare as the level rises. Wagstaff argued that the weighting schemes then proposed for QALYs captured distributional concerns other than a concern about inequality, and proposed a social welfare function over individuals' health that incorporates both, which turns the choice between total QALYs and their distribution into an explicit equity and efficiency trade-off.

Arrow's 1951 monograph Social Choice and Individual Values set a limit on this programme. In the form usually stated, which follows the 1963 second edition, the theorem holds that with at least two people and three or more alternatives, no rule for turning individual rankings into a complete and transitive social ranking can satisfy four conditions at once: it must work for any pattern of individual preferences, rank one option above another whenever everyone does, make the social ranking of two options depend only on how individuals rank those two options, and avoid dictatorship. The 1951 original used a related set of conditions from which the unanimity condition follows. Any practical social welfare function therefore relaxes at least one condition, for example by allowing interpersonal comparisons of wellbeing, as distributional weights and extra-welfarist health maximisation both do.

Welfarism and extra-welfarism

The largest methodological divide in health economic evaluation concerns what the social welfare function should contain. Welfarism, a term health economists take from Sen, judges the goodness of a situation solely by the utility levels individuals reach, as assessed by those individuals. In the account given by Brouwer, Culyer, van Exel and Rutten, the dominant welfarist framework rests on four tenets: utility maximisation, individual sovereignty over what contributes to utility, consequentialism, and welfarism in the narrow sense of confining the evaluative space to individual utility.

Extra-welfarism widens that space. Culyer's 1989 paper on the normative economics of health care finance and provision is the usual starting point for the approach in health economics: it moved attention from the utility people derive from goods to characteristics of people themselves, of which health is the one most directly affected by health care. Brouwer and colleagues identify four ways in which extra-welfarism differs from welfarism.

  1. Outcomes. It permits outcomes other than utility, such as health.
  2. Sources of valuation. It permits valuations from sources other than the affected individuals, such as a representative sample of the public or a decision-maker.
  3. Weighting. It permits outcomes to be weighted by principles that need not be preference-based.
  4. Interpersonal comparisons. It permits comparisons of wellbeing between people across several dimensions, which moves beyond Paretian economics.

Sen's capability approach, set out in Commodities and Capabilities, is one root of this position. Sen distinguished the commodities a person has, the functionings they achieve (what they manage to do or be) and their capabilities (the set of functionings they could achieve), and argued that utility, understood as happiness or desire fulfilment, can mislead because people adapt their expectations to deprivation. Extra-welfarism in health economics took from this the idea that the evaluative space can be something other than utility. Coast, Smith and Lorgelly note, however, that extra-welfarism as Culyer developed it is not a direct application of the capability approach, because its focus remains narrow: the concern is with functioning, and maximisation is retained.

The divide maps onto the main forms of economic evaluation. Cost-benefit analysis is welfarist: benefits are measured as individuals' own willingness to pay, and the decision rule is the Kaldor-Hicks potential Pareto improvement. Cost-utility analysis, as NICE and many other health technology assessment bodies apply it, is usually classed as extra-welfarist. The NICE technology appraisal manual (PMG36) justifies cost-effectiveness analysis by NICE's focus on maximising health gains from a fixed NHS and personal social services budget, measures health effects in QALYs, draws preference data for valuing health states from a representative sample of the UK population rather than from the patients who benefit, and gives an additional QALY the same weight regardless of the other characteristics of the people receiving it, except in specific circumstances. On each of the four dimensions above, that reference case departs from welfarism.

Trading equity against efficiency: distributional weights

Neither the compensation tests nor unweighted QALY maximisation says anything about who gains. Both count a unit of benefit equally whoever receives it, and so both favour the option with the largest total even when its gains go to people who are already better off. The equity page sets out the principles of fairness that may justify departing from this, and the link to the objective of maximising total benefit is set out on the efficiency page.

Welfare economics offers a direct way to handle the trade-off: attach weights to gains according to who receives them, as a social welfare function with inequality aversion implies. In cost-benefit analysis, distributional weights can give more weight to monetary gains of lower-income groups to offset the influence of ability to pay. In health, equity-weighted analysis applies weights to QALYs, and distributional cost-effectiveness analysis models the distribution of health under each option and evaluates it against the joint objectives of improving total health and reducing unfair inequality.

Weights express value judgements, even when they are elicited from surveys of public preferences. Good practice presents weighted and unweighted results side by side, states the principle behind the weights, and reports how large the weights must be to change the decision, which the worked example below illustrates.

Worked example: one reconfiguration, three verdicts

The figures in this example are illustrative. A health authority is considering concentrating a specialist service at one city hospital. Two groups of 1,000 people each are affected: Group A lives near the city hospital and gains faster access, and Group B lives in an outlying area and loses a local clinic. The example assumes no income effects, so each person's compensating and equivalent variations are equal. It also assumes that every option costs the health authority the same, so resource costs cancel and each comparison turns only on the gains and losses of the two groups.

Step 1: the Pareto test

Each member of Group A would pay up to £300 for the change, and each member of Group B would need £200 in compensation to accept it. Because every member of Group B is made worse off, the change is not a Pareto improvement, and the Pareto criterion cannot rank it against the status quo.

Step 2: the Kaldor and Hicks tests

The compensation tests set the monetary gains against the monetary losses. Each total is the number of people in the group multiplied by the compensating variation per person. In pounds:

Gains to Group A: 1,000 × 300 = 300,000

Losses to Group B: 1,000 × 200 = 200,000

Net gain: 300,000 − 200,000 = 100,000

If each member of Group A paid £200 to a member of Group B, a total of £200,000, every member of Group B would be exactly as well off as before and each member of Group A would keep a gain of 300 − 200 = 100 pounds. The gainers could compensate the losers and remain better off, so the change passes the Kaldor test. Once the change has happened, the most Group B could offer to reverse it is £200,000, which is less than the £300,000 Group A would need to give it up, so the losers cannot profitably bribe the gainers and the reverse (Hicks) test is also passed. The change therefore meets the Scitovsky double criterion as well. Because costs are equal across options, a cost-benefit analysis would report a net benefit of £100,000. Unless the compensation is actually paid, however, 1,000 people are worse off.

Step 3: utilitarian and Rawlsian social welfare functions

The authority also has estimates of total expected QALYs for each group under three options: the status quo (S), the reconfiguration (X), and an alternative (Y) that uses the same budget to keep the clinic and add an outreach service for Group B. Because the groups are the same size and everyone within a group is assumed to be affected alike, comparing group totals ranks the options in the same way as comparing QALYs per person, including for the Rawlsian function.

OptionGroup A total QALYsGroup B total QALYsUtilitarian sumRawlsian minimum
S (status quo)20,00012,00032,00012,000
X (reconfiguration)20,03011,99032,02011,990
Y (outreach)20,00012,01532,01512,015

The utilitarian function adds the two groups' QALYs: 20,030 + 11,990 = 32,020 for X, 20,000 + 12,015 = 32,015 for Y and 20,000 + 12,000 = 32,000 for S. It ranks X first, then Y, then S, in line with the cost-benefit result. The Rawlsian function looks only at the worse-off group, which has 11,990 QALYs under X, 12,015 under Y and 12,000 under S. It ranks Y first, then S, then X, so the option the utilitarian function prefers comes last.

The monetary and health measures need not move together. Group A's £300,000 corresponds to a gain of 20,030 − 20,000 = 30 QALYs, or £10,000 per QALY, while Group B's £200,000 corresponds to a loss of 12,000 − 11,990 = 10 QALYs, or £20,000 per QALY, because Group B's valuation also reflects longer journeys and other effects that the QALY does not capture.

The Pareto criterion settles only one of these comparisons. In health terms Y is a Pareto improvement on S, because Group A is unchanged and Group B gains 12,015 − 12,000 = 15 QALYs, but X cannot be ranked against either S or Y without a distributional judgement.

Step 4: the weight at which the decision switches

Between the utilitarian and Rawlsian extremes lies a family of weighted functions. Each gives a QALY accruing to Group B a weight w relative to a QALY accruing to Group A, and the question is how large w must be before the ranking of X and Y changes. The weighted function is:

W = QA + w × QB

where W is weighted social welfare, QA and QB are the total QALYs of Groups A and B, and w is the distributional weight on Group B. X and Y tie when 20,030 + 11,990w equals 20,000 + 12,015w, which gives w = 30 ÷ 25 = 1.2. With a weight of 2, for example, X scores 20,030 + 2 × 11,990 = 44,010 and Y scores 20,000 + 2 × 12,015 = 44,030, so Y is preferred.

The result can be read in welfare-economic terms. The cost-benefit analysis and the utilitarian function both favour the reconfiguration because they give every unit of gain the same weight. Any weight on Group B's health above 1.2, including the extreme Rawlsian case, reverses that ranking. Reporting the switching weight lets a decision-maker see how strong a concern for the worse-off group must be to change the decision, without the analyst choosing the weight on their behalf.

How the welfare frameworks map onto evaluation methods

The same decision problem produces different answers under different welfare frameworks, and many disagreements between analysts are in fact disagreements about framework. The table summarises how the main approaches used in health economics differ along the dimensions Brouwer and colleagues use.

FrameworkWhat counts as a benefitWhose valuationHow gains are aggregatedTypical method
Paretian welfarismIndividual utilityAffected individualsOnly unanimous gains countRarely usable alone
Kaldor-Hicks welfarismIndividual utility in moneyAffected individuals, through willingness to payUnweighted sum of monetary gains and lossesCost-benefit analysis
Bergson-Samuelson welfarismIndividual utilityAffected individualsExplicit social welfare functionCost-benefit analysis with distributional weights
Extra-welfarismHealth, or other characteristics of peoplePublic or decision-makerQALY sum, with or without equity weightsCost-utility analysis, distributional cost-effectiveness analysis

No framework is neutral. Choosing cost-utility analysis over cost-benefit analysis is a choice to value health on behalf of the public rather than let each person value their own gains, and choosing cost-benefit analysis is a choice to let ability to pay influence the result unless weights are applied. Both choices can be defended, and both should be stated.

Common misreadings of welfare-economic results

A frequent error is to read a positive net benefit as proof that a policy makes society better off. It shows that the gainers could compensate the losers, not that they will, and it relies on counting each pound equally whoever gains or loses it.

A second is to treat Pareto efficiency as the same thing as a desirable allocation. Many Pareto efficient allocations exist, some of them highly unequal, and choosing between them needs a distributional judgement that the Pareto criterion cannot supply.

A third is to assume that cost-per-QALY analysis avoids value judgements because it does not use money to measure benefits. It embeds its own judgements about the maximand, the source of valuation and the equal weighting of QALYs. A fourth is to read Arrow's impossibility theorem as showing that social welfare functions are meaningless; it shows only that any usable function must give up at least one of Arrow's conditions, for example by allowing interpersonal comparisons.

A final error is to cite market failure as a sufficient case for any particular intervention. Arrow's analysis explains why health care markets depart from the competitive model and why non-market institutions arise, but whether a specific policy improves welfare still depends on its costs, its effects and its distribution.

Sources

  • Arrow KJ. Social Choice and Individual Values. Cowles Commission Monograph No. 12. New York: John Wiley and Sons; 1951 (2nd edition 1963; 3rd edition, Yale University Press, 2012).
  • Arrow KJ. Uncertainty and the welfare economics of medical care. American Economic Review. 1963;53(5):941-973.
  • Bergson A. A reformulation of certain aspects of welfare economics. Quarterly Journal of Economics. 1938;52(2):310-334.
  • Brouwer WBF, Culyer AJ, van Exel NJA, Rutten FFH. Welfarism vs. extra-welfarism. Journal of Health Economics. 2008;27(2):325-338.
  • Coast J, Smith RD, Lorgelly P. Welfarism, extra-welfarism and capability: the spread of ideas in health economics. Social Science and Medicine. 2008;67(7):1190-1198.
  • Culyer AJ. The normative economics of health care finance and provision. Oxford Review of Economic Policy. 1989;5(1):34-58.
  • Drummond MF, Sculpher MJ, Claxton K, Stoddart GL, Torrance GW. Methods for the Economic Evaluation of Health Care Programmes. 4th edition. Oxford: Oxford University Press; 2015.
  • Hanemann WM. Willingness to pay and willingness to accept: how much can they differ? American Economic Review. 1991;81(3):635-647.
  • Hicks JR. The foundations of welfare economics. Economic Journal. 1939;49(196):696-712.
  • Johannesson M. Theory and Methods of Economic Evaluation of Health Care. Developments in Health Economics and Public Policy. Boston: Springer; 1996.
  • Kaldor N. Welfare propositions of economics and interpersonal comparisons of utility. Economic Journal. 1939;49(195):549-552.
  • National Institute for Health and Care Excellence. NICE technology appraisal and highly specialised technologies guidance: the manual (PMG36). Published 31 January 2022, last updated 31 March 2026.
  • Pigou AC. The Economics of Welfare. London: Macmillan; 1920.
  • Rawls J. A Theory of Justice. Cambridge, MA: Belknap Press of Harvard University Press; 1971.
  • Samuelson PA. Foundations of Economic Analysis. Harvard Economic Studies, Volume 80. Cambridge, MA: Harvard University Press; 1947.
  • Scitovszky T de. A note on welfare propositions in economics. Review of Economic Studies. 1941;9(1):77-88.
  • Sen A. Commodities and Capabilities. Amsterdam: North-Holland; 1985.
  • Wagstaff A. QALYs and the equity-efficiency trade-off. Journal of Health Economics. 1991;10(1):21-41.
  • Willig RD. Consumer's surplus without apology. American Economic Review. 1976;66(4):589-597.

Trust Record

Verified by Dr Darrin Baines

British health economist

Professional identity: darrinbaines.org

Verification date: 29 Sep 2026

Content version: 1.0.0

Canonical Identity

Term code
HE-EE-WE-028

Stable URI · Machine-readable · Resolvable · CC BY 4.0