Standardised best-minus-worst score in object case best-worst scaling

Divides the difference between the number of times an item was chosen best and the number of times it was chosen worst by the number of times it could have been chosen: the number of respondents times its appearances in each respondent's design. Scores run from minus 1, always worst, to plus 1, always best, and across all items the differences sum to zero.

Signature

S_i = (B_i - W_i) / (N * r_i)
Inputs
InputsDefinitionUnit
B_iNumber of times item i was chosen best across all respondentscount
W_iNumber of times item i was chosen worst across all respondentscount
NNumber of respondents who completed the taskscount
r_iNumber of sets in which item i appears for each respondent; r in a balanced designcount
Output
S_iBest-minus-worst count of item i per opportunity to be chosen, from minus 1 to 1none

Function

Best-worst scaling design, scoring and choice probability function

Maps a planned series of choice sets, each answered with a best and a worst choice, to a scale of preference or priority. The design is usually a balanced incomplete block design, the simplest analysis counts best and worst choices and standardises their difference, and choice models place items on a latent utility scale through the probability of each best and worst pair. The notation follows the Best-Worst Scaling article.

Try this function

Implementations

  • Excel

    Standardised best-minus-worst score from named counts

    With the counts in BestCount and WorstCount, the number of respondents in Respondents and the item's appearances per respondent in Appearances, the formula returns the score.

    =(BestCount-WorstCount)/(Respondents*Appearances)

Assumptions

  • Every respondent completes the same design

    Each respondent answers every set and sees item i r_i times, so N times r_i is the number of opportunities. With several design versions or missing tasks, the opportunities are counted directly.

  • Scores read as an ordering summary

    Best-minus-worst scores summarise ordering. Ratio-scale properties come from a choice model such as the maxdiff model, not from the counts.

Worked examples

  • Object A always chosen best in the seven-set design

    In the article's example one respondent sees each object four times; object A is chosen best four times and never worst, a score of 1.

    B_i = 4; W_i = 0; N = 1; r_i = 4; S_i = 1
  • Object D chosen best once

    Object D is chosen best once, in the only set where it is the strongest object, and never worst, a score of 0.25, above C's score of 0 although C ranks higher in the respondent's true order.

    B_i = 1; W_i = 0; N = 1; r_i = 4; S_i = 0.25
  • Object F chosen worst twice

    Object F is chosen worst twice and never best, a score of minus 0.5.

    B_i = 0; W_i = 2; N = 1; r_i = 4; S_i = -0.5
  • Sample of 100 respondents

    If 100 respondents each see an item four times and choose it best 180 times and worst 60 times, its score is 120 / 400 = 0.3 (computed here for illustration).

    B_i = 180; W_i = 60; N = 100; r_i = 4; S_i = 0.3

Common errors

  • Reading best-minus-worst scores as ratios

    Scores can be negative and sum to zero across items, so a statement that one item is twice as important as another has no meaning on this scale. The square root of best over worst counts is undefined for items never chosen worst, as for A, B, C and D in the article's example.

  • Trusting one respondent's counts for middle-ranked items

    In the article's example C scores 0 and D 0.25 although the respondent prefers C, because every set with C also held A or B. Middle items get less information than extreme ones, so a sample, a model or both are needed.

Sources

  • Best-worst count scores and their standardisation

    Mühlbacher AC, Kaczynski A, Zweifel P, Johnson FR. Experimental measurement of preferences in health and healthcare using best-worst scaling: an overview. Health Economics Review. 2016;6:2. Count analysis: a best-worst score is Total(Best) minus Total(Worst); standardisation divides it by the product of the frequency of occurrence and the sample size; some authors take the square root of Total(Best) over Total(Worst).

    View source →

  • Design used for the best-minus-worst worked example

    Hollin IL, Paskett J, Schuster ALR, Crossnohere NL, Bridges JFP. Best-worst scaling and the prioritization of objects in health: a systematic review. PharmacoEconomics. 2022;40(9):883-899. Introduction: the seven-set design (acge, fgbc, ebaf, gefd, dfca, cdeb, badg) with each object shown four times.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0