Weighted squared distance between two units with variable weights

Multiplies each squared difference by a weight that sets the variable's influence. Each variable's influence on squared Euclidean distance is proportional to its variance, so weights in inverse proportion to the variance give equal influence; this is what standardising the variables does. Multiplying every weight by the same constant rescales all distances equally and leaves the grouping unchanged.

Signature

D_ij = w_1 * (x_i1 - x_j1)^2 + w_2 * (x_i2 - x_j2)^2
Inputs
InputsDefinitionUnit
w_1Weight on the first variable, such as 1 divided by its sample variance; zero or aboveinverse squared units of the first variable
x_i1Value of the first variable for unit ias measured
x_j1Value of the first variable for unit jas measured
w_2Weight on the second variable; zero or aboveinverse squared units of the second variable
x_i2Value of the second variable for unit ias measured
x_j2Value of the second variable for unit jas measured
Output
D_ijWeighted dissimilarity between the profiles of units i and jdepends on the weights; none with standardising weights

Function

Grouping patients or other units by dissimilarity in cluster analysis

Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.

Try this function

Implementations

  • Excel

    Weighted squared distance from named cells

    With weights in WeightOne and WeightTwo and values in VarOneI, VarOneJ, VarTwoI and VarTwoJ, the formula returns the weighted distance, held in WeightedDistance. A standardising weight is =1/VAR.S(range) over the variable's values; for p variables, =SUMPRODUCT(Weights,(ProfileI-ProfileJ)^2) gives the same result.

    =WeightOne*(VarOneI-VarOneJ)^2+WeightTwo*(VarTwoI-VarTwoJ)^2

Assumptions

  • Weights chosen for the purpose of the segmentation

    Equal influence is not always right: standardising can obscure groups that only some variables separate, so the weights are a stated analyst choice.

  • Standardising weights computed on the clustered sample

    Inverse-variance weights use the variance of each variable in the units being clustered; weights from another sample give a different balance.

Worked examples

  • Patients P1 and P5 on emergency attendances and cost without weights

    With weights of 1, the squared difference of 36 in attendances is lost against the squared cost difference of 51,840,000 pounds squared, giving 51,840,036 (computed here for illustration).

    w_1 = 1; x_i1 = 0; x_j1 = 6; w_2 = 1; x_i2 = 400; x_j2 = 7600; D_ij = 51840036
  • Patients P1 and P5 with standardising weights

    In the six patients the sample variance is 5.6 for attendances and 10,359,000 for cost; weights of 1 / 5.6 and 1 / 10,359,000 give 6.4286 plus 5.0043, about 11.4329, so both variables now count (computed here for illustration).

    w_1 = 0.178571; x_i1 = 0; x_j1 = 6; w_2 = 0.0000000965344; x_i2 = 400; x_j2 = 7600; D_ij = 11.4329

Common errors

  • Assuming equal weights give each variable equal influence

    Weights of 1 leave each variable's influence proportional to its variance; in the example cost carries over 99.9 per cent of the unweighted distance between P1 and P5.

  • Standardising by default when only some variables separate the groups

    Rescaling every variable to unit variance can hide two well-separated groups that one variable defines, so the weighting is a substantive choice, not a routine step.

Sources

  • Weighted dissimilarity and the influence of each variable's variance

    Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd ed. New York: Springer; 2009. Section 14.3.3, equations 14.24 to 14.27, pp. 505-506: overall dissimilarity is a weighted combination of attribute dissimilarities; with squared error the average dissimilarity on attribute j is twice its variance, so the relative importance of each variable is proportional to its variance and weights in inverse proportion give equal influence; Figure 14.5: standardisation, equivalent to weights 1/[2 var(X_j)], obscured two well-separated groups.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0