Signature
D_ij = w_1 * (x_i1 - x_j1)^2 + w_2 * (x_i2 - x_j2)^2
| Inputs | Definition | Unit |
|---|---|---|
w_1 | Weight on the first variable, such as 1 divided by its sample variance; zero or above | inverse squared units of the first variable |
x_i1 | Value of the first variable for unit i | as measured |
x_j1 | Value of the first variable for unit j | as measured |
w_2 | Weight on the second variable; zero or above | inverse squared units of the second variable |
x_i2 | Value of the second variable for unit i | as measured |
x_j2 | Value of the second variable for unit j | as measured |
D_ij | Weighted dissimilarity between the profiles of units i and j | depends on the weights; none with standardising weights |
|---|
Function
Grouping patients or other units by dissimilarity in cluster analysis
Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.
Try this function
Implementations
Excel
Weighted squared distance from named cells
With weights in WeightOne and WeightTwo and values in VarOneI, VarOneJ, VarTwoI and VarTwoJ, the formula returns the weighted distance, held in WeightedDistance. A standardising weight is =1/VAR.S(range) over the variable's values; for p variables, =SUMPRODUCT(Weights,(ProfileI-ProfileJ)^2) gives the same result.
=WeightOne*(VarOneI-VarOneJ)^2+WeightTwo*(VarTwoI-VarTwoJ)^2
Assumptions
Weights chosen for the purpose of the segmentation
Equal influence is not always right: standardising can obscure groups that only some variables separate, so the weights are a stated analyst choice.
Standardising weights computed on the clustered sample
Inverse-variance weights use the variance of each variable in the units being clustered; weights from another sample give a different balance.
Worked examples
Patients P1 and P5 on emergency attendances and cost without weights
With weights of 1, the squared difference of 36 in attendances is lost against the squared cost difference of 51,840,000 pounds squared, giving 51,840,036 (computed here for illustration).
w_1 = 1; x_i1 = 0; x_j1 = 6; w_2 = 1; x_i2 = 400; x_j2 = 7600; D_ij = 51840036
Patients P1 and P5 with standardising weights
In the six patients the sample variance is 5.6 for attendances and 10,359,000 for cost; weights of 1 / 5.6 and 1 / 10,359,000 give 6.4286 plus 5.0043, about 11.4329, so both variables now count (computed here for illustration).
w_1 = 0.178571; x_i1 = 0; x_j1 = 6; w_2 = 0.0000000965344; x_i2 = 400; x_j2 = 7600; D_ij = 11.4329
Common errors
Assuming equal weights give each variable equal influence
Weights of 1 leave each variable's influence proportional to its variance; in the example cost carries over 99.9 per cent of the unweighted distance between P1 and P5.
Standardising by default when only some variables separate the groups
Rescaling every variable to unit variance can hide two well-separated groups that one variable defines, so the weighting is a substantive choice, not a routine step.
Sources
Weighted dissimilarity and the influence of each variable's variance
Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd ed. New York: Springer; 2009. Section 14.3.3, equations 14.24 to 14.27, pp. 505-506: overall dissimilarity is a weighted combination of attribute dissimilarities; with squared error the average dissimilarity on attribute j is twice its variance, so the relative importance of each variable is proportional to its variance and weights in inverse proportion give equal influence; Figure 14.5: standardisation, equivalent to weights 1/[2 var(X_j)], obscured two well-separated groups.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0