Signature
d_ij = (x_i1 - x_j1)^2 + (x_i2 - x_j2)^2
| Inputs | Definition | Unit |
|---|---|---|
x_i1 | Value of the first clustering variable for unit i, such as emergency department attendances in a year | as measured |
x_j1 | Value of the first clustering variable for unit j | as measured |
x_i2 | Value of the second clustering variable for unit i, such as outpatient appointments in a year | as measured |
x_j2 | Value of the second clustering variable for unit j | as measured |
d_ij | Dissimilarity between the profiles of units i and j | squared units of the clustering variables |
|---|
Function
Grouping patients or other units by dissimilarity in cluster analysis
Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.
Try this function
Implementations
Excel
Squared distance between two patient profiles from named cells
With the two units' values in EdI, EdJ, OpI and OpJ, the formula returns the squared distance, held in SqDistance. For p variables held in two equal-length ranges ProfileI and ProfileJ, =SUMXMY2(ProfileI,ProfileJ) gives the same sum.
=(EdI-EdJ)^2+(OpI-OpJ)^2
Assumptions
Quantitative clustering variables on comparable scales
Both variables are quantitative and on similar ranges, as with the two counts in the article's example; otherwise the variable with the larger variance dominates the distance.
Clustering variables chosen on subject-matter grounds
The variables define what similar means: segmenting on care use and segmenting on diagnoses group the same patients differently, so the choice is stated before any algorithm is run.
Worked examples
Patient P2 against patient P3 in the six-patient example
P2 has 1 emergency attendance and 2 outpatient appointments and P3 has 2 and 0, so the squared distance is 1 plus 4, which is 5, as in the article.
x_i1 = 1; x_j1 = 2; x_i2 = 2; x_j2 = 0; d_ij = 5
Patient P2 against patient P4 in the six-patient example
P4 has 5 attendances and 4 appointments, so the squared distance from P2 is 16 plus 4, which is 20; P2 is nearer P3 and joins its cluster.
x_i1 = 1; x_j1 = 5; x_i2 = 2; x_j2 = 4; d_ij = 20
Patient P2 against the converged mean of cluster A
After the first update the mean of cluster A is (1, 1), so P2 at (1, 2) is at squared distance 1 from it, against 25 from the cluster B mean (5, 5), and stays in A.
x_i1 = 1; x_j1 = 1; x_i2 = 2; x_j2 = 1; d_ij = 1
Common errors
Clustering on raw cost alongside counts of care
Had annual cost in pounds been a third variable, the squared cost difference between P1 and P5 alone, 51,840,000, would dwarf their squared difference of 61 on the two counts, and the clusters would simply be cost bands; cost is better rescaled or described after clustering.
Adding a clustering variable that repeats the others
Adding total contacts, the sum of the two counts, as a third variable counts the same care use twice: the squared distance from P2 rises from 5 to 6 for P3 but from 20 to 56 for P4 (computed here for illustration), so the shared dimension gains weight without any new information.
Mixing squared and unsquared distances in later steps
K-means works with squared distances, but the silhouette uses ordinary Euclidean distances; using squared distances for P2 gives a silhouette of about 0.87 instead of 0.64 (computed here for illustration).
Sources
Squared distance as the most common dissimilarity for quantitative variables
Hastie T, Tibshirani R, Friedman J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction. 2nd ed. New York: Springer; 2009. Section 14.3.2, equations 14.20 and 14.21: the dissimilarity between objects sums attribute dissimilarities, and by far the most common choice is squared distance; section 14.3, p. 502: the definition of similarity can only come from subject-matter considerations; section 14.3.6: k-means uses squared Euclidean distance as the dissimilarity.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0