Grouping patients or other units by dissimilarity in cluster analysis
d(x_i, x_j) = sum_(v=1)^p (x_iv - x_jv)^2; W(K) = sum_(k=1)^K sum_(i in C_k) ||x_i - xbar_k||^2
Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.
Squared Euclidean distance between two patient profiles on two clustering variables
d_ij = (x_i1 - x_j1)^2 + (x_i2 - x_j2)^2
Weighted squared distance between two units with variable weights
D_ij = w_1 * (x_i1 - x_j1)^2 + w_2 * (x_i2 - x_j2)^2
Share of total scatter removed by a K-cluster k-means solution
R_K = 1 - W_K / W_1
Posterior probability of latent class membership with two classes
post_1 = pi_1 * f_1 / (pi_1 * f_1 + pi_2 * f_2)
Silhouette of one unit in a cluster solution
s_i = (b_i - a_i) / max(a_i, b_i)