Functions & Formulae

Each applied formula has its own function page, with a signature, implementations, and tests.

Grouping patients or other units by dissimilarity in cluster analysis

d(x_i, x_j) = sum_(v=1)^p (x_iv - x_jv)^2; W(K) = sum_(k=1)^K sum_(i in C_k) ||x_i - xbar_k||^2

Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.

  • Squared Euclidean distance between two patient profiles on two clustering variables

    d_ij = (x_i1 - x_j1)^2 + (x_i2 - x_j2)^2

    Adds the squared differences between two units on each clustering variable. It is the dissimilarity k-means minimises, written here for two variables, such as annual emergency department attendances and outpatient appointments; with p variables the sum has p terms. Because differences are squared, variables with larger spread carry more weight, which is why scaling is a separate decision (HE-FM-CLU-002).

  • Weighted squared distance between two units with variable weights

    D_ij = w_1 * (x_i1 - x_j1)^2 + w_2 * (x_i2 - x_j2)^2

    Multiplies each squared difference by a weight that sets the variable's influence. Each variable's influence on squared Euclidean distance is proportional to its variance, so weights in inverse proportion to the variance give equal influence; this is what standardising the variables does. Multiplying every weight by the same constant rescales all distances equally and leaves the grouping unchanged.

  • Share of total scatter removed by a K-cluster k-means solution

    R_K = 1 - W_K / W_1

    Compares the within-cluster sum of squares W(K), the squared distances of units from their own cluster mean that k-means minimises, with its value for one cluster, the total sum of squares about the overall mean. The share removed rises as clusters are added, so the number of clusters is judged by where the gains flatten (the kink or elbow), not by maximising it.

  • Posterior probability of latent class membership with two classes

    post_1 = pi_1 * f_1 / (pi_1 * f_1 + pi_2 * f_2)

    Applies Bayes' rule to give the probability that a unit belongs to a class, given its observed data: the class share times the likelihood of the data under that class, divided by the same quantity summed over classes. Units are commonly allocated to their most probable class, but the probability shows how uncertain the allocation is. With K classes the denominator has K terms.

  • Silhouette of one unit in a cluster solution

    s_i = (b_i - a_i) / max(a_i, b_i)

    Compares a unit's mean distance to the other members of its own cluster, a_i, with its mean distance to the members of the nearest other cluster, b_i. Values run from minus 1 to 1: near 1 the unit sits well inside a dense, well-separated cluster, around 0 the clusters overlap and below 0 the unit is closer to another cluster than to its own. The mean over all units summarises a solution.