Silhouette of one unit in a cluster solution

Compares a unit's mean distance to the other members of its own cluster, a_i, with its mean distance to the members of the nearest other cluster, b_i. Values run from minus 1 to 1: near 1 the unit sits well inside a dense, well-separated cluster, around 0 the clusters overlap and below 0 the unit is closer to another cluster than to its own. The mean over all units summarises a solution.

Signature

s_i = (b_i - a_i) / max(a_i, b_i)
Inputs
InputsDefinitionUnit
b_iAverage Euclidean distance from unit i to the units of the nearest cluster it does not belong tounits of the clustering variables
a_iAverage Euclidean distance from unit i to the other units in its own clusterunits of the clustering variables
Output
s_iSilhouette value of unit inone

Function

Grouping patients or other units by dissimilarity in cluster analysis

Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.

Try this function

Implementations

  • Excel

    Silhouette from named mean distances

    With the mean distances in MeanOwn and MeanOther, the formula returns the silhouette, held in Silhouette; =AVERAGE over the column of silhouettes gives the mean for the solution.

    =(MeanOther-MeanOwn)/MAX(MeanOwn,MeanOther)

Assumptions

  • Silhouette distances measured as in the clustering

    a_i and b_i use ordinary Euclidean distances on the same scaled variables as the clustering; the unit's cluster has at least one other member.

  • Nearest other cluster chosen by mean distance

    With more than two clusters, b_i is the smallest of the mean distances to each other cluster.

Worked examples

  • Silhouette of patient P2

    P2's distances are 1.414 and 2.236 to P1 and P3, mean 1.825, and 4.472, 6.403 and 4.243 to cluster B, mean 5.039, so s is 3.214 / 5.039, about 0.6378 (0.638 in the article).

    a_i = 1.8251; b_i = 5.0393; s_i = 0.6378
  • Silhouette of patient P1

    P1's mean distance to P2 and P3 is 1.825 and to cluster B 6.433, giving about 0.7163, the highest of the six (computed here for illustration).

    a_i = 1.8251; b_i = 6.4327; s_i = 0.7163
  • Unit closer to another cluster than its own

    A unit with a mean distance of 3 to its own cluster and 2 to the nearest other cluster has a silhouette of minus 1 / 3, about minus 0.3333, a sign that it is misallocated (computed here for illustration).

    a_i = 3; b_i = 2; s_i = -0.3333

Common errors

  • Computing the silhouette from squared distances

    Squared distances for P2 give a mean of 3.5 to its own cluster and 26.3 to cluster B, a silhouette of about 0.87 against 0.64 from Euclidean distances (computed here for illustration); values are not comparable with published silhouettes.

  • Using the mean distance to all other units as b

    With more than two clusters, b_i is the mean distance to the nearest other cluster only; averaging over every unit outside the cluster inflates b and the silhouette.

Sources

  • Silhouette coefficient from the mean intra-cluster and nearest-cluster distances

    scikit-learn developers. Clustering. scikit-learn User Guide, section 2.3, version 1.9.1. Accessed 3 October 2026. Section 2.3.11.5, Silhouette Coefficient: a is the mean distance between a sample and all other points in the same class, b the mean distance between a sample and all other points in the next nearest cluster, and s = (b minus a) / max(a, b); the score for a set of samples is the mean; it is bounded between minus 1 for incorrect clustering and plus 1 for highly dense clustering, and scores around zero indicate overlapping clusters.

    View source →

Canonical Identity

Stable URI · Machine-readable · Resolvable · CC BY 4.0