Signature
s_i = (b_i - a_i) / max(a_i, b_i)
| Inputs | Definition | Unit |
|---|---|---|
b_i | Average Euclidean distance from unit i to the units of the nearest cluster it does not belong to | units of the clustering variables |
a_i | Average Euclidean distance from unit i to the other units in its own cluster | units of the clustering variables |
s_i | Silhouette value of unit i | none |
|---|
Function
Grouping patients or other units by dissimilarity in cluster analysis
Maps the clustering variables measured on each unit, such as a patient's counts of emergency and outpatient care, to a set of groups found in the data without an outcome variable. Combinatorial methods such as k-means assign each unit to one group by minimising the dissimilarity within groups; mixture models such as latent class analysis give each unit a probability of belonging to each class. The notation follows the Cluster Analysis article, whose six-patient example is used throughout.
Try this function
Implementations
Excel
Silhouette from named mean distances
With the mean distances in MeanOwn and MeanOther, the formula returns the silhouette, held in Silhouette; =AVERAGE over the column of silhouettes gives the mean for the solution.
=(MeanOther-MeanOwn)/MAX(MeanOwn,MeanOther)
Assumptions
Silhouette distances measured as in the clustering
a_i and b_i use ordinary Euclidean distances on the same scaled variables as the clustering; the unit's cluster has at least one other member.
Nearest other cluster chosen by mean distance
With more than two clusters, b_i is the smallest of the mean distances to each other cluster.
Worked examples
Silhouette of patient P2
P2's distances are 1.414 and 2.236 to P1 and P3, mean 1.825, and 4.472, 6.403 and 4.243 to cluster B, mean 5.039, so s is 3.214 / 5.039, about 0.6378 (0.638 in the article).
a_i = 1.8251; b_i = 5.0393; s_i = 0.6378
Silhouette of patient P1
P1's mean distance to P2 and P3 is 1.825 and to cluster B 6.433, giving about 0.7163, the highest of the six (computed here for illustration).
a_i = 1.8251; b_i = 6.4327; s_i = 0.7163
Unit closer to another cluster than its own
A unit with a mean distance of 3 to its own cluster and 2 to the nearest other cluster has a silhouette of minus 1 / 3, about minus 0.3333, a sign that it is misallocated (computed here for illustration).
a_i = 3; b_i = 2; s_i = -0.3333
Common errors
Computing the silhouette from squared distances
Squared distances for P2 give a mean of 3.5 to its own cluster and 26.3 to cluster B, a silhouette of about 0.87 against 0.64 from Euclidean distances (computed here for illustration); values are not comparable with published silhouettes.
Using the mean distance to all other units as b
With more than two clusters, b_i is the mean distance to the nearest other cluster only; averaging over every unit outside the cluster inflates b and the silhouette.
Sources
Silhouette coefficient from the mean intra-cluster and nearest-cluster distances
scikit-learn developers. Clustering. scikit-learn User Guide, section 2.3, version 1.9.1. Accessed 3 October 2026. Section 2.3.11.5, Silhouette Coefficient: a is the mean distance between a sample and all other points in the same class, b the mean distance between a sample and all other points in the next nearest cluster, and s = (b minus a) / max(a, b); the score for a set of samples is the mean; it is bounded between minus 1 for incorrect clustering and plus 1 for highly dense clustering, and scores around zero indicate overlapping clusters.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0