Home Glossary K-Means Clustering

K-Means Clustering - Page 2

K-means clustering separates unlabeled data into a chosen number of groups based on similarity. The algorithm starts with k centroids, assigns each observation to the nearest centroid, recalculates the center of every group, and repeats until assignments stabilize. It is widely used for customer segmentation, document grouping, image compression, and exploratory analysis. Results depend on the selected value of k, feature scaling, initial centroids, and the distance metric. K-means works best with compact, similarly sized clusters and can be distorted by outliers or irregular shapes. Because the algorithm always produces groups, analysts must verify that the resulting clusters are meaningful rather than artifacts of the chosen settings.