Computing Atlas

How Computing Was Built
Sign In
Text size
100%
Theme
Concept

K-Means Clustering

Algorithm

An unsupervised learning algorithm that partitions a set of points into a chosen number of clusters by alternating between assigning each point to its nearest cluster center and moving each center to the average of the points assigned to it.

Facts
Origin Year
1957 1
The term k-means itself was coined by James MacQueen in 1967; Edward Forgy published essentially the same method independently in 1965.
Core Principle
An iterative clustering method that partitions n observations into k clusters by repeatedly assigning each point to the cluster whose mean is nearest and then recomputing each cluster's mean, converging to a local optimum that minimizes within-cluster squared distance. 1
Connections

In Field

Sources
1. Wikipedia: K-means clustering
Wikimedia Foundation
  • History section, first paragraph
    The standard algorithm was first proposed by Stuart Lloyd of Bell Labs in 1957 as a technique for pulse-code modulation, although it was not published as a journal article until 1982.
  • Lead section, first paragraph
    k-means clustering minimizes within-cluster variances (squared Euclidean distances), but not regular Euclidean distances
View the Source
Frequently Asked Questions

Who first proposed k-means clustering and what was it originally for?

Stuart Lloyd of Bell Labs proposed it in 1957 for pulse-code modulation.

The standard algorithm was first proposed by Stuart Lloyd of Bell Labs in 1957 as a technique for pulse-code modulation, so it began as a signal-processing method rather than a machine learning one. Its objective is to minimize within-cluster variances, meaning squared Euclidean distances, and not the plain Euclidean distances themselves.
Comments (0)
No comments yet. Be the first to share a thought.
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.