An unsupervised learning algorithm that partitions a set of points into a chosen number of clusters by alternating between assigning each point to its nearest cluster center and moving each center to the average of the points assigned to it.
Facts
Origin YearThe term k-means itself was coined by James MacQueen in 1967; Edward Forgy published essentially the same method independently in 1965. Core PrincipleAn iterative clustering method that partitions n observations into k clusters by repeatedly assigning each point to the cluster whose mean is nearest and then recomputing each cluster's mean, converging to a local optimum that minimizes within-cluster squared distance. 1 Connections
Sources
1. Wikipedia: K-means clustering
Wikimedia FoundationHistory section, first paragraph
The standard algorithm was first proposed by Stuart Lloyd of Bell Labs in 1957 as a technique for pulse-code modulation, although it was not published as a journal article until 1982.
Lead section, first paragraph
k-means clustering minimizes within-cluster variances (squared Euclidean distances), but not regular Euclidean distances
View the Source Frequently Asked Questions
Who first proposed k-means clustering and what was it originally for?
Stuart Lloyd of Bell Labs proposed it in 1957 for pulse-code modulation.
The standard algorithm was first proposed by Stuart Lloyd of Bell Labs in 1957 as a technique for pulse-code modulation, so it began as a signal-processing method rather than a machine learning one. Its objective is to minimize within-cluster variances, meaning squared Euclidean distances, and not the plain Euclidean distances themselves.
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.