In plain words: Clustering groups similar data points without labels, and this review explains how to represent data, run common algorithms, and check whether the results make sense. It sorts popular methods into two families—linking nearby points versus fitting centers—and compares how each behaves.
Abstract · Clustering -- Basic concepts and methods
We review clustering as an analysis tool and the underlying concepts from an introductory perspective. What is clustering and how can clusterings be realised programmatically? How can data be represented and prepared for a clustering task? And how can clustering results be validated? Connectivity-based versus prototype-based approaches are reflected in the context of several popular methods: single-linkage, spectral embedding, k-means, and Gaussian mixtures are discussed as well as the density-based protocols (H)DBSCAN, Jarvis-Patrick, CommonNN, and density-peaks.
Jan-Oliver Felix Kapp-Joswig, Bettina G. Keller
arXiv:2212.01248 · cs.LG · submitted Dec 1, 2022
abstract · pdf · Two chapters adapted from a doctoral thesis (J.-O. Kapp-Joswig, "Applications of Molecular Dynamics simulations for biomolecular systems and improvements to density-based clustering in the analysis", 2022, FU Berlin), 59 pages, 30 figures