Normalised Clustering Accuracy: An Asymmetric External Cluster Validity Measure

Normalised Clustering Accuracy: An Asymmetric External Cluster Validity Measure

Author

Listed:

Marek Gagolewski
(Warsaw University of Technology
Polish Academy of Sciences)

Abstract

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underperform. Clustering algorithms are traditionally evaluated using either internal or external validity measures. Internal measures quantify different aspects of the obtained partitions, e.g., the average degree of cluster compactness or point separability. However, their validity is questionable because the clusterings they endorse can sometimes be meaningless. External measures, on the other hand, compare the algorithms’ outputs to fixed ground truth groupings provided by experts. In this paper, we argue that the commonly used classical partition similarity scores, such as the normalised mutual information, Fowlkes–Mallows, or adjusted Rand index, miss some desirable properties. In particular, they do not identify worst-case scenarios correctly, nor are they easily interpretable. As a consequence, the evaluation of clustering algorithms on diverse benchmark datasets can be difficult. To remedy these issues, we propose and analyse a new measure: a version of the optimal set-matching accuracy, which is normalised, monotonic with respect to some similarity relation, scale-invariant, and corrected for the imbalancedness of cluster sizes (but neither symmetric nor adjusted for chance).

Suggested Citation

Marek Gagolewski, 2025. "Normalised Clustering Accuracy: An Asymmetric External Cluster Validity Measure," Journal of Classification, Springer;The Classification Society, vol. 42(1), pages 2-30, March.

Handle: RePEc:spr:jclass:v:42:y:2025:i:1:d:10.1007_s00357-024-09482-2
DOI: 10.1007/s00357-024-09482-2

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

Jeffrey L. Andrews & Ryan Browne & Chelsey D. Hvingelby, 2022. "On Assessments of Agreement Between Fuzzy Partitions," Journal of Classification, Springer;The Classification Society, vol. 39(2), pages 326-342, July.
Lawrence Hubert & Phipps Arabie, 1985. "Comparing partitions," Journal of Classification, Springer;The Classification Society, vol. 2(1), pages 193-218, December.
Antonio D’Ambrosio & Sonia Amodio & Carmela Iorio & Giuseppe Pandolfo & Roberta Siciliano, 2021. "Adjusted Concordance Index: an Extensionl of the Adjusted Rand Index to Fuzzy Partitions," Journal of Classification, Springer;The Classification Society, vol. 38(1), pages 112-128, April.
Michel Grabisch & Jean-Luc Marichal & Radko Mesiar & Endre Pap, 2009. "Aggregation functions," Université Paris1 Panthéon-Sorbonne (Post-Print and Working Papers) halshs-00445120, HAL.
- Michel Grabisch & Jean-Luc Marichal & Radko Mesiar & Endre Pap, 2009. "Aggregation functions," Post-Print halshs-00445120, HAL.
Irene Charon & Lucile Denoeud & Alain Guenoche & Olivier Hudry, 2006. "Maximum Transfer Distance Between Partitions," Journal of Classification, Springer;The Classification Society, vol. 23(1), pages 103-121, June.
Matthijs J. Warrens & Hanneke Hoef, 2022. "Understanding the Adjusted Rand Index and Other Partition Comparison Indices Based on Counting Object Pairs," Journal of Classification, Springer;The Classification Society, vol. 39(3), pages 487-509, November.
Glenn Milligan & Martha Cooper, 1985. "An examination of procedures for determining the number of clusters in a data set," Psychometrika, Springer;The Psychometric Society, vol. 50(2), pages 159-179, June.

Full references (including those not matched with items on IDEAS)

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Ryan DeWolfe & Jeffrey L. Andrews, 2025. "Random models for adjusting fuzzy rand index extensions," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 19(2), pages 361-385, June.
Marek Gagolewski & Anna Cena & Maciej Bartoszuk & Łukasz Brzozowski, 2025. "Clustering with Minimum Spanning Trees: How Good Can It Be?," Journal of Classification, Springer;The Classification Society, vol. 42(1), pages 90-112, March.
Li, Pai-Ling & Chiou, Jeng-Min, 2011. "Identifying cluster number for subspace projected functional data clustering," Computational Statistics & Data Analysis, Elsevier, vol. 55(6), pages 2090-2103, June.
Jesús Pineda & Sergi Masó-Orriols & Montse Masoliver & Joan Bertran & Mattias Goksör & Giovanni Volpe & Carlo Manzo, 2025. "Enhanced spatial clustering of single-molecule localizations with graph neural networks," Nature Communications, Nature, vol. 16(1), pages 1-17, December.
Ana Helena Tavares & Jakob Raymaekers & Peter J. Rousseeuw & Paula Brito & Vera Afreixo, 2020. "Clustering genomic words in human DNA using peaks and trends of distributions," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 14(1), pages 57-76, March.
Wang, Ketong & Porter, Michael D., 2018. "Optimal Bayesian clustering using non-negative matrix factorization," Computational Statistics & Data Analysis, Elsevier, vol. 128(C), pages 395-411.
Douglas L. Steinley & M. J. Brusco, 2019. "Using an Iterative Reallocation Partitioning Algorithm to Verify Test Multidimensionality," Journal of Classification, Springer;The Classification Society, vol. 36(3), pages 397-413, October.
José E. Chacón & Ana I. Rastrojo, 2023. "Minimum adjusted Rand index for two clusterings of a given size," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 17(1), pages 125-133, March.
Hennig, Christian, 2008. "Dissolution point and isolation robustness: Robustness criteria for general cluster analysis methods," Journal of Multivariate Analysis, Elsevier, vol. 99(6), pages 1154-1176, July.
J. Fernando Vera & Rodrigo Macías, 2021. "On the Behaviour of K-Means Clustering of a Dissimilarity Matrix by Means of Full Multidimensional Scaling," Psychometrika, Springer;The Psychometric Society, vol. 86(2), pages 489-513, June.
Lucile Denœud, 2008. "Transfer distance between partitions," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 2(3), pages 279-294, December.
Michael Brusco & Douglas Steinley, 2015. "Affinity Propagation and Uncapacitated Facility Location Problems," Journal of Classification, Springer;The Classification Society, vol. 32(3), pages 443-480, October.
Giuseppe Pandolfo & Antonio D’ambrosio, 2023. "Clustering directional data through depth functions," Computational Statistics, Springer, vol. 38(3), pages 1487-1506, September.
Boztug, Yasemin & Reutterer, Thomas, 2008. "A combined approach for segment-specific market basket analysis," European Journal of Operational Research, Elsevier, vol. 187(1), pages 294-312, May.
Zhiguang Huo & Li Zhu & Tianzhou Ma & Hongcheng Liu & Song Han & Daiqing Liao & Jinying Zhao & George Tseng, 2020. "Two-Way Horizontal and Vertical Omics Integration for Disease Subtype Discovery," Statistics in Biosciences, Springer;International Chinese Statistical Association, vol. 12(1), pages 1-22, April.
Donggyu Kim & Minseog Oh, 2024. "Property of Inverse Covariance Matrix-based Financial Adjacency Matrix for Detecting Local Groups," Working Papers 202420, University of California at Riverside, Department of Economics.
Herz, Denise C. & Eastman, Andrea Lane & Suthar, Himal & McCroskey, Jacquelyn, 2025. "Exploring child welfare pathways and dual system involvement," Journal of Criminal Justice, Elsevier, vol. 98(C).
Wang, Xiaogang & Qiu, Weiliang & Zamar, Ruben H., 2007. "CLUES: A non-parametric clustering method based on local shrinking," Computational Statistics & Data Analysis, Elsevier, vol. 52(1), pages 286-298, September.
Julian Rossbroich & Jeffrey Durieux & Tom F. Wilderjans, 2022. "Model Selection Strategies for Determining the Optimal Number of Overlapping Clusters in Additive Overlapping Partitional Clustering," Journal of Classification, Springer;The Classification Society, vol. 39(2), pages 264-301, July.
Michio Yamamoto, 2012. "Clustering of functional data in a low-dimensional subspace," Advances in Data Analysis and Classification, Springer;German Classification Society - Gesellschaft für Klassifikation (GfKl);Japanese Classification Society (JCS);Classification and Data Analysis Group of the Italian Statistical Society (CLADAG);International Federation of Classification Societies (IFCS), vol. 6(3), pages 219-247, October.

More about this item

Keywords

; ; ; ; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:spr:jclass:v:42:y:2025:i:1:d:10.1007_s00357-024-09482-2. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Sonal Shukla or Springer Nature Abstracting and Indexing (email available below). General contact details of provider: http://www.springer.com .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Normalised Clustering Accuracy: An Asymmetric External Cluster Validity Measure

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data