Estimation of partition similarity measures

The problem of measuring the similarity between two partitions of a set arises in many applications. For example, one way to assess the quality of a clustering method is to compare its output to a known ground truth partition. Unfortunately the sets that occur in tasks like document classification a...

Descripción completa

Detalles Bibliográficos
Autor: Sanz González, Hugo
Tipo de recurso: tesis de maestría
Fecha de publicación:2025
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/449910
Acceso en línea:https://hdl.handle.net/2117/449910
Access Level:acceso abierto
Palabra clave:Partitions (Mathematics)
Sampling (Statistics)
Similitud
Particions
Estimació
Mostreig
Biaix
Variança
Similarity
Partitions
Estimation
Sampling
Bias
Variance
Particions (Matemàtica)
Mostreig (Estadística)
Àrees temàtiques de la UPC::Matemàtiques i estadística::Estadística matemàtica::Inferència estadística
Descripción
Sumario:The problem of measuring the similarity between two partitions of a set arises in many applications. For example, one way to assess the quality of a clustering method is to compare its output to a known ground truth partition. Unfortunately the sets that occur in tasks like document classification and market segmentation are usually very large, making it infeasible for a human expert to construct such a ground truth. We investigate how to use a random sample of t elements or t element pairs from a set to estimate the similarity between two of its partitions. This formulation allows the expert to annotate just a small number of elements, or to judge whether the two members of each pair should go into the same cluster or not. Dozens of partition similarity measures have been invented, yet most of them are based on the same principles of pair counting and information theory. We find that virtually all the measures in the former group, and some of the most popular ones in the latter, can be estimated accurately and quickly. In particular, we show that their estimators either are unbiased or have a bias of O(1/t), and have a variance of O(1/t). Our derivations hold under rather weak hypotheses, allowing us to generalize the previous results to estimators of various similarity measures on sets, permutations, strings... Experiments with synthetic and real-life partitions suggest that the bounds on the bias and variance are reasonably tight. Armed with efficient methods for sampling k-combinations, k at most 2, the estimates can be obtained in O(t) time on the average, using O(t) space. These costs sometimes include an extra term of order c · c', where c and c' are the respective numbers of clusters in the two partitions being compared.