Towards a fairer vision: addressing class imbalance in image gender recognition

This paper addresses the growing integration of machine learning and computer vision with a particular emphasis on classifying facial features. Although these technologies hold great potential, there are concerns regarding biases that must be mitigated to ensure more impartial results. This study pr...

Descripción completa

Detalles Bibliográficos
Autores: Barahona, Daniel, Quijano Sánchez, Lara, Liberatore, Federico
Tipo de recurso: artículo
Fecha de publicación:2025
País:España
Institución:Universidad Autónoma de Madrid
Repositorio:Biblos-e Archivo. Repositorio Institucional de la UAM
Idioma:inglés
OAI Identifier:oai:repositorio.uam.es:10486/728980
Acceso en línea:https://hdl.handle.net/10486/728980
https://dx.doi.org/10.1007/s00138-025-01732-6
Access Level:acceso abierto
Palabra clave:Bias Mitigation
Class Imbalance
Human Data
Informática
Descripción
Sumario:This paper addresses the growing integration of machine learning and computer vision with a particular emphasis on classifying facial features. Although these technologies hold great potential, there are concerns regarding biases that must be mitigated to ensure more impartial results. This study primarily addresses the issue of bias caused by class imbalance, which can be mitigated through algorithmic or data-level approaches. Notably, the literature presents gaps, including the lack of comprehensive studies comparing these two types of mitigation techniques and understanding the circumstances in which each is more suitable. Moreover, the influence of imbalance conditions and data complexity on mitigation meth¬ods remains underexplored. To address these gaps, this research formulates three key research questions and conducts experiments using two datasets, UTKFace and PlantVillage, known for varying complexities. Various imbalance scenarios are simulated in these datasets. Additionally, a novel algorithm-level mitigation method named “Diffuse Focal Loss” is introduced. Results indicate the high effectiveness of synthetic oversampling methods, specifically using the Wasserstein variant of Generative Adversarial Networks, compared to algorithm-level approaches. Among the latter, the proposed novel method outperforms others in terms of metrics. However, it is worth noting that algorithmic techniques are more practical and quicker to apply in low-complexity scenarios