When is resampling beneficial for feature selection with imbalanced wide data?

This paper studies the effects that combinations of balancing and feature selection techniques have on wide data (many more attributes than instances) when different classifiers are used. For this, an extensive study is done using 14 datasets, 3 balancing strategies, and 7 feature selection algorith...

Descripción completa

Detalles Bibliográficos
Autores: Ramos Pérez, Ismael, Arnaiz González, Álvar, Rodríguez Diez, Juan José, García Osorio, César
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2022
País:España
Institución:Universidad de Burgos (UBU)
Repositorio:Repositorio Institucional de la Universidad de Burgos (RIUBU)
OAI Identifier:oai:riubu.ubu.es:10259/7326
Acceso en línea:http://hdl.handle.net/10259/7326
Access Level:acceso abierto
Palabra clave:Feature selection
Wide data
High dimensional data
Very low sample size
Unbalanced
Machine learning
Informática
Computer science
Descripción
Sumario:This paper studies the effects that combinations of balancing and feature selection techniques have on wide data (many more attributes than instances) when different classifiers are used. For this, an extensive study is done using 14 datasets, 3 balancing strategies, and 7 feature selection algorithms. The evaluation is carried out using 5 classification algorithms, analyzing the results for different percentages of selected features, and establishing the statistical significance using Bayesian tests. Some general conclusions of the study are that it is better to use RUS before the feature selection, while ROS and SMOTE offer better results when applied afterwards. Additionally, specific results are also obtained depending on the classifier used, for example, for Gaussian SVM the best performance is obtained when the feature selection is done with SVM-RFE before balancing the data with RUS.