Experimental evaluation of ensemble classifiers for imbalance in Big Data

Datasets are growing in size and complexity at a pace never seen before, forming ever larger datasets known as Big Data. A common problem for classification, especially in Big Data, is that the numerous examples of the different classes might not be balanced. Some decades ago, imbalanced classificat...

Descripción completa

Detalles Bibliográficos
Autores: Juez Gil, Mario, Arnaiz González, Álvar, Rodríguez Diez, Juan José, García Osorio, César
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2021
País:España
Institución:Universidad de Burgos (UBU)
Repositorio:Repositorio Institucional de la Universidad de Burgos (RIUBU)
OAI Identifier:oai:riubu.ubu.es:10259/5766
Acceso en línea:http://hdl.handle.net/10259/5766
Access Level:acceso abierto
Palabra clave:Unbalance
Imbalance
Ensemble
Resampling
Big Data
Spark
Informática
Computer science
id ES_3c1c524da290a58cd0ffda12a355df82
oai_identifier_str oai:riubu.ubu.es:10259/5766
network_acronym_str ES
network_name_str España
repository_id_str
spelling Experimental evaluation of ensemble classifiers for imbalance in Big DataJuez Gil, MarioArnaiz González, ÁlvarRodríguez Diez, Juan JoséGarcía Osorio, CésarUnbalanceImbalanceEnsembleResamplingBig DataSparkInformáticaComputer scienceDatasets are growing in size and complexity at a pace never seen before, forming ever larger datasets known as Big Data. A common problem for classification, especially in Big Data, is that the numerous examples of the different classes might not be balanced. Some decades ago, imbalanced classification was therefore introduced, to correct the tendency of classifiers that show bias in favor of the majority class and that ignore the minority one. To date, although the number of imbalanced classification methods have increased, they continue to focus on normal-sized datasets and not on the new reality of Big Data. In this paper, in-depth experimentation with ensemble classifiers is conducted in the context of imbalanced Big Data classification, using two popular ensemble families (Bagging and Boosting) and different resampling methods. All the experimentation was launched in Spark clusters, comparing ensemble performance and execution times with statistical test results, including the newest ones based on the Bayesian approach. One very interesting conclusion from the study was that simpler methods applied to unbalanced datasets in the context of Big Data provided better results than complex methods. The additional complexity of some of the sophisticated methods, which appear necessary to process and to reduce imbalance in normal-sized datasets were not effective for imbalanced Big Data.“la Caixa” Foundation, Spain, under agreement LCF/PR/PR18/51130007. This work was supported by the Junta de Castilla y León, Spain under project BU055P20 (JCyL/FEDER, UE) co-financed through European Union FEDER funds, and by the Consejería de Educación of the Junta de Castilla y León and the European Social Fund, Spain through a pre-doctoral grant (EDU/1100/2017).Elsevier202120212021info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfhttp://hdl.handle.net/10259/5766reponame:Repositorio Institucional de la Universidad de Burgos (RIUBU)instname:Universidad de Burgos (UBU)InglésApplied Soft Computing. 2021, V. 108, 107447https://doi.org/10.1016/j.asoc.2021.107447info:eu-repo/grantAgreement/Fundación Bancaria Caixa d'Estalvis i Pensions de Barcelona//LCF%2FPR%2FPR18%2F51130007info:eu-repo/grantAgreement/Junta de Castilla y León//BU055P20Attribution-NonCommercial-NoDerivatives 4.0 Internacionalhttp://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:riubu.ubu.es:10259/57662026-05-28T07:56:11Z
dc.title.none.fl_str_mv Experimental evaluation of ensemble classifiers for imbalance in Big Data
title Experimental evaluation of ensemble classifiers for imbalance in Big Data
spellingShingle Experimental evaluation of ensemble classifiers for imbalance in Big Data
Juez Gil, Mario
Unbalance
Imbalance
Ensemble
Resampling
Big Data
Spark
Informática
Computer science
title_short Experimental evaluation of ensemble classifiers for imbalance in Big Data
title_full Experimental evaluation of ensemble classifiers for imbalance in Big Data
title_fullStr Experimental evaluation of ensemble classifiers for imbalance in Big Data
title_full_unstemmed Experimental evaluation of ensemble classifiers for imbalance in Big Data
title_sort Experimental evaluation of ensemble classifiers for imbalance in Big Data
dc.creator.none.fl_str_mv Juez Gil, Mario
Arnaiz González, Álvar
Rodríguez Diez, Juan José
García Osorio, César
author Juez Gil, Mario
author_facet Juez Gil, Mario
Arnaiz González, Álvar
Rodríguez Diez, Juan José
García Osorio, César
author_role author
author2 Arnaiz González, Álvar
Rodríguez Diez, Juan José
García Osorio, César
author2_role author
author
author
dc.subject.none.fl_str_mv Unbalance
Imbalance
Ensemble
Resampling
Big Data
Spark
Informática
Computer science
topic Unbalance
Imbalance
Ensemble
Resampling
Big Data
Spark
Informática
Computer science
description Datasets are growing in size and complexity at a pace never seen before, forming ever larger datasets known as Big Data. A common problem for classification, especially in Big Data, is that the numerous examples of the different classes might not be balanced. Some decades ago, imbalanced classification was therefore introduced, to correct the tendency of classifiers that show bias in favor of the majority class and that ignore the minority one. To date, although the number of imbalanced classification methods have increased, they continue to focus on normal-sized datasets and not on the new reality of Big Data. In this paper, in-depth experimentation with ensemble classifiers is conducted in the context of imbalanced Big Data classification, using two popular ensemble families (Bagging and Boosting) and different resampling methods. All the experimentation was launched in Spark clusters, comparing ensemble performance and execution times with statistical test results, including the newest ones based on the Bayesian approach. One very interesting conclusion from the study was that simpler methods applied to unbalanced datasets in the context of Big Data provided better results than complex methods. The additional complexity of some of the sophisticated methods, which appear necessary to process and to reduce imbalance in normal-sized datasets were not effective for imbalanced Big Data.
publishDate 2021
dc.date.none.fl_str_mv 2021
2021
2021
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv http://hdl.handle.net/10259/5766
url http://hdl.handle.net/10259/5766
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Applied Soft Computing. 2021, V. 108, 107447
https://doi.org/10.1016/j.asoc.2021.107447
info:eu-repo/grantAgreement/Fundación Bancaria Caixa d'Estalvis i Pensions de Barcelona//LCF%2FPR%2FPR18%2F51130007
info:eu-repo/grantAgreement/Junta de Castilla y León//BU055P20
dc.rights.none.fl_str_mv Attribution-NonCommercial-NoDerivatives 4.0 Internacional
http://creativecommons.org/licenses/by-nc-nd/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv Attribution-NonCommercial-NoDerivatives 4.0 Internacional
http://creativecommons.org/licenses/by-nc-nd/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Repositorio Institucional de la Universidad de Burgos (RIUBU)
instname:Universidad de Burgos (UBU)
instname_str Universidad de Burgos (UBU)
reponame_str Repositorio Institucional de la Universidad de Burgos (RIUBU)
collection Repositorio Institucional de la Universidad de Burgos (RIUBU)
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869406348632915968
score 15,301603