Efficient Detection of Botnet Traffic by Features Selection and Decision Trees

[EN] Botnets are one of the online threats with the most significant presence, causing billionaire losses to global economies. Nowadays, the increasing number of devices connected to the Internet makes it necessary to analyze extensive network traffic data. In this work, we focus on increasing the p...

ver descrição completa

Detalhes bibliográficos
Autores: Velasco Mata, Javier, González Castro, Víctor, Fidalgo Fernández, Eduardo, Alegre Gutiérrez, Enrique
Formato: artículo
Estado:Versión actualizada desde la publicación
Fecha de publicación:2021
País:España
Recursos:Universidad de León
Repositorio:BULERIA. Repositorio Institucional de la Universidad de León
OAI Identifier:oai:buleria.unileon.es:10612/22946
Acesso em linha:https://ieeexplore.ieee.org/document/9523853
https://hdl.handle.net/10612/22946
Access Level:acceso abierto
Palavra-chave:Informática
Ingeniería de sistemas
Computational and artificial intelligence
Computer applications
Decision tree
Machine learning algorithms
Optimization methods
1203.04 Inteligencia Artificial
1207.03 Cibernética
1209.03 Análisis de Datos
1203.17 Informática
Descrição
Resumo:[EN] Botnets are one of the online threats with the most significant presence, causing billionaire losses to global economies. Nowadays, the increasing number of devices connected to the Internet makes it necessary to analyze extensive network traffic data. In this work, we focus on increasing the performance of botnet traffic classification by selecting those features that further increase the detection rate. For this purpose, we use two feature selection techniques, i.e., Information Gain and Gini Importance, which led to three pre-selected subsets of five, six and seven features. Then, we evaluate the three feature subsets and three models, i.e., Decision Tree, Random Forest and k-Nearest Neighbors. To test the performance of the three feature vectors and the three models, we generate two datasets based on the CTU-13 dataset, namely QB-CTU13 and EQB-CTU13. Finally, we measure the performance as the macro averaged F1 score over the computational time required to classify a sample. The results show that the highest performance is achieved by Decision Trees using a five feature set, which obtained a mean F1 score of 85% classifying each sample in an average time of 0.78 microseconds.