Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning

Current state-of-the-art research for tackling the problem of malware detection and classification is centered on the design, implementation and deployment of systems powered by machine learning because of its ability to generalize to never-before-seen malware families and polymorphic mutations. How...

Descripción completa

Detalles Bibliográficos
Autores: Gibert Llauradó, Daniel, Fredrikson, Matt, Mateu Piñol, Carles, Planes Cid, Jordi
Tipo de recurso: artículo
Estado:Versión enviada para evaluación y publicación
Fecha de publicación:2022
País:España
Institución:Universitat de Lleida (UdL)
Repositorio:Repositori Obert UdL
OAI Identifier:oai:repositori.udl.cat:10459.1/72778
Acceso en línea:https://doi.org/10.1016/j.cose.2021.102543
http://hdl.handle.net/10459.1/72778
Access Level:acceso abierto
Palabra clave:Malware Classification
Assembly Language Source Code
Obfuscation
Reinforcement Learning
Deep Q-Network
id ES_7ac1ec36c33d02e6600aeee699653ac0
oai_identifier_str oai:repositori.udl.cat:10459.1/72778
network_acronym_str ES
network_name_str España
repository_id_str
spelling Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learningGibert Llauradó, DanielFredrikson, MattMateu Piñol, CarlesPlanes Cid, JordiMalware ClassificationAssembly Language Source CodeObfuscationReinforcement LearningDeep Q-NetworkCurrent state-of-the-art research for tackling the problem of malware detection and classification is centered on the design, implementation and deployment of systems powered by machine learning because of its ability to generalize to never-before-seen malware families and polymorphic mutations. However, it has been shown that machine learning models, in partidular deep neural networks, lack robustness against crafted inputs (adversarial examples). In this work, we have investigated the vulnerability of a state-of-the-art shallow convolutional neural network malware classifier against the deat code insertion technique. We propose a general framework powered by a Double Q-network to induce misclassification over malware families. The framework trains an agent through a convolutional neural network to select the optimal positions in a code sequence to insert dead code instructions so that the machine learning classifier mislabels the resulting executable. The experiments show that the proposed method significantly drops the classification accuracy of the classifier to 56.53% while having an evasion rate of 100% for the samples belonging to Kelihos_ver3, Simda, and Kelihos_ver1 families. In addition, the average number of instructions needed to mislabel malware in comparison to a random agent decreased by 33%.This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No. 847402. This research has been partially funded by the Spanish MICINN Projects TIN2015-71799-C2-2-P, ENE2015-64117-C5-1-R, PID2019-111544GB-C22, and supported by the University of Lleida.Elsevier2022info:eu-repo/semantics/articleinfo:eu-repo/semantics/submittedVersionhttps://doi.org/10.1016/j.cose.2021.102543http://hdl.handle.net/10459.1/72778reponame:Repositori Obert UdL instname:Universitat de Lleida (UdL)Inglésinfo:eu-repo/grantAgreement/MINECO//TIN2015-71799-C2-2-Pinfo:eu-repo/grantAgreement/MINECO//ENE2015-64117-C5-1-Rinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/PID2019-111544GB-C22Versió preprint del document publicat a https://doi.org/10.1016/j.cose.2021.102543Computers and Security, 2022, vol. 113, 102543info:eu-repo/grantAgreement/EC/H2020/847402(c) Elsevier, 2021info:eu-repo/semantics/openAccessoai:repositori.udl.cat:10459.1/727782026-06-24T12:42:17Z
dc.title.none.fl_str_mv Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
title Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
spellingShingle Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
Gibert Llauradó, Daniel
Malware Classification
Assembly Language Source Code
Obfuscation
Reinforcement Learning
Deep Q-Network
title_short Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
title_full Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
title_fullStr Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
title_full_unstemmed Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
title_sort Enhancing the insertion of NOP instructions to obfuscate malware via deep reinforcement learning
dc.creator.none.fl_str_mv Gibert Llauradó, Daniel
Fredrikson, Matt
Mateu Piñol, Carles
Planes Cid, Jordi
author Gibert Llauradó, Daniel
author_facet Gibert Llauradó, Daniel
Fredrikson, Matt
Mateu Piñol, Carles
Planes Cid, Jordi
author_role author
author2 Fredrikson, Matt
Mateu Piñol, Carles
Planes Cid, Jordi
author2_role author
author
author
dc.subject.none.fl_str_mv Malware Classification
Assembly Language Source Code
Obfuscation
Reinforcement Learning
Deep Q-Network
topic Malware Classification
Assembly Language Source Code
Obfuscation
Reinforcement Learning
Deep Q-Network
description Current state-of-the-art research for tackling the problem of malware detection and classification is centered on the design, implementation and deployment of systems powered by machine learning because of its ability to generalize to never-before-seen malware families and polymorphic mutations. However, it has been shown that machine learning models, in partidular deep neural networks, lack robustness against crafted inputs (adversarial examples). In this work, we have investigated the vulnerability of a state-of-the-art shallow convolutional neural network malware classifier against the deat code insertion technique. We propose a general framework powered by a Double Q-network to induce misclassification over malware families. The framework trains an agent through a convolutional neural network to select the optimal positions in a code sequence to insert dead code instructions so that the machine learning classifier mislabels the resulting executable. The experiments show that the proposed method significantly drops the classification accuracy of the classifier to 56.53% while having an evasion rate of 100% for the samples belonging to Kelihos_ver3, Simda, and Kelihos_ver1 families. In addition, the average number of instructions needed to mislabel malware in comparison to a random agent decreased by 33%.
publishDate 2022
dc.date.none.fl_str_mv 2022
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/submittedVersion
format article
status_str submittedVersion
dc.identifier.none.fl_str_mv https://doi.org/10.1016/j.cose.2021.102543
http://hdl.handle.net/10459.1/72778
url https://doi.org/10.1016/j.cose.2021.102543
http://hdl.handle.net/10459.1/72778
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv info:eu-repo/grantAgreement/MINECO//TIN2015-71799-C2-2-P
info:eu-repo/grantAgreement/MINECO//ENE2015-64117-C5-1-R
info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/PID2019-111544GB-C22
Versió preprint del document publicat a https://doi.org/10.1016/j.cose.2021.102543
Computers and Security, 2022, vol. 113, 102543
info:eu-repo/grantAgreement/EC/H2020/847402
dc.rights.none.fl_str_mv (c) Elsevier, 2021
info:eu-repo/semantics/openAccess
rights_invalid_str_mv (c) Elsevier, 2021
eu_rights_str_mv openAccess
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Repositori Obert UdL
instname:Universitat de Lleida (UdL)
instname_str Universitat de Lleida (UdL)
reponame_str Repositori Obert UdL
collection Repositori Obert UdL
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869411463159873536
score 15.812429