Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
In an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for...
| Autores: | , , , , |
|---|---|
| Formato: | artículo |
| Fecha de publicación: | 2021 |
| País: | España |
| Recursos: | Universidad Autónoma de Madrid |
| Repositorio: | Biblos-e Archivo. Repositorio Institucional de la UAM |
| Idioma: | inglés |
| OAI Identifier: | oai:repositorio.uam.es:10486/700706 |
| Acesso em linha: | http://hdl.handle.net/10486/700706 https://dx.doi.org/10.1016/j.comnet.2021.108357 |
| Access Level: | acceso abierto |
| Palavra-chave: | Deep learning Internet monitoring Natural language processing Traffic monetization Users analytics Web browsing Telecomunicaciones |
| id |
ES_b115d2a0649ca6edace538e564fc9b2c |
|---|---|
| oai_identifier_str |
oai:repositorio.uam.es:10486/700706 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunitiesPerdices Burrero, DanielRamos de Santiago, Fco. JavierGarcía Dorado, José LuisGonzález Martínez, IvánLópez de Vergara Méndez, Jorge EnriqueDeep learningInternet monitoringNatural language processingTraffic monetizationUsers analyticsWeb browsingTelecomunicacionesIn an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks’ links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF–IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrizationThis research has been partially funded by the Spanish State Research Agency under the project AgileMon (AEI PID2019-104451RBC21) and by the Spanish Ministry of Science, Innovation and Universities under the program for the training of university lecturers (Grant number: FPU19/05678)ElsevierDepartamento de Tecnología Electrónica y de las ComunicacionesEscuela Politécnica Superior20212021-10-24research articlehttp://purl.org/coar/resource_type/c_2df8fbb1VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10486/700706https://dx.doi.org/10.1016/j.comnet.2021.108357reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:repositorio.uam.es:10486/7007062026-06-23T12:46:27Z |
| dc.title.none.fl_str_mv |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| title |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| spellingShingle |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities Perdices Burrero, Daniel Deep learning Internet monitoring Natural language processing Traffic monetization Users analytics Web browsing Telecomunicaciones |
| title_short |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| title_full |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| title_fullStr |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| title_full_unstemmed |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| title_sort |
Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities |
| dc.creator.none.fl_str_mv |
Perdices Burrero, Daniel Ramos de Santiago, Fco. Javier García Dorado, José Luis González Martínez, Iván López de Vergara Méndez, Jorge Enrique |
| author |
Perdices Burrero, Daniel |
| author_facet |
Perdices Burrero, Daniel Ramos de Santiago, Fco. Javier García Dorado, José Luis González Martínez, Iván López de Vergara Méndez, Jorge Enrique |
| author_role |
author |
| author2 |
Ramos de Santiago, Fco. Javier García Dorado, José Luis González Martínez, Iván López de Vergara Méndez, Jorge Enrique |
| author2_role |
author author author author |
| dc.contributor.none.fl_str_mv |
Departamento de Tecnología Electrónica y de las Comunicaciones Escuela Politécnica Superior |
| dc.subject.none.fl_str_mv |
Deep learning Internet monitoring Natural language processing Traffic monetization Users analytics Web browsing Telecomunicaciones |
| topic |
Deep learning Internet monitoring Natural language processing Traffic monetization Users analytics Web browsing Telecomunicaciones |
| description |
In an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks’ links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF–IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrization |
| publishDate |
2021 |
| dc.date.none.fl_str_mv |
2021 2021-10-24 |
| dc.type.none.fl_str_mv |
research article http://purl.org/coar/resource_type/c_2df8fbb1 VoR http://purl.org/coar/version/c_970fb48d4fbd8a85 |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/article |
| format |
article |
| dc.identifier.none.fl_str_mv |
http://hdl.handle.net/10486/700706 https://dx.doi.org/10.1016/j.comnet.2021.108357 |
| url |
http://hdl.handle.net/10486/700706 https://dx.doi.org/10.1016/j.comnet.2021.108357 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf |
| dc.publisher.none.fl_str_mv |
Elsevier |
| publisher.none.fl_str_mv |
Elsevier |
| dc.source.none.fl_str_mv |
reponame:Biblos-e Archivo. Repositorio Institucional de la UAM instname:Universidad Autónoma de Madrid |
| instname_str |
Universidad Autónoma de Madrid |
| reponame_str |
Biblos-e Archivo. Repositorio Institucional de la UAM |
| collection |
Biblos-e Archivo. Repositorio Institucional de la UAM |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869416881404772352 |
| score |
15.301603 |