Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities

In an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for...

ver descrição completa

Detalhes bibliográficos
Autores: Perdices Burrero, Daniel, Ramos de Santiago, Fco. Javier, García Dorado, José Luis, González Martínez, Iván, López de Vergara Méndez, Jorge Enrique
Formato: artículo
Fecha de publicación:2021
País:España
Recursos:Universidad Autónoma de Madrid
Repositorio:Biblos-e Archivo. Repositorio Institucional de la UAM
Idioma:inglés
OAI Identifier:oai:repositorio.uam.es:10486/700706
Acesso em linha:http://hdl.handle.net/10486/700706
https://dx.doi.org/10.1016/j.comnet.2021.108357
Access Level:acceso abierto
Palavra-chave:Deep learning
Internet monitoring
Natural language processing
Traffic monetization
Users analytics
Web browsing
Telecomunicaciones
id ES_b115d2a0649ca6edace538e564fc9b2c
oai_identifier_str oai:repositorio.uam.es:10486/700706
network_acronym_str ES
network_name_str España
repository_id_str
spelling Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunitiesPerdices Burrero, DanielRamos de Santiago, Fco. JavierGarcía Dorado, José LuisGonzález Martínez, IvánLópez de Vergara Méndez, Jorge EnriqueDeep learningInternet monitoringNatural language processingTraffic monetizationUsers analyticsWeb browsingTelecomunicacionesIn an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks’ links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF–IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrizationThis research has been partially funded by the Spanish State Research Agency under the project AgileMon (AEI PID2019-104451RBC21) and by the Spanish Ministry of Science, Innovation and Universities under the program for the training of university lecturers (Grant number: FPU19/05678)ElsevierDepartamento de Tecnología Electrónica y de las ComunicacionesEscuela Politécnica Superior20212021-10-24research articlehttp://purl.org/coar/resource_type/c_2df8fbb1VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfhttp://hdl.handle.net/10486/700706https://dx.doi.org/10.1016/j.comnet.2021.108357reponame:Biblos-e Archivo. Repositorio Institucional de la UAMinstname:Universidad Autónoma de MadridInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:repositorio.uam.es:10486/7007062026-06-23T12:46:27Z
dc.title.none.fl_str_mv Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
title Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
spellingShingle Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
Perdices Burrero, Daniel
Deep learning
Internet monitoring
Natural language processing
Traffic monetization
Users analytics
Web browsing
Telecomunicaciones
title_short Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
title_full Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
title_fullStr Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
title_full_unstemmed Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
title_sort Natural language processing for web browsing analytics: Challenges, lessons learned, and opportunities
dc.creator.none.fl_str_mv Perdices Burrero, Daniel
Ramos de Santiago, Fco. Javier
García Dorado, José Luis
González Martínez, Iván
López de Vergara Méndez, Jorge Enrique
author Perdices Burrero, Daniel
author_facet Perdices Burrero, Daniel
Ramos de Santiago, Fco. Javier
García Dorado, José Luis
González Martínez, Iván
López de Vergara Méndez, Jorge Enrique
author_role author
author2 Ramos de Santiago, Fco. Javier
García Dorado, José Luis
González Martínez, Iván
López de Vergara Méndez, Jorge Enrique
author2_role author
author
author
author
dc.contributor.none.fl_str_mv Departamento de Tecnología Electrónica y de las Comunicaciones
Escuela Politécnica Superior
dc.subject.none.fl_str_mv Deep learning
Internet monitoring
Natural language processing
Traffic monetization
Users analytics
Web browsing
Telecomunicaciones
topic Deep learning
Internet monitoring
Natural language processing
Traffic monetization
Users analytics
Web browsing
Telecomunicaciones
description In an Internet arena where the search engines and other digital marketing firms’ revenues peak, other actors still have open opportunities to monetize their users’ data. After the convenient anonymization, aggregation, and agreement, the set of websites users visit may result in exploitable data for ISPs. Uses cover from assessing the scope of advertising campaigns to reinforcing user fidelity among other marketing approaches, as well as security issues. However, sniffers based on HTTP, DNS, TLS or flow features do not suffice for this task. Modern websites are designed for preloading and prefetching some contents in addition to embedding banners, social networks’ links, images, and scripts from other websites. This self-triggered traffic makes it confusing to assess which websites users visited on purpose. Moreover, DNS caches prevent some queries of actively visited websites to be even sent. On this limited input, we propose to handle such domains as words and the sequences of domains as documents. This way, it is possible to identify the visited websites by translating this problem to a text classification context and applying the most promising techniques of the natural language processing and neural networks fields. After applying different representation methods such as TF–IDF, Word2vec, Doc2vec, and custom neural networks in diverse scenarios and with several datasets, we can state websites visited on purpose with accuracy figures over 90%, with peaks close to 100%, being processes that are fully automated and free of any human parametrization
publishDate 2021
dc.date.none.fl_str_mv 2021
2021-10-24
dc.type.none.fl_str_mv research article
http://purl.org/coar/resource_type/c_2df8fbb1
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv http://hdl.handle.net/10486/700706
https://dx.doi.org/10.1016/j.comnet.2021.108357
url http://hdl.handle.net/10486/700706
https://dx.doi.org/10.1016/j.comnet.2021.108357
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:Biblos-e Archivo. Repositorio Institucional de la UAM
instname:Universidad Autónoma de Madrid
instname_str Universidad Autónoma de Madrid
reponame_str Biblos-e Archivo. Repositorio Institucional de la UAM
collection Biblos-e Archivo. Repositorio Institucional de la UAM
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869416881404772352
score 15.301603