PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam
Spam web pages have become a problem for Information Retrieval systems due to the negative effects that this phenomenon can cause in their results. In this work we tackle the problem of detecting these pages with a propagation algorithm that, taking as input a web graph, chooses a set of spam and no...
| Autores: | , , , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión enviada para evaluación y publicación |
| Fecha de publicación: | 2012 |
| País: | España |
| Institución: | Universidad de Sevilla (US) |
| Repositorio: | idUS. Depósito de Investigación de la Universidad de Sevilla |
| OAI Identifier: | oai:idus.us.es:11441/130681 |
| Acceso en línea: | https://hdl.handle.net/11441/130681 |
| Access Level: | acceso abierto |
| Palabra clave: | Information retrieval Web spam detection Graph algorithms PageRank Web search |
| id |
ES_9e0b5cdee82e0841c145ccc02ee46d1a |
|---|---|
| oai_identifier_str |
oai:idus.us.es:11441/130681 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web SpamOrtega Rodríguez, Francisco JavierTroyano Jiménez, José AntonioCruz Mata, FermínGarcía Vallejo, Carlos AntonioInformation retrievalWeb spam detectionGraph algorithmsPageRankWeb searchSpam web pages have become a problem for Information Retrieval systems due to the negative effects that this phenomenon can cause in their results. In this work we tackle the problem of detecting these pages with a propagation algorithm that, taking as input a web graph, chooses a set of spam and not-spam web pages in order to spread their spam likelihood over the rest of the network. Thus we take advantage of the links between pages to obtain a ranking of pages according to their relevance and their spam likelihood. Our intuition consists in giving a high reputation to those pages related to relevant ones, and giving a high spam likelihood to the pages linked to spam web pages. We introduce the novelty of including the content of the web pages in the computation of an a priori estimation of the spam likelihood of the pages, and propagate this information. Our graph-based algorithm computes two scores for each node in the graph. Intuitively, these values represent how bad or good (spam-like or not) is a web page, according to its textual content and its relations in the graph. The experimental results show that our method outperforms other techniques for spam detectionMinisterio de Educación y Ciencia HUM2007-66607-C04-04ICIC InternationalLenguajes y Sistemas InformáticosTIC134: Sistemas InformáticosMinisterio de Educación y Ciencia (MEC). España2012info:eu-repo/semantics/articleinfo:eu-repo/semantics/submittedVersionapplication/pdfapplication/pdfhttps://hdl.handle.net/11441/130681reponame:idUS. Depósito de Investigación de la Universidad de Sevillainstname:Universidad de Sevilla (US)InglésInternational Journal of Innovative Computing, Information and Control, 8 (4), 2915-2928.HUM2007-66607-C04-04http://www.ijicic.org/contents.htminfo:eu-repo/semantics/openAccessoai:idus.us.es:11441/1306812026-06-17T12:51:07Z |
| dc.title.none.fl_str_mv |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| title |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| spellingShingle |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam Ortega Rodríguez, Francisco Javier Information retrieval Web spam detection Graph algorithms PageRank Web search |
| title_short |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| title_full |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| title_fullStr |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| title_full_unstemmed |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| title_sort |
PolaritySpam: Propagating Content-based Information Through a Web-Graph to Detect Web Spam |
| dc.creator.none.fl_str_mv |
Ortega Rodríguez, Francisco Javier Troyano Jiménez, José Antonio Cruz Mata, Fermín García Vallejo, Carlos Antonio |
| author |
Ortega Rodríguez, Francisco Javier |
| author_facet |
Ortega Rodríguez, Francisco Javier Troyano Jiménez, José Antonio Cruz Mata, Fermín García Vallejo, Carlos Antonio |
| author_role |
author |
| author2 |
Troyano Jiménez, José Antonio Cruz Mata, Fermín García Vallejo, Carlos Antonio |
| author2_role |
author author author |
| dc.contributor.none.fl_str_mv |
Lenguajes y Sistemas Informáticos TIC134: Sistemas Informáticos Ministerio de Educación y Ciencia (MEC). España |
| dc.subject.none.fl_str_mv |
Information retrieval Web spam detection Graph algorithms PageRank Web search |
| topic |
Information retrieval Web spam detection Graph algorithms PageRank Web search |
| description |
Spam web pages have become a problem for Information Retrieval systems due to the negative effects that this phenomenon can cause in their results. In this work we tackle the problem of detecting these pages with a propagation algorithm that, taking as input a web graph, chooses a set of spam and not-spam web pages in order to spread their spam likelihood over the rest of the network. Thus we take advantage of the links between pages to obtain a ranking of pages according to their relevance and their spam likelihood. Our intuition consists in giving a high reputation to those pages related to relevant ones, and giving a high spam likelihood to the pages linked to spam web pages. We introduce the novelty of including the content of the web pages in the computation of an a priori estimation of the spam likelihood of the pages, and propagate this information. Our graph-based algorithm computes two scores for each node in the graph. Intuitively, these values represent how bad or good (spam-like or not) is a web page, according to its textual content and its relations in the graph. The experimental results show that our method outperforms other techniques for spam detection |
| publishDate |
2012 |
| dc.date.none.fl_str_mv |
2012 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/submittedVersion |
| format |
article |
| status_str |
submittedVersion |
| dc.identifier.none.fl_str_mv |
https://hdl.handle.net/11441/130681 |
| url |
https://hdl.handle.net/11441/130681 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
International Journal of Innovative Computing, Information and Control, 8 (4), 2915-2928. HUM2007-66607-C04-04 http://www.ijicic.org/contents.htm |
| dc.rights.none.fl_str_mv |
info:eu-repo/semantics/openAccess |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.publisher.none.fl_str_mv |
ICIC International |
| publisher.none.fl_str_mv |
ICIC International |
| dc.source.none.fl_str_mv |
reponame:idUS. Depósito de Investigación de la Universidad de Sevilla instname:Universidad de Sevilla (US) |
| instname_str |
Universidad de Sevilla (US) |
| reponame_str |
idUS. Depósito de Investigación de la Universidad de Sevilla |
| collection |
idUS. Depósito de Investigación de la Universidad de Sevilla |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869414793715122176 |
| score |
15.301629 |