Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
[EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online...
| Autores: | , , , |
|---|---|
| Tipo de recurso: | artículo |
| Fecha de publicación: | 2020 |
| País: | España |
| Institución: | Universitat Politècnica de València (UPV) |
| Repositorio: | RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| Idioma: | inglés |
| OAI Identifier: | oai:riunet.upv.es:10251/171117 |
| Acceso en línea: | https://riunet.upv.es/handle/10251/171117 |
| Access Level: | acceso abierto |
| Palabra clave: | Deception detection Crowdsourcing Ground truth Probabilistic labeling LENGUAJES Y SISTEMAS INFORMATICOS |
| id |
ES_ebc7e8d513ab866b9fc7bbc81616475b |
|---|---|
| oai_identifier_str |
oai:riunet.upv.es:10251/171117 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| dc.title.none.fl_str_mv |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| title |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| spellingShingle |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? Fornaciari, Tommaso Deception detection Crowdsourcing Ground truth Probabilistic labeling LENGUAJES Y SISTEMAS INFORMATICOS |
| title_short |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| title_full |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| title_fullStr |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| title_full_unstemmed |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| title_sort |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data? |
| dc.creator.none.fl_str_mv |
Fornaciari, Tommaso Cagnina, Leticia Poesio, Massimo Rosso, Paolo |
| author |
Fornaciari, Tommaso |
| author_facet |
Fornaciari, Tommaso Cagnina, Leticia Poesio, Massimo Rosso, Paolo |
| author_role |
author |
| author2 |
Cagnina, Leticia Poesio, Massimo Rosso, Paolo |
| author2_role |
author author author |
| dc.contributor.none.fl_str_mv |
Departamento de Sistemas Informáticos y Computación Escuela Técnica Superior de Ingeniería Informática Centro de Investigación Pattern Recognition and Human Language Technology Agencia Estatal de Investigación UK Research and Innovation European Regional Development Fund Ministerio de Economía y Competitividad Economic and Social Research Council, Reino Unido Consejo Nacional de Investigaciones Científicas y Técnicas, Argentina Repositorio Institucional de la Universitat Politècnica de València Riunet |
| dc.subject.none.fl_str_mv |
Deception detection Crowdsourcing Ground truth Probabilistic labeling LENGUAJES Y SISTEMAS INFORMATICOS |
| topic |
Deception detection Crowdsourcing Ground truth Probabilistic labeling LENGUAJES Y SISTEMAS INFORMATICOS |
| description |
[EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and adding them to a set of deceptive reviews obtained through crowdsourcing services. Models trained this way are generally successful at discriminating between `genuine¿ online reviews and the crowdsourced deceptive reviews. It has been argued that the deceptive reviews obtained via crowdsourcing are very different from real fake reviews, but the claim has never been properly tested. In this paper, we compare (false) crowdsourced reviews with a set of `real¿ fake reviews published on line. We evaluate their degree of similarity and their usefulness in training models for the detection of untrustworthy reviews. We find that the deceptive reviews collected via crowdsourcing are significantly different from the fake reviews published online. In the case of the artificially produced deceptive texts, it turns out that their domain similarity with the targets affects the models¿ performance, much more than their untruthfulness. This suggests that the use of crowdsourced datasets for opinion spam detection may not result in models applicable to the real task of detecting deceptive reviews. As an alternative method to create large-size datasets for the fake reviews detection task, we propose methods based on the probabilistic annotation of unlabeled texts, relying on the use of meta-information generally available on the e-commerce sites. Such methods are independent from the content of the reviews and allow to train reliable models for the detection of fake reviews. |
| publishDate |
2020 |
| dc.date.none.fl_str_mv |
2020 2020-12-01 |
| dc.type.none.fl_str_mv |
journal article http://purl.org/coar/resource_type/c_6501 VoR http://purl.org/coar/version/c_970fb48d4fbd8a85 |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/article |
| format |
article |
| dc.identifier.none.fl_str_mv |
https://riunet.upv.es/handle/10251/171117 |
| url |
https://riunet.upv.es/handle/10251/171117 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.relation.none.fl_str_mv |
UK Research and Innovation https://doi.org/10.13039/100014013 ES%2FM010236%2F1 Human Rights and Information Technology in the Era of Big Data Ministerio de Economía y Competitividad http://dx.doi.org/10.13039/501100003329 TIN2015-71147-C2-1-P COMPRENSION DEL LENGUAJE EN LOS MEDIOS DE COMUNICACION SOCIAL - REPRESENTANDO CONTEXTOS DE FORMA CONTINUA Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 PGC2018-096212-B-C31 DESINFORMACION Y AGRESIVIDAD EN SOCIAL MEDIA: AGREGANDO INFORMACION Y ANALIZANDO EL LENGUAJE |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Reserva de todos los derechos http://rightsstatements.org/vocab/InC/1.0/ |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Reserva de todos los derechos http://rightsstatements.org/vocab/InC/1.0/ |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.publisher.none.fl_str_mv |
Springer-Verlag |
| publisher.none.fl_str_mv |
Springer-Verlag |
| dc.source.none.fl_str_mv |
reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia instname:Universitat Politècnica de València (UPV) |
| instname_str |
Universitat Politècnica de València (UPV) |
| reponame_str |
RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| collection |
RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869423257783894016 |
| spelling |
Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?Fornaciari, TommasoCagnina, LeticiaPoesio, MassimoRosso, PaoloDeception detectionCrowdsourcingGround truthProbabilistic labelingLENGUAJES Y SISTEMAS INFORMATICOS[EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and adding them to a set of deceptive reviews obtained through crowdsourcing services. Models trained this way are generally successful at discriminating between `genuine¿ online reviews and the crowdsourced deceptive reviews. It has been argued that the deceptive reviews obtained via crowdsourcing are very different from real fake reviews, but the claim has never been properly tested. In this paper, we compare (false) crowdsourced reviews with a set of `real¿ fake reviews published on line. We evaluate their degree of similarity and their usefulness in training models for the detection of untrustworthy reviews. We find that the deceptive reviews collected via crowdsourcing are significantly different from the fake reviews published online. In the case of the artificially produced deceptive texts, it turns out that their domain similarity with the targets affects the models¿ performance, much more than their untruthfulness. This suggests that the use of crowdsourced datasets for opinion spam detection may not result in models applicable to the real task of detecting deceptive reviews. As an alternative method to create large-size datasets for the fake reviews detection task, we propose methods based on the probabilistic annotation of unlabeled texts, relying on the use of meta-information generally available on the e-commerce sites. Such methods are independent from the content of the reviews and allow to train reliable models for the detection of fake reviews.Leticia Cagnina thanks CONICET for the continued financial support. This work was funded by MINECO/FEDER (Grant No. SomEMBED TIN2015-71147-C2-1-P). The work of Paolo Rosso was partially funded by the MISMIS-FAKEnHATE Spanish MICINN research project (PGC2018-096212-B-C31). Massimo Poesio was in part supported by the UK Economic and Social Research Council (Grant Number ES/M010236/1).Springer-VerlagDepartamento de Sistemas Informáticos y ComputaciónEscuela Técnica Superior de Ingeniería InformáticaCentro de Investigación Pattern Recognition and Human Language TechnologyAgencia Estatal de InvestigaciónUK Research and InnovationEuropean Regional Development FundMinisterio de Economía y CompetitividadEconomic and Social Research Council, Reino UnidoConsejo Nacional de Investigaciones Científicas y Técnicas, ArgentinaRepositorio Institucional de la Universitat Politècnica de València Riunet20202020-12-01journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfapplication/pdfhttps://riunet.upv.es/handle/10251/171117reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valénciainstname:Universitat Politècnica de València (UPV)InglésengUK Research and Innovation https://doi.org/10.13039/100014013 ES%2FM010236%2F1 Human Rights and Information Technology in the Era of Big DataMinisterio de Economía y Competitividad http://dx.doi.org/10.13039/501100003329 TIN2015-71147-C2-1-P COMPRENSION DEL LENGUAJE EN LOS MEDIOS DE COMUNICACION SOCIAL - REPRESENTANDO CONTEXTOS DE FORMA CONTINUAAgencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 PGC2018-096212-B-C31 DESINFORMACION Y AGRESIVIDAD EN SOCIAL MEDIA: AGREGANDO INFORMACION Y ANALIZANDO EL LENGUAJEopen accesshttp://purl.org/coar/access_right/c_abf2Reserva de todos los derechoshttp://rightsstatements.org/vocab/InC/1.0/info:eu-repo/semantics/openAccessoai:riunet.upv.es:10251/1711172026-06-13T07:49:27Z |
| score |
15.301629 |