Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?

[EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online...

Descripción completa

Detalles Bibliográficos
Autores: Fornaciari, Tommaso, Cagnina, Leticia, Poesio, Massimo, Rosso, Paolo
Tipo de recurso: artículo
Fecha de publicación:2020
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/171117
Acceso en línea:https://riunet.upv.es/handle/10251/171117
Access Level:acceso abierto
Palabra clave:Deception detection
Crowdsourcing
Ground truth
Probabilistic labeling
LENGUAJES Y SISTEMAS INFORMATICOS
id ES_ebc7e8d513ab866b9fc7bbc81616475b
oai_identifier_str oai:riunet.upv.es:10251/171117
network_acronym_str ES
network_name_str España
repository_id_str
dc.title.none.fl_str_mv Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
title Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
spellingShingle Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
Fornaciari, Tommaso
Deception detection
Crowdsourcing
Ground truth
Probabilistic labeling
LENGUAJES Y SISTEMAS INFORMATICOS
title_short Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
title_full Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
title_fullStr Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
title_full_unstemmed Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
title_sort Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?
dc.creator.none.fl_str_mv Fornaciari, Tommaso
Cagnina, Leticia
Poesio, Massimo
Rosso, Paolo
author Fornaciari, Tommaso
author_facet Fornaciari, Tommaso
Cagnina, Leticia
Poesio, Massimo
Rosso, Paolo
author_role author
author2 Cagnina, Leticia
Poesio, Massimo
Rosso, Paolo
author2_role author
author
author
dc.contributor.none.fl_str_mv Departamento de Sistemas Informáticos y Computación
Escuela Técnica Superior de Ingeniería Informática
Centro de Investigación Pattern Recognition and Human Language Technology
Agencia Estatal de Investigación
UK Research and Innovation
European Regional Development Fund
Ministerio de Economía y Competitividad
Economic and Social Research Council, Reino Unido
Consejo Nacional de Investigaciones Científicas y Técnicas, Argentina
Repositorio Institucional de la Universitat Politècnica de València Riunet
dc.subject.none.fl_str_mv Deception detection
Crowdsourcing
Ground truth
Probabilistic labeling
LENGUAJES Y SISTEMAS INFORMATICOS
topic Deception detection
Crowdsourcing
Ground truth
Probabilistic labeling
LENGUAJES Y SISTEMAS INFORMATICOS
description [EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and adding them to a set of deceptive reviews obtained through crowdsourcing services. Models trained this way are generally successful at discriminating between `genuine¿ online reviews and the crowdsourced deceptive reviews. It has been argued that the deceptive reviews obtained via crowdsourcing are very different from real fake reviews, but the claim has never been properly tested. In this paper, we compare (false) crowdsourced reviews with a set of `real¿ fake reviews published on line. We evaluate their degree of similarity and their usefulness in training models for the detection of untrustworthy reviews. We find that the deceptive reviews collected via crowdsourcing are significantly different from the fake reviews published online. In the case of the artificially produced deceptive texts, it turns out that their domain similarity with the targets affects the models¿ performance, much more than their untruthfulness. This suggests that the use of crowdsourced datasets for opinion spam detection may not result in models applicable to the real task of detecting deceptive reviews. As an alternative method to create large-size datasets for the fake reviews detection task, we propose methods based on the probabilistic annotation of unlabeled texts, relying on the use of meta-information generally available on the e-commerce sites. Such methods are independent from the content of the reviews and allow to train reliable models for the detection of fake reviews.
publishDate 2020
dc.date.none.fl_str_mv 2020
2020-12-01
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://riunet.upv.es/handle/10251/171117
url https://riunet.upv.es/handle/10251/171117
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.relation.none.fl_str_mv UK Research and Innovation https://doi.org/10.13039/100014013 ES%2FM010236%2F1 Human Rights and Information Technology in the Era of Big Data
Ministerio de Economía y Competitividad http://dx.doi.org/10.13039/501100003329 TIN2015-71147-C2-1-P COMPRENSION DEL LENGUAJE EN LOS MEDIOS DE COMUNICACION SOCIAL - REPRESENTANDO CONTEXTOS DE FORMA CONTINUA
Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 PGC2018-096212-B-C31 DESINFORMACION Y AGRESIVIDAD EN SOCIAL MEDIA: AGREGANDO INFORMACION Y ANALIZANDO EL LENGUAJE
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reserva de todos los derechos
http://rightsstatements.org/vocab/InC/1.0/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reserva de todos los derechos
http://rightsstatements.org/vocab/InC/1.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Springer-Verlag
publisher.none.fl_str_mv Springer-Verlag
dc.source.none.fl_str_mv reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
instname:Universitat Politècnica de València (UPV)
instname_str Universitat Politècnica de València (UPV)
reponame_str RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
collection RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869423257783894016
spelling Fake Opinion Detection: How Similar are Crowdsourced Datasets to Real Data?Fornaciari, TommasoCagnina, LeticiaPoesio, MassimoRosso, PaoloDeception detectionCrowdsourcingGround truthProbabilistic labelingLENGUAJES Y SISTEMAS INFORMATICOS[EN] Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and adding them to a set of deceptive reviews obtained through crowdsourcing services. Models trained this way are generally successful at discriminating between `genuine¿ online reviews and the crowdsourced deceptive reviews. It has been argued that the deceptive reviews obtained via crowdsourcing are very different from real fake reviews, but the claim has never been properly tested. In this paper, we compare (false) crowdsourced reviews with a set of `real¿ fake reviews published on line. We evaluate their degree of similarity and their usefulness in training models for the detection of untrustworthy reviews. We find that the deceptive reviews collected via crowdsourcing are significantly different from the fake reviews published online. In the case of the artificially produced deceptive texts, it turns out that their domain similarity with the targets affects the models¿ performance, much more than their untruthfulness. This suggests that the use of crowdsourced datasets for opinion spam detection may not result in models applicable to the real task of detecting deceptive reviews. As an alternative method to create large-size datasets for the fake reviews detection task, we propose methods based on the probabilistic annotation of unlabeled texts, relying on the use of meta-information generally available on the e-commerce sites. Such methods are independent from the content of the reviews and allow to train reliable models for the detection of fake reviews.Leticia Cagnina thanks CONICET for the continued financial support. This work was funded by MINECO/FEDER (Grant No. SomEMBED TIN2015-71147-C2-1-P). The work of Paolo Rosso was partially funded by the MISMIS-FAKEnHATE Spanish MICINN research project (PGC2018-096212-B-C31). Massimo Poesio was in part supported by the UK Economic and Social Research Council (Grant Number ES/M010236/1).Springer-VerlagDepartamento de Sistemas Informáticos y ComputaciónEscuela Técnica Superior de Ingeniería InformáticaCentro de Investigación Pattern Recognition and Human Language TechnologyAgencia Estatal de InvestigaciónUK Research and InnovationEuropean Regional Development FundMinisterio de Economía y CompetitividadEconomic and Social Research Council, Reino UnidoConsejo Nacional de Investigaciones Científicas y Técnicas, ArgentinaRepositorio Institucional de la Universitat Politècnica de València Riunet20202020-12-01journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfapplication/pdfhttps://riunet.upv.es/handle/10251/171117reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valénciainstname:Universitat Politècnica de València (UPV)InglésengUK Research and Innovation https://doi.org/10.13039/100014013 ES%2FM010236%2F1 Human Rights and Information Technology in the Era of Big DataMinisterio de Economía y Competitividad http://dx.doi.org/10.13039/501100003329 TIN2015-71147-C2-1-P COMPRENSION DEL LENGUAJE EN LOS MEDIOS DE COMUNICACION SOCIAL - REPRESENTANDO CONTEXTOS DE FORMA CONTINUAAgencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 PGC2018-096212-B-C31 DESINFORMACION Y AGRESIVIDAD EN SOCIAL MEDIA: AGREGANDO INFORMACION Y ANALIZANDO EL LENGUAJEopen accesshttp://purl.org/coar/access_right/c_abf2Reserva de todos los derechoshttp://rightsstatements.org/vocab/InC/1.0/info:eu-repo/semantics/openAccessoai:riunet.upv.es:10251/1711172026-06-13T07:49:27Z
score 15.301629