Efficiency of automatic text generators for online review content generation

The evolution of Artificial Intelligence has led to the appearance of automatic text generators able to closely resemble human writing, endangering the development of e-commerce and the consumer confidence. Thus, it is critical to deeply understand how these text generators work to present the prese...

ver descrição completa

Detalhes bibliográficos
Autores: Pérez-Castro, A., Martínez Torres, María del Rocío, Toral, S. L.
Formato: artículo
Estado:Versión publicada
Fecha de publicación:2023
País:España
Recursos:Universidad de Sevilla (US)
Repositorio:idUS. Depósito de Investigación de la Universidad de Sevilla
OAI Identifier:oai:idus.us.es:11441/154160
Acesso em linha:https://hdl.handle.net/11441/154160
https://doi.org/10.1016/j.techfore.2023.122380
Access Level:acceso abierto
Palavra-chave:Deceptive reviews generation
Word-based encoding
Context-based encoding
Pretrained models
Transfer learning
Descrição
Resumo:The evolution of Artificial Intelligence has led to the appearance of automatic text generators able to closely resemble human writing, endangering the development of e-commerce and the consumer confidence. Thus, it is critical to deeply understand how these text generators work to present the presence of deceptive reviews. This paper analyzes one of the most popular text generators, GPT2 (Generative Pre-trained Transformer 2), and studies its effectivity compared to human-generated reviews using previously published classifiers trained to distinguish between real and deceptive reviews. One parameter of the model is the so-called temperature, which determines how deterministic the model is. The temperature adjusts the probability distribution of the words in the model, so that a higher temperature translates into a higher degree of inventiveness in the generation of the texts. Findings reveal (i) that automatically-generated deceptive reviews worsen the accuracy of existing classifiers, this effect being accentuated by the degree of inventiveness; (ii) that their performance depends on the data used to train the generator; and (iii) that the sentiment polarity has no effect on the performance of detection classifiers.