Determining and Characterizing the Reused Text for Plagiarism Detection

An important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficu...

ver descrição completa

Detalhes bibliográficos
Autores: Sánchez-Vega, Fernando, Villatoro-Tello, Esaú, Montes-y-Gómez, Manuel, Villaseñor-Pineda; Luis, Rosso, Paolo
Formato: artículo
Fecha de publicación:2013
País:España
Recursos:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/38255
Acesso em linha:https://riunet.upv.es/handle/10251/38255
Access Level:acceso abierto
Palavra-chave:Plagiarism detection
Text reuse
Machine learning
Supervised classification
LENGUAJES Y SISTEMAS INFORMATICOS
id ES_fe39b30f30a3a4806d4026feebe5bd8d
oai_identifier_str oai:riunet.upv.es:10251/38255
network_acronym_str ES
network_name_str España
repository_id_str
spelling Determining and Characterizing the Reused Text for Plagiarism DetectionSánchez-Vega, FernandoVillatoro-Tello, EsaúMontes-y-Gómez, ManuelVillaseñor-Pineda; LuisRosso, PaoloPlagiarism detectionText reuseMachine learningSupervised classificationLENGUAJES Y SISTEMAS INFORMATICOSAn important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficulty is that not all similar text fragments are examples of plagiarism, since thematic coincidences also tend to produce portions of similar text. In order to tackle these problems, we propose a novel method for detecting likely portions of reused text. This method is able to detect common actions performed by plagiarists such as word deletion, insertion and transposition, allowing to obtain plausible portions of reused text. We also propose representing the identified reused text by means of a set of features that denote its degree of plagiarism, relevance and fragmentation. This new representation aims to facilitate the recognition of plagiarism by considering diverse characteristics of the reused text during the classification phase. Experimental results employing a supervised classification strategy showed that the proposed method is able to outperform traditionally used approaches. 2012 Elsevier Ltd. All rights reserved.This work was done under partial support of CONACyT project Grants: 134186, and Scholarships: 258345/224483. This work is the result of the collaboration in the framework of the WIQEI IRSES project (Grant No. 269180) within the FP 7 Marie Curie. The work of the last author was in the framework of the VLC/CAMPUS Microcluster on Multimodal Interaction in Intelligent Systems.ElsevierDepartamento de Sistemas Informáticos y ComputaciónEscuela Técnica Superior de Ingeniería InformáticaCentro de Investigación Pattern Recognition and Human Language TechnologyConsejo Nacional de Ciencia y Tecnología, MéxicoEuropean CommissionRepositorio Institucional de la Universitat Politècnica de València Riunet20132013-04-01journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfapplication/pdfhttps://riunet.upv.es/handle/10251/38255reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valénciainstname:Universitat Politècnica de València (UPV)InglésengConsejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 134186Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 258345%2F224483European Commission https://doi.org/10.13039/501100000780 FP7 269180 Web Information Quality Evaluation Initiativeopen accesshttp://purl.org/coar/access_right/c_abf2Reserva de todos los derechoshttp://rightsstatements.org/vocab/InC/1.0/info:eu-repo/semantics/openAccessoai:riunet.upv.es:10251/382552026-06-13T07:49:27Z
dc.title.none.fl_str_mv Determining and Characterizing the Reused Text for Plagiarism Detection
title Determining and Characterizing the Reused Text for Plagiarism Detection
spellingShingle Determining and Characterizing the Reused Text for Plagiarism Detection
Sánchez-Vega, Fernando
Plagiarism detection
Text reuse
Machine learning
Supervised classification
LENGUAJES Y SISTEMAS INFORMATICOS
title_short Determining and Characterizing the Reused Text for Plagiarism Detection
title_full Determining and Characterizing the Reused Text for Plagiarism Detection
title_fullStr Determining and Characterizing the Reused Text for Plagiarism Detection
title_full_unstemmed Determining and Characterizing the Reused Text for Plagiarism Detection
title_sort Determining and Characterizing the Reused Text for Plagiarism Detection
dc.creator.none.fl_str_mv Sánchez-Vega, Fernando
Villatoro-Tello, Esaú
Montes-y-Gómez, Manuel
Villaseñor-Pineda; Luis
Rosso, Paolo
author Sánchez-Vega, Fernando
author_facet Sánchez-Vega, Fernando
Villatoro-Tello, Esaú
Montes-y-Gómez, Manuel
Villaseñor-Pineda; Luis
Rosso, Paolo
author_role author
author2 Villatoro-Tello, Esaú
Montes-y-Gómez, Manuel
Villaseñor-Pineda; Luis
Rosso, Paolo
author2_role author
author
author
author
dc.contributor.none.fl_str_mv Departamento de Sistemas Informáticos y Computación
Escuela Técnica Superior de Ingeniería Informática
Centro de Investigación Pattern Recognition and Human Language Technology
Consejo Nacional de Ciencia y Tecnología, México
European Commission
Repositorio Institucional de la Universitat Politècnica de València Riunet
dc.subject.none.fl_str_mv Plagiarism detection
Text reuse
Machine learning
Supervised classification
LENGUAJES Y SISTEMAS INFORMATICOS
topic Plagiarism detection
Text reuse
Machine learning
Supervised classification
LENGUAJES Y SISTEMAS INFORMATICOS
description An important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficulty is that not all similar text fragments are examples of plagiarism, since thematic coincidences also tend to produce portions of similar text. In order to tackle these problems, we propose a novel method for detecting likely portions of reused text. This method is able to detect common actions performed by plagiarists such as word deletion, insertion and transposition, allowing to obtain plausible portions of reused text. We also propose representing the identified reused text by means of a set of features that denote its degree of plagiarism, relevance and fragmentation. This new representation aims to facilitate the recognition of plagiarism by considering diverse characteristics of the reused text during the classification phase. Experimental results employing a supervised classification strategy showed that the proposed method is able to outperform traditionally used approaches. 2012 Elsevier Ltd. All rights reserved.
publishDate 2013
dc.date.none.fl_str_mv 2013
2013-04-01
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://riunet.upv.es/handle/10251/38255
url https://riunet.upv.es/handle/10251/38255
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.relation.none.fl_str_mv Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 134186
Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 258345%2F224483
European Commission https://doi.org/10.13039/501100000780 FP7 269180 Web Information Quality Evaluation Initiative
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reserva de todos los derechos
http://rightsstatements.org/vocab/InC/1.0/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reserva de todos los derechos
http://rightsstatements.org/vocab/InC/1.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/pdf
dc.publisher.none.fl_str_mv Elsevier
publisher.none.fl_str_mv Elsevier
dc.source.none.fl_str_mv reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
instname:Universitat Politècnica de València (UPV)
instname_str Universitat Politècnica de València (UPV)
reponame_str RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
collection RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869425665103626240
score 15.301603