Determining and Characterizing the Reused Text for Plagiarism Detection
An important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficu...
| Autores: | , , , , |
|---|---|
| Formato: | artículo |
| Fecha de publicación: | 2013 |
| País: | España |
| Recursos: | Universitat Politècnica de València (UPV) |
| Repositorio: | RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| Idioma: | inglés |
| OAI Identifier: | oai:riunet.upv.es:10251/38255 |
| Acesso em linha: | https://riunet.upv.es/handle/10251/38255 |
| Access Level: | acceso abierto |
| Palavra-chave: | Plagiarism detection Text reuse Machine learning Supervised classification LENGUAJES Y SISTEMAS INFORMATICOS |
| id |
ES_fe39b30f30a3a4806d4026feebe5bd8d |
|---|---|
| oai_identifier_str |
oai:riunet.upv.es:10251/38255 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Determining and Characterizing the Reused Text for Plagiarism DetectionSánchez-Vega, FernandoVillatoro-Tello, EsaúMontes-y-Gómez, ManuelVillaseñor-Pineda; LuisRosso, PaoloPlagiarism detectionText reuseMachine learningSupervised classificationLENGUAJES Y SISTEMAS INFORMATICOSAn important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficulty is that not all similar text fragments are examples of plagiarism, since thematic coincidences also tend to produce portions of similar text. In order to tackle these problems, we propose a novel method for detecting likely portions of reused text. This method is able to detect common actions performed by plagiarists such as word deletion, insertion and transposition, allowing to obtain plausible portions of reused text. We also propose representing the identified reused text by means of a set of features that denote its degree of plagiarism, relevance and fragmentation. This new representation aims to facilitate the recognition of plagiarism by considering diverse characteristics of the reused text during the classification phase. Experimental results employing a supervised classification strategy showed that the proposed method is able to outperform traditionally used approaches. 2012 Elsevier Ltd. All rights reserved.This work was done under partial support of CONACyT project Grants: 134186, and Scholarships: 258345/224483. This work is the result of the collaboration in the framework of the WIQEI IRSES project (Grant No. 269180) within the FP 7 Marie Curie. The work of the last author was in the framework of the VLC/CAMPUS Microcluster on Multimodal Interaction in Intelligent Systems.ElsevierDepartamento de Sistemas Informáticos y ComputaciónEscuela Técnica Superior de Ingeniería InformáticaCentro de Investigación Pattern Recognition and Human Language TechnologyConsejo Nacional de Ciencia y Tecnología, MéxicoEuropean CommissionRepositorio Institucional de la Universitat Politècnica de València Riunet20132013-04-01journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfapplication/pdfhttps://riunet.upv.es/handle/10251/38255reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valénciainstname:Universitat Politècnica de València (UPV)InglésengConsejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 134186Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 258345%2F224483European Commission https://doi.org/10.13039/501100000780 FP7 269180 Web Information Quality Evaluation Initiativeopen accesshttp://purl.org/coar/access_right/c_abf2Reserva de todos los derechoshttp://rightsstatements.org/vocab/InC/1.0/info:eu-repo/semantics/openAccessoai:riunet.upv.es:10251/382552026-06-13T07:49:27Z |
| dc.title.none.fl_str_mv |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| title |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| spellingShingle |
Determining and Characterizing the Reused Text for Plagiarism Detection Sánchez-Vega, Fernando Plagiarism detection Text reuse Machine learning Supervised classification LENGUAJES Y SISTEMAS INFORMATICOS |
| title_short |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| title_full |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| title_fullStr |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| title_full_unstemmed |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| title_sort |
Determining and Characterizing the Reused Text for Plagiarism Detection |
| dc.creator.none.fl_str_mv |
Sánchez-Vega, Fernando Villatoro-Tello, Esaú Montes-y-Gómez, Manuel Villaseñor-Pineda; Luis Rosso, Paolo |
| author |
Sánchez-Vega, Fernando |
| author_facet |
Sánchez-Vega, Fernando Villatoro-Tello, Esaú Montes-y-Gómez, Manuel Villaseñor-Pineda; Luis Rosso, Paolo |
| author_role |
author |
| author2 |
Villatoro-Tello, Esaú Montes-y-Gómez, Manuel Villaseñor-Pineda; Luis Rosso, Paolo |
| author2_role |
author author author author |
| dc.contributor.none.fl_str_mv |
Departamento de Sistemas Informáticos y Computación Escuela Técnica Superior de Ingeniería Informática Centro de Investigación Pattern Recognition and Human Language Technology Consejo Nacional de Ciencia y Tecnología, México European Commission Repositorio Institucional de la Universitat Politècnica de València Riunet |
| dc.subject.none.fl_str_mv |
Plagiarism detection Text reuse Machine learning Supervised classification LENGUAJES Y SISTEMAS INFORMATICOS |
| topic |
Plagiarism detection Text reuse Machine learning Supervised classification LENGUAJES Y SISTEMAS INFORMATICOS |
| description |
An important task in plagiarism detection is determining and measuring similar text portions between a given pair of documents. One of the main difficulties of this task resides on the fact that reused text is commonly modified with the aim of covering or camouflaging the plagiarism. Another difficulty is that not all similar text fragments are examples of plagiarism, since thematic coincidences also tend to produce portions of similar text. In order to tackle these problems, we propose a novel method for detecting likely portions of reused text. This method is able to detect common actions performed by plagiarists such as word deletion, insertion and transposition, allowing to obtain plausible portions of reused text. We also propose representing the identified reused text by means of a set of features that denote its degree of plagiarism, relevance and fragmentation. This new representation aims to facilitate the recognition of plagiarism by considering diverse characteristics of the reused text during the classification phase. Experimental results employing a supervised classification strategy showed that the proposed method is able to outperform traditionally used approaches. 2012 Elsevier Ltd. All rights reserved. |
| publishDate |
2013 |
| dc.date.none.fl_str_mv |
2013 2013-04-01 |
| dc.type.none.fl_str_mv |
journal article http://purl.org/coar/resource_type/c_6501 VoR http://purl.org/coar/version/c_970fb48d4fbd8a85 |
| dc.type.openaire.fl_str_mv |
info:eu-repo/semantics/article |
| format |
article |
| dc.identifier.none.fl_str_mv |
https://riunet.upv.es/handle/10251/38255 |
| url |
https://riunet.upv.es/handle/10251/38255 |
| dc.language.none.fl_str_mv |
Inglés eng |
| language_invalid_str_mv |
Inglés |
| language |
eng |
| dc.relation.none.fl_str_mv |
Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 134186 Consejo Nacional de Ciencia y Tecnología, México https://doi.org/10.13039/501100003141 258345%2F224483 European Commission https://doi.org/10.13039/501100000780 FP7 269180 Web Information Quality Evaluation Initiative |
| dc.rights.none.fl_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Reserva de todos los derechos http://rightsstatements.org/vocab/InC/1.0/ |
| dc.rights.openaire.fl_str_mv |
info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
open access http://purl.org/coar/access_right/c_abf2 Reserva de todos los derechos http://rightsstatements.org/vocab/InC/1.0/ |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
application/pdf application/pdf |
| dc.publisher.none.fl_str_mv |
Elsevier |
| publisher.none.fl_str_mv |
Elsevier |
| dc.source.none.fl_str_mv |
reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia instname:Universitat Politècnica de València (UPV) |
| instname_str |
Universitat Politècnica de València (UPV) |
| reponame_str |
RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| collection |
RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869425665103626240 |
| score |
15.301603 |