Clasificación automática de textos considerando el estilo de redacción

Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more p...

Descripción completa

Detalles Bibliográficos
Autor: ROSA MARIA COYOTL MORALES
Tipo de recurso: tesis de maestría
Estado:Versión aceptada para publicación
Fecha de publicación:2007
País:México
Institución:Instituto Nacional de Astrofísica, Óptica y Electrónica
Repositorio:Repositorio Institucional del INAOE
Idioma:español
OAI Identifier:oai:inaoe.repositorioinstitucional.mx:1009/587
Acceso en línea:http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/587
Access Level:acceso abierto
Palabra clave:info:eu-repo/classification/Aprendizaje automático/Machine learning
info:eu-repo/classification/Clasificación/Classification
info:eu-repo/classification/Análisis de la información/Information analysis
info:eu-repo/classification/cti/1
info:eu-repo/classification/cti/12
info:eu-repo/classification/cti/1203
info:eu-repo/classification/cti/330405
Descripción
Sumario:Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more predefined categories. Most work in this field focuses on categorizing documents by their topic. However, a document can be also classified by its written style (non-topic classification). Basically, nontopic classification considers tasks such as sentiment classification, plagiarism detection, authorship attribution, genre classification, etc. The main objective of this thesis is to propose methods for determining the lexical features that allow characterizing the written style of documents. The proposed methods consider the characterization of documents by sets of word sequences that combine content and functional words. The usefulness of this kind of characterization is demonstrated by its application in the tasks of authorship attribution and genre classification.