Clasificación automática de textos considerando el estilo de redacción

Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more p...

ver descrição completa

Detalhes bibliográficos
Autor: ROSA MARIA COYOTL MORALES
Tipo de documento: dissertação
Estado:Versión aceptada para publicación
Data de publicação:2007
País:México
Recursos:Instituto Nacional de Astrofísica, Óptica y Electrónica
Repositório:Repositorio Institucional del INAOE
Idioma:espanhol
OAI Identifier:oai:inaoe.repositorioinstitucional.mx:1009/587
Acesso em linha:http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/587
Access Level:Acceso aberto
Palavra-chave:info:eu-repo/classification/Aprendizaje automático/Machine learning
info:eu-repo/classification/Clasificación/Classification
info:eu-repo/classification/Análisis de la información/Information analysis
info:eu-repo/classification/cti/1
info:eu-repo/classification/cti/12
info:eu-repo/classification/cti/1203
info:eu-repo/classification/cti/330405
Descrição
Resumo:Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more predefined categories. Most work in this field focuses on categorizing documents by their topic. However, a document can be also classified by its written style (non-topic classification). Basically, nontopic classification considers tasks such as sentiment classification, plagiarism detection, authorship attribution, genre classification, etc. The main objective of this thesis is to propose methods for determining the lexical features that allow characterizing the written style of documents. The proposed methods consider the characterization of documents by sets of word sequences that combine content and functional words. The usefulness of this kind of characterization is demonstrated by its application in the tasks of authorship attribution and genre classification.