Clasificación automática de textos considerando el estilo de redacción
Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more p...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis de maestría |
| Estado: | Versión aceptada para publicación |
| Fecha de publicación: | 2007 |
| País: | México |
| Institución: | Instituto Nacional de Astrofísica, Óptica y Electrónica |
| Repositorio: | Repositorio Institucional del INAOE |
| Idioma: | español |
| OAI Identifier: | oai:inaoe.repositorioinstitucional.mx:1009/587 |
| Acceso en línea: | http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/587 |
| Access Level: | acceso abierto |
| Palabra clave: | info:eu-repo/classification/Aprendizaje automático/Machine learning info:eu-repo/classification/Clasificación/Classification info:eu-repo/classification/Análisis de la información/Information analysis info:eu-repo/classification/cti/1 info:eu-repo/classification/cti/12 info:eu-repo/classification/cti/1203 info:eu-repo/classification/cti/330405 |
| Sumario: | Nowadays there is a large amount of information available in digital format. All this information is useless if we do not have adequate mechanisms for its access, classification and analysis. In particular, text classification concerns the automatic assignment of free text documents to one or more predefined categories. Most work in this field focuses on categorizing documents by their topic. However, a document can be also classified by its written style (non-topic classification). Basically, nontopic classification considers tasks such as sentiment classification, plagiarism detection, authorship attribution, genre classification, etc. The main objective of this thesis is to propose methods for determining the lexical features that allow characterizing the written style of documents. The proposed methods consider the characterization of documents by sets of word sequences that combine content and functional words. The usefulness of this kind of characterization is demonstrated by its application in the tasks of authorship attribution and genre classification. |
|---|