Descubrimiento automático de hipónimos a partir de texto no estructurado

Nowadays, thanks to the Web, we dispose of a huge number of electronic texts. Given the availability and easy access to these texts, it has emerged an interest for manipulating them in an automatic way with the aim to extract prominent information. The extracted information can be used to create or...

Descripción completa

Detalles Bibliográficos
Autor: ROSA MARIA ORTEGA MENDOZA
Tipo de recurso: tesis de maestría
Estado:Versión aceptada para publicación
Fecha de publicación:2007
País:México
Institución:Instituto Nacional de Astrofísica, Óptica y Electrónica
Repositorio:Repositorio Institucional del INAOE
Idioma:español
OAI Identifier:oai:inaoe.repositorioinstitucional.mx:1009/639
Acceso en línea:http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/639
Access Level:acceso abierto
Palabra clave:info:eu-repo/classification/Idiomas naturales/Natural languages
info:eu-repo/classification/Análisis de texto/Text analysis
info:eu-repo/classification/Aplicaciones computacionales/Computer applications
info:eu-repo/classification/cti/1
info:eu-repo/classification/cti/12
info:eu-repo/classification/cti/1203
info:eu-repo/classification/cti/330405
Descripción
Sumario:Nowadays, thanks to the Web, we dispose of a huge number of electronic texts. Given the availability and easy access to these texts, it has emerged an interest for manipulating them in an automatic way with the aim to extract prominent information. The extracted information can be used to create or to enrich lexical resources. In general, this type of resources contains knowledge about the language’s words. Typically, it proposes methods that extract semantic relationships from texts for building automatically these resources. The present investigation work is located inside the automatic construction of lexical resources. In particular, this work is focused on the construction of a hyponyms catalog. Basically, the proposed method is based on the use of patterns to treat the automatic extraction of hyponyms in non-structured texts Traditionally, methods that use patterns to solve the problem involve morphological or syntactic information in the patterns’ definition. In contrast with these methods, we work without this type of information. Therefore, the patterns are defined exclusively at a lexical level. This way, the proposed method achieves language independence and domain independence. In addition, the use of linguistic tools characteristic of a language is avoided (for example: taggers, syntactic analyzers, etc.). However, the extraction of incorrect information is favored. The proposed method confronts this inconvenience by applying two approaches in order to estimate the confidence of the extracted hyponym-hypernym couples. Finally, for showing the utility of the proposed method we evaluated the precision of the obtained catalog. The achieved results are encouraging and they show the feasibility of using lexical patterns to extract automatically hyponyms from non-structured texts.