Extracting compound terms from domain corpora

The need for domain ontologies motivates the research on structured information extraction from texts. A foundational part of this process is the identification of domain relevant compound terms. This paper presents an evaluation of compound terms extraction from a corpus of the domain of Pediatrics...

ver descrição completa

Detalhes bibliográficos
Autores: Lopes, Lucelene, Vieira, Renata, Finatto, Maria José Bocorny, Martins, Daniel
Tipo de documento: artigo
Estado:Versão publicada
Data de publicação:2010
País:Brasil
Recursos:Universidade Federal do Rio Grande do Sul (UFRGS)
Repositório:Repositório Institucional da UFRGS
Idioma:inglês
OAI Identifier:oai:www.lume.ufrgs.br:10183/174302
Acesso em linha:http://hdl.handle.net/10183/174302
Access Level:Acceso aberto
Palavra-chave:Ontologia
Terminologia
Term extraction
Statistical and linguistic methods
Ontology automatic construction
Extraction from corpora
Descrição
Resumo:The need for domain ontologies motivates the research on structured information extraction from texts. A foundational part of this process is the identification of domain relevant compound terms. This paper presents an evaluation of compound terms extraction from a corpus of the domain of Pediatrics. Bigrams and trigrams were automatically extracted from a corpus composed by 283 texts from a Portuguese journal, Jornal de Pediatria, using three different extraction methods. Considering that these methods generate an elevated number of candidates, we analyzed the quality of the resulting terms according to different methods and cut-off points. The evaluation is reported by metrics such as precision, recall and f-measure, which are computed on the basis of a hand-made reference list of domain relevant compounds.