Survey of Word Co-occurrence Measures for Collocation Detection

This paper presents a detailed survey of word co-occurrence measures used in natural language processing. Word co-occurrence information is vital for accurate computational text treatment, it is important to distinguish words which can combine freely with other words from other words whose preferenc...

Descripción completa

Detalles Bibliográficos
Autor: Olga Kolesnikova
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2016
País:México
Institución:Instituto Politécnico Nacional
Repositorio:Redalyc-IPN
OAI Identifier:oai:redalyc.org:61547469003
Acceso en línea:https://www.redalyc.org/articulo.oa?id=61547469003
Access Level:acceso abierto
Palabra clave:Computación
Word co
occurrence
collocation
occurrence measure
association measure
Descripción
Sumario:This paper presents a detailed survey of word co-occurrence measures used in natural language processing. Word co-occurrence information is vital for accurate computational text treatment, it is important to distinguish words which can combine freely with other words from other words whose preferences to generate phrases are restricted. The latter words together with their typical co-occurring companions are called collocations. To detect collocations, many word cooccurrence measures, also called association measures, are used to determine a high degree of cohesion between words in collocations as opposed to a low degree of cohesion in free word combinations. We describe such association measures grouping them in classes depending on approaches and mathematical models used to formalize word co-occurrence.