Language variety identification using distributed representations of words and documents

In this work we focus on the use of distributed representations of words and documents using the continuous Skip-gram model. We compare this model with three recent approaches: Information Gain Word-Patterns, TF-IDF graphs and Emotion-labeled Graphs, in addition to several baselines. We evaluate the...

Descripción completa

Detalles Bibliográficos
Autores: Franco Salvador, Marc, Rangel, Francisco, Taulé, Mariona, Martí, M. Antònia, Rosso, Paolo
Tipo de recurso: capítulo de libro
Fecha de publicación:2015
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/64372
Acceso en línea:https://riunet.upv.es/handle/10251/64372
Access Level:acceso abierto
Palabra clave:Author profiling
Language variety identification
Distributed representations
Information Gain Word-Patterns
TF-IDF graphs
Emotion-labeled Graphs
LENGUAJES Y SISTEMAS INFORMATICOS
Descripción
Sumario:In this work we focus on the use of distributed representations of words and documents using the continuous Skip-gram model. We compare this model with three recent approaches: Information Gain Word-Patterns, TF-IDF graphs and Emotion-labeled Graphs, in addition to several baselines. We evaluate the models introducing the Hispablogs dataset, a new collection of Spanish blogs from five different countries: Argentina, Chile, Mexico, Peru and Spain. Experimental results show state-of-the-art performance in language variety identification.