Iarg-AnCora: Spanish corpus annotated with implicit arguments

This article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually...

Descripción completa

Detalles Bibliográficos
Autores: Taulé Delor, Mariona, Peris Morant, Aina, Rodríguez Hontoria, Horacio
Tipo de recurso: artículo
Estado:Versión aceptada para publicación
Fecha de publicación:2016
País:España
Institución:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
Repositorio:Recercat. Dipósit de la Recerca de Catalunya
OAI Identifier:oai:recercat.cat:2445/171322
Acceso en línea:https://hdl.handle.net/2445/171322
Access Level:acceso abierto
Palabra clave:Corpus (Lingüística)
Semàntica
Castellà (Llengua)
Corpora (Linguistics)
Semantics
Spanish language
id ES_7ef497f16633633ceb93f1df9964e63f
oai_identifier_str oai:recercat.cat:2445/171322
network_acronym_str ES
network_name_str España
repository_id_str
spelling Iarg-AnCora: Spanish corpus annotated with implicit argumentsTaulé Delor, MarionaPeris Morant, AinaRodríguez Hontoria, HoracioCorpus (Lingüística)SemànticaCastellà (Llengua)Corpora (Linguistics)SemanticsSpanish languageThis article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually annotated and an interannotator agreement test was conducted (81 % observed agreement) in order to ensure the reliability of the final resource. The annotation of implicit arguments results in an important gain in argument and thematic role coverage (128 % on average). It is the first corpus annotated with implicit arguments for the Spanish language with a wide coverage that is freely available. This corpus can subsequently be used by machine learning-based semantic role labeling systems, and for the linguistic analysis of implicit arguments grounded on real data. Semantic analyzers are essential components of current language technology applications, which need to obtain a deeper understanding of the text in order to make inferences at the highest level to obtain qualitative improvements in the results.Springer Verlag2020202020162020info:eu-repo/semantics/articleinfo:eu-repo/semantics/acceptedVersion28 p.application/pdfhttps://hdl.handle.net/2445/171322Articles publicats en revistes (Filologia Catalana i Lingüística General)reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésVersió postprint del document publicat a: https://doi.org/10.1007/s10579-015-9334-3Language Resources And Evaluation, 2016, vol. 50, num. 3, p. 549-584https://doi.org/10.1007/s10579-015-9334-3(c) Springer Verlag, 2016info:eu-repo/semantics/openAccessoai:recercat.cat:2445/1713222026-05-29T05:05:01Z
dc.title.none.fl_str_mv Iarg-AnCora: Spanish corpus annotated with implicit arguments
title Iarg-AnCora: Spanish corpus annotated with implicit arguments
spellingShingle Iarg-AnCora: Spanish corpus annotated with implicit arguments
Taulé Delor, Mariona
Corpus (Lingüística)
Semàntica
Castellà (Llengua)
Corpora (Linguistics)
Semantics
Spanish language
title_short Iarg-AnCora: Spanish corpus annotated with implicit arguments
title_full Iarg-AnCora: Spanish corpus annotated with implicit arguments
title_fullStr Iarg-AnCora: Spanish corpus annotated with implicit arguments
title_full_unstemmed Iarg-AnCora: Spanish corpus annotated with implicit arguments
title_sort Iarg-AnCora: Spanish corpus annotated with implicit arguments
dc.creator.none.fl_str_mv Taulé Delor, Mariona
Peris Morant, Aina
Rodríguez Hontoria, Horacio
author Taulé Delor, Mariona
author_facet Taulé Delor, Mariona
Peris Morant, Aina
Rodríguez Hontoria, Horacio
author_role author
author2 Peris Morant, Aina
Rodríguez Hontoria, Horacio
author2_role author
author
dc.subject.none.fl_str_mv Corpus (Lingüística)
Semàntica
Castellà (Llengua)
Corpora (Linguistics)
Semantics
Spanish language
topic Corpus (Lingüística)
Semàntica
Castellà (Llengua)
Corpora (Linguistics)
Semantics
Spanish language
description This article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually annotated and an interannotator agreement test was conducted (81 % observed agreement) in order to ensure the reliability of the final resource. The annotation of implicit arguments results in an important gain in argument and thematic role coverage (128 % on average). It is the first corpus annotated with implicit arguments for the Spanish language with a wide coverage that is freely available. This corpus can subsequently be used by machine learning-based semantic role labeling systems, and for the linguistic analysis of implicit arguments grounded on real data. Semantic analyzers are essential components of current language technology applications, which need to obtain a deeper understanding of the text in order to make inferences at the highest level to obtain qualitative improvements in the results.
publishDate 2016
dc.date.none.fl_str_mv 2016
2020
2020
2020
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/acceptedVersion
format article
status_str acceptedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/2445/171322
url https://hdl.handle.net/2445/171322
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.relation.none.fl_str_mv Versió postprint del document publicat a: https://doi.org/10.1007/s10579-015-9334-3
Language Resources And Evaluation, 2016, vol. 50, num. 3, p. 549-584
https://doi.org/10.1007/s10579-015-9334-3
dc.rights.none.fl_str_mv (c) Springer Verlag, 2016
info:eu-repo/semantics/openAccess
rights_invalid_str_mv (c) Springer Verlag, 2016
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv 28 p.
application/pdf
dc.publisher.none.fl_str_mv Springer Verlag
publisher.none.fl_str_mv Springer Verlag
dc.source.none.fl_str_mv Articles publicats en revistes (Filologia Catalana i Lingüística General)
reponame:Recercat. Dipósit de la Recerca de Catalunya
instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
instname_str Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)
reponame_str Recercat. Dipósit de la Recerca de Catalunya
collection Recercat. Dipósit de la Recerca de Catalunya
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869411788169150464
score 15,812429