Iarg-AnCora: Spanish corpus annotated with implicit arguments
This article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually...
| Autores: | , , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión aceptada para publicación |
| Fecha de publicación: | 2016 |
| País: | España |
| Institución: | Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| Repositorio: | Recercat. Dipósit de la Recerca de Catalunya |
| OAI Identifier: | oai:recercat.cat:2445/171322 |
| Acceso en línea: | https://hdl.handle.net/2445/171322 |
| Access Level: | acceso abierto |
| Palabra clave: | Corpus (Lingüística) Semàntica Castellà (Llengua) Corpora (Linguistics) Semantics Spanish language |
| id |
ES_7ef497f16633633ceb93f1df9964e63f |
|---|---|
| oai_identifier_str |
oai:recercat.cat:2445/171322 |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Iarg-AnCora: Spanish corpus annotated with implicit argumentsTaulé Delor, MarionaPeris Morant, AinaRodríguez Hontoria, HoracioCorpus (Lingüística)SemànticaCastellà (Llengua)Corpora (Linguistics)SemanticsSpanish languageThis article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually annotated and an interannotator agreement test was conducted (81 % observed agreement) in order to ensure the reliability of the final resource. The annotation of implicit arguments results in an important gain in argument and thematic role coverage (128 % on average). It is the first corpus annotated with implicit arguments for the Spanish language with a wide coverage that is freely available. This corpus can subsequently be used by machine learning-based semantic role labeling systems, and for the linguistic analysis of implicit arguments grounded on real data. Semantic analyzers are essential components of current language technology applications, which need to obtain a deeper understanding of the text in order to make inferences at the highest level to obtain qualitative improvements in the results.Springer Verlag2020202020162020info:eu-repo/semantics/articleinfo:eu-repo/semantics/acceptedVersion28 p.application/pdfhttps://hdl.handle.net/2445/171322Articles publicats en revistes (Filologia Catalana i Lingüística General)reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)InglésVersió postprint del document publicat a: https://doi.org/10.1007/s10579-015-9334-3Language Resources And Evaluation, 2016, vol. 50, num. 3, p. 549-584https://doi.org/10.1007/s10579-015-9334-3(c) Springer Verlag, 2016info:eu-repo/semantics/openAccessoai:recercat.cat:2445/1713222026-05-29T05:05:01Z |
| dc.title.none.fl_str_mv |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| title |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| spellingShingle |
Iarg-AnCora: Spanish corpus annotated with implicit arguments Taulé Delor, Mariona Corpus (Lingüística) Semàntica Castellà (Llengua) Corpora (Linguistics) Semantics Spanish language |
| title_short |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| title_full |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| title_fullStr |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| title_full_unstemmed |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| title_sort |
Iarg-AnCora: Spanish corpus annotated with implicit arguments |
| dc.creator.none.fl_str_mv |
Taulé Delor, Mariona Peris Morant, Aina Rodríguez Hontoria, Horacio |
| author |
Taulé Delor, Mariona |
| author_facet |
Taulé Delor, Mariona Peris Morant, Aina Rodríguez Hontoria, Horacio |
| author_role |
author |
| author2 |
Peris Morant, Aina Rodríguez Hontoria, Horacio |
| author2_role |
author author |
| dc.subject.none.fl_str_mv |
Corpus (Lingüística) Semàntica Castellà (Llengua) Corpora (Linguistics) Semantics Spanish language |
| topic |
Corpus (Lingüística) Semàntica Castellà (Llengua) Corpora (Linguistics) Semantics Spanish language |
| description |
This article presents the Spanish Iarg-AnCora corpus (400 k-words, 13,883 sentences) annotated with the implicit arguments of deverbal nominalizations (18,397 occurrences). We describe the methodology used to create it, focusing on the annotation scheme and criteria adopted. The corpus was manually annotated and an interannotator agreement test was conducted (81 % observed agreement) in order to ensure the reliability of the final resource. The annotation of implicit arguments results in an important gain in argument and thematic role coverage (128 % on average). It is the first corpus annotated with implicit arguments for the Spanish language with a wide coverage that is freely available. This corpus can subsequently be used by machine learning-based semantic role labeling systems, and for the linguistic analysis of implicit arguments grounded on real data. Semantic analyzers are essential components of current language technology applications, which need to obtain a deeper understanding of the text in order to make inferences at the highest level to obtain qualitative improvements in the results. |
| publishDate |
2016 |
| dc.date.none.fl_str_mv |
2016 2020 2020 2020 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/acceptedVersion |
| format |
article |
| status_str |
acceptedVersion |
| dc.identifier.none.fl_str_mv |
https://hdl.handle.net/2445/171322 |
| url |
https://hdl.handle.net/2445/171322 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
Versió postprint del document publicat a: https://doi.org/10.1007/s10579-015-9334-3 Language Resources And Evaluation, 2016, vol. 50, num. 3, p. 549-584 https://doi.org/10.1007/s10579-015-9334-3 |
| dc.rights.none.fl_str_mv |
(c) Springer Verlag, 2016 info:eu-repo/semantics/openAccess |
| rights_invalid_str_mv |
(c) Springer Verlag, 2016 |
| eu_rights_str_mv |
openAccess |
| dc.format.none.fl_str_mv |
28 p. application/pdf |
| dc.publisher.none.fl_str_mv |
Springer Verlag |
| publisher.none.fl_str_mv |
Springer Verlag |
| dc.source.none.fl_str_mv |
Articles publicats en revistes (Filologia Catalana i Lingüística General) reponame:Recercat. Dipósit de la Recerca de Catalunya instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| instname_str |
Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| reponame_str |
Recercat. Dipósit de la Recerca de Catalunya |
| collection |
Recercat. Dipósit de la Recerca de Catalunya |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869411788169150464 |
| score |
15,812429 |