Testing the Reasoning for Question Answering Validation

Question answering (QA) is a task that deserves more collaboration between natural language processing (NLP) and knowledge representation (KR) communities, not only to introduce reasoning when looking for answers or making use of answer type taxonomies and encyclopaedic knowledge, but also, as discu...

Full description

Bibliographic Details
Authors: Peñas Padilla, Anselmo, Rodrigo Yuste, Álvaro, Sama, Valentín, Verdejo Maíllo, María Felisa
Format: article
Publication Date:2007
Country:España
Institution:Universidad Nacional de Educación a Distancia
Repository:e-spacio. Repositorio Institucional de la UNED
Language:English
OAI Identifier:oai:dnet:espacio_____::b4f3491e5ea4a6169d63f2b99099ea23
Online Access:https://hdl.handle.net/20.500.14468/32544
Access Level:Open access
Keyword:33 Ciencias Tecnológicas
Textual Entailment
Test Collections
Question Answering
Answer Validation
Evaluation
id ES_c063cc72b10a336ca5c4e5b27103cb9b
oai_identifier_str oai:dnet:espacio_____::b4f3491e5ea4a6169d63f2b99099ea23
network_acronym_str ES
network_name_str España
repository_id_str
spelling Testing the Reasoning for Question Answering ValidationPeñas Padilla, AnselmoRodrigo Yuste, ÁlvaroSama, ValentínVerdejo Maíllo, María Felisa33 Ciencias TecnológicasTextual EntailmentTest CollectionsQuestion AnsweringAnswer ValidationEvaluationQuestion answering (QA) is a task that deserves more collaboration between natural language processing (NLP) and knowledge representation (KR) communities, not only to introduce reasoning when looking for answers or making use of answer type taxonomies and encyclopaedic knowledge, but also, as discussed here, for answer validation (AV), that is to say, to decide whether the responses of a QA system are correct or not. This was one of the motivations for the first Answer Validation Exercise at CLEF 2006 (AVE 2006). The starting point for the AVE 2006 was the reformulation of the answer validation as a recognizing textual entailment (RTE) problem, under the assumption that a hypothesis can be automatically generated instantiating a hypothesis pattern with a QA system answer. The test collections that we developed in seven different languages at AVE 2006 are specially oriented to the development and evaluation of answer validation systems. We show in this article the methodology followed for developing these collections taking advantage of the human assessments already made in the evaluation of QA systems. We also propose an evaluation framework for AV linked to a QA evaluation track. We quantify and discuss the source of errors introduced by the reformulation of the answer validation problem in terms of textual entailment (around 2%, in the range of inter-annotator disagreement). We also show the evaluation results of the first answer validation exercise at CLEF 2006 where 11 groups have participated with 38 runs in seven different languages. The most extensively used techniques were Machine Learning and overlapping measures, but systems with broader knowledge resources and richer representation formalisms obtained the best results.Oxford University PressMinisterio de Ciencia y TecnologíaComunidad de Madride-Spacio UNED20262026-05-0820072007-12-1320072007-12-13journal articlehttp://purl.org/coar/resource_type/c_6501info:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/20.500.14468/32544reponame:e-spacio. Repositorio Institucional de la UNEDinstname:Universidad Nacional de Educación a DistanciaInglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by-nc-nd/4.0/deed.esoai:dnet:espacio_____::b4f3491e5ea4a6169d63f2b99099ea232026-06-06T12:38:31Z
dc.title.none.fl_str_mv Testing the Reasoning for Question Answering Validation
title Testing the Reasoning for Question Answering Validation
spellingShingle Testing the Reasoning for Question Answering Validation
Peñas Padilla, Anselmo
33 Ciencias Tecnológicas
Textual Entailment
Test Collections
Question Answering
Answer Validation
Evaluation
title_short Testing the Reasoning for Question Answering Validation
title_full Testing the Reasoning for Question Answering Validation
title_fullStr Testing the Reasoning for Question Answering Validation
title_full_unstemmed Testing the Reasoning for Question Answering Validation
title_sort Testing the Reasoning for Question Answering Validation
dc.creator.none.fl_str_mv Peñas Padilla, Anselmo
Rodrigo Yuste, Álvaro
Sama, Valentín
Verdejo Maíllo, María Felisa
author Peñas Padilla, Anselmo
author_facet Peñas Padilla, Anselmo
Rodrigo Yuste, Álvaro
Sama, Valentín
Verdejo Maíllo, María Felisa
author_role author
author2 Rodrigo Yuste, Álvaro
Sama, Valentín
Verdejo Maíllo, María Felisa
author2_role author
author
author
dc.contributor.none.fl_str_mv Ministerio de Ciencia y Tecnología
Comunidad de Madrid
e-Spacio UNED
dc.subject.none.fl_str_mv 33 Ciencias Tecnológicas
Textual Entailment
Test Collections
Question Answering
Answer Validation
Evaluation
topic 33 Ciencias Tecnológicas
Textual Entailment
Test Collections
Question Answering
Answer Validation
Evaluation
description Question answering (QA) is a task that deserves more collaboration between natural language processing (NLP) and knowledge representation (KR) communities, not only to introduce reasoning when looking for answers or making use of answer type taxonomies and encyclopaedic knowledge, but also, as discussed here, for answer validation (AV), that is to say, to decide whether the responses of a QA system are correct or not. This was one of the motivations for the first Answer Validation Exercise at CLEF 2006 (AVE 2006). The starting point for the AVE 2006 was the reformulation of the answer validation as a recognizing textual entailment (RTE) problem, under the assumption that a hypothesis can be automatically generated instantiating a hypothesis pattern with a QA system answer. The test collections that we developed in seven different languages at AVE 2006 are specially oriented to the development and evaluation of answer validation systems. We show in this article the methodology followed for developing these collections taking advantage of the human assessments already made in the evaluation of QA systems. We also propose an evaluation framework for AV linked to a QA evaluation track. We quantify and discuss the source of errors introduced by the reformulation of the answer validation problem in terms of textual entailment (around 2%, in the range of inter-annotator disagreement). We also show the evaluation results of the first answer validation exercise at CLEF 2006 where 11 groups have participated with 38 runs in seven different languages. The most extensively used techniques were Machine Learning and overlapping measures, but systems with broader knowledge resources and richer representation formalisms obtained the best results.
publishDate 2007
dc.date.none.fl_str_mv 2007
2007-12-13
2007
2007-12-13
2026
2026-05-08
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/20.500.14468/32544
url https://hdl.handle.net/20.500.14468/32544
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
info:eu-repo/semantics/openAccess
http://creativecommons.org/licenses/by-nc-nd/4.0/deed.es
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
http://creativecommons.org/licenses/by-nc-nd/4.0/deed.es
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Oxford University Press
publisher.none.fl_str_mv Oxford University Press
dc.source.none.fl_str_mv reponame:e-spacio. Repositorio Institucional de la UNED
instname:Universidad Nacional de Educación a Distancia
instname_str Universidad Nacional de Educación a Distancia
reponame_str e-spacio. Repositorio Institucional de la UNED
collection e-spacio. Repositorio Institucional de la UNED
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869418473500704768
score 15.812429