Automated multiple-choice question generation in Spanish using neural language models

This research presents an approach to automatic multiple-choice question (MCQ) generation in the Spanish language, using mT5-based models. The process encompasses three crucial tasks: candidate answer extraction, answer-aware question generation, and distractor generation. A methodical pipeline is s...

Descripción completa

Detalles Bibliográficos
Autores: Fitero Domínguez, David de, García Cabot, Antonio|||0000-0002-0298-3237, García López, Eva|||0000-0002-7598-3289
Tipo de recurso: artículo
Fecha de publicación:2024
País:España
Institución:Universidad de Alcalá (UAH)
Repositorio:e_Buah Biblioteca Digital Universidad de Alcalá
Idioma:inglés
OAI Identifier:oai:ebuah.uah.es:10017/67416
Acceso en línea:http://hdl.handle.net/10017/67416
https://dx.doi.org/10.1007/s00521-024-10076-7
Access Level:acceso abierto
Palabra clave:Machine learning
Transformers
Natural language processing
Text generation
Distractor generation
Informática
Computer science
Descripción
Sumario:This research presents an approach to automatic multiple-choice question (MCQ) generation in the Spanish language, using mT5-based models. The process encompasses three crucial tasks: candidate answer extraction, answer-aware question generation, and distractor generation. A methodical pipeline is structured to seamlessly integrate these tasks, converting an input text into a systematic questionnaire. For model fine-tuning, the Stanford Question Answering Dataset is employed for the first two tasks, while a combination of three different multiple-choice question datasets, translated automatically into Spanish, is used for the distractor generation task. The efficiency of the models is then evaluated by using a triad of metrics, namely BLEU, ROUGE-L, and cosine similarity. The outcomes indicate a marginal deviation from the baseline model in the question generation task but demonstrate superior performance in the distractor generation task. Importantly, this research emphasizes the potential and effectiveness of language models for automating MCQ generation, providing a valuable contribution to the field and enhancing the understanding and application of such models in the context of the Spanish language.