Joint generation of distractors for multiple-choice questions: a text-to-text approach

Generation of good-quality distractors is a key and time-consuming task associated with multiple-choice questions (MCQs), one of the assessment items that have dominated the educational field for years. Recent advances in language models and architectures present an opportunity for helping teachers...

Descripción completa

Detalles Bibliográficos
Autores: Rodriguez Torrealba, Ricardo, García López, Eva|||0000-0002-7598-3289, García Cabot, Antonio|||0000-0002-0298-3237
Tipo de recurso: artículo
Fecha de publicación:2025
País:España
Institución:Universidad de Alcalá (UAH)
Repositorio:e_Buah Biblioteca Digital Universidad de Alcalá
Idioma:inglés
OAI Identifier:oai:ebuah.uah.es:10017/67456
Acceso en línea:http://hdl.handle.net/10017/67456
https://dx.doi.org/10.32604/cmc.2025.062004
Access Level:acceso abierto
Palabra clave:Text-to-text
Distractor generation
Fine-tuning
FlanT5
LongT5
Multiple-choice
Questionnaire
Informática
Computer science
Descripción
Sumario:Generation of good-quality distractors is a key and time-consuming task associated with multiple-choice questions (MCQs), one of the assessment items that have dominated the educational field for years. Recent advances in language models and architectures present an opportunity for helping teachers to generate and update these elements to the required speed and scale of widespread increase in online education. This study focuses on a text-to-text approach for joints generation of distractors for MCQs, where the context, question and correct answer are used as input, while the set of distractors corresponds to the output, allowing the generation of three distractors in a singlemodel inference. By fine-tuning FlanT5 models and LongT5 with TGlobal attention using a RACE-based dataset, the potential of this approach is explored, demonstrating an improvement in the BLEU and ROUGE-L metrics when compared to previous works and a GPT-3.5 baseline. Additionally, BERTScore is introduced in the evaluation, showing that the fine-tuned models generate distractors semantically close to the reference, but the GPT-3.5 baseline still outperforms in this area. A tendency toward duplicating distractors is noted, although models fine-tuned with Low-Rank Adaptation (LoRA) and 4-bit quantization showcased a significant reduction in duplicated distractors.