Improving Radiology Report Generation Quality and Diversity through Reinforcement Learning and Text Augmentation

[EN] Deep learning is revolutionizing radiology report generation (RRG) with the adoption of vision encoder--decoder (VED) frameworks, which transform radiographs into detailed medical reports. Traditional methods, however, often generate reports of limited diversity and struggle with generalization...

Descripción completa

Detalles Bibliográficos
Autores: Parres-Montoya, Daniel, Albiol Colomer, Alberto|||0000-0002-1970-3289, Paredes Palacios, Roberto
Tipo de recurso: artículo
Fecha de publicación:2024
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/208655
Acceso en línea:https://riunet.upv.es/handle/10251/208655
Access Level:acceso abierto
Palabra clave:Radiology report generation
Reinforcement learning
Text augmentation
Machine learning
Deep learning
Vision transformer
Chest X-rays
Medical image
Text generation
LENGUAJES Y SISTEMAS INFORMATICOS
TEORÍA DE LA SEÑAL Y COMUNICACIONES
Descripción
Sumario:[EN] Deep learning is revolutionizing radiology report generation (RRG) with the adoption of vision encoder--decoder (VED) frameworks, which transform radiographs into detailed medical reports. Traditional methods, however, often generate reports of limited diversity and struggle with generalization. Our research introduces reinforcement learning and text augmentation to tackle these issues, significantly improving report quality and variability. By employing RadGraph as a reward metric and innovating in text augmentation, we surpass existing benchmarks like BLEU4, ROUGE-L, F1CheXbert, and RadGraph, setting new standards for report accuracy and diversity on MIMIC-CXR and Open-i datasets. Our VED model achieves F1-scores of 66.2 for CheXbert and 37.8 for RadGraph on the MIMIC-CXR dataset, and 54.7 and 45.6, respectively, on Open-i. These outcomes represent a significant breakthrough in the RRG field. The findings and implementation of the proposed approach, aimed at enhancing diagnostic precision and radiological interpretations in clinical settings, are publicly available on GitHub to encourage further advancements in the field.