A deep attention-based encoder for the prediction of type 2 diabetes longitudinal outcomes from routinely collected health care data

Recent evidence indicates that Type 2 Diabetes Mellitus (T2DM) is a complex and highly heterogeneous disease involving various pathophysiological and genetic pathways, which presents clinicians with challenges in disease management. While deep learning models have made significant progress in helpin...

Descripción completa

Detalles Bibliográficos
Autores: Manzini, E, Vlacho, B, Franch-Nadal, J, Escudero, J, Génova, A, Reixach, E, Andrés, E, Pizarro, I, Mauricio, D, Perera-Lluna, A
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2025
País:España
Institución:Institut d’Investigació Biomèdica Sant Pau (IIB Sant Pau)
Repositorio:r-IIB SANT PAU. Repositorio Institucional de Producción Científica del Instituto de Investigación Biomédica Sant Pau
OAI Identifier:oai:iibsantpau.fundanetsuite.com:p19389
Acceso en línea:https://iibsantpau.fundanetsuite.com/Publicaciones/ProdCientif/PublicacionFrw.aspx?id=19389
Access Level:acceso abierto
Palabra clave:Deep learning
Transformer
Type 2 diabetes
Diabetes complications
Electronic health records
Descripción
Sumario:Recent evidence indicates that Type 2 Diabetes Mellitus (T2DM) is a complex and highly heterogeneous disease involving various pathophysiological and genetic pathways, which presents clinicians with challenges in disease management. While deep learning models have made significant progress in helping practitioners manage T2DM treatments, several important limitations persist. In this paper we propose DARE, a model based on the transformer encoder, designed for analyzing longitudinal heterogeneous diabetes data. The model can be easily fine-tuned for various clinical prediction tasks, enabling a computational approach to assist clinicians in the management of the disease. We trained DARE using data from over 200,000 diabetic subjects from the primary healthcare SIDIAP database, which includes diagnosis and drug codes, along with various clinical and analytical measurements. After an unsupervised pre-training phase, we fine-tuned the model for predicting three specific clinical outcomes: i) occurrence of comorbidity, ii) achievement of target glycemic control (defined as glycated hemoglobin <7%) and iii) changes in glucose-lowering treatment. In cross-validation, the embedding vectors generated by DARE outperformed those from baseline models (comorbidities prediction task AUC = 0.88, treatment prediction task AUC = 0.91, HbA1c target prediction task AUC = 0.82). Our findings suggest that attention-based encoders improve results with respect to different deep learning and classical baseline models when used to predict different clinical relevant outcomes from T2DM longitudinal data.