Mapping layperson medical terminology into the Human Phenotype Ontology using neural machine translation models

In the medical domain there exists a terminological gap between patients and caregivers and the healthcare professionals. This gap may hinder the success of the communication between healthcare consumers and professionals in the field, with negative emotional and clinical consequences. In this work,...

Descripción completa

Detalles Bibliográficos
Autores: Manzini, E, Garrido-Aguirre, J, Fonollosa, J, Lluna, AP
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2022
País:España
Institución:Fundació Sant Joan de Déu
Repositorio:r-FSJD. Repositorio Institucional de Producción Científica de la Fundació Sant Joan de Déu
OAI Identifier:oai:fsjd.fundanetsuite.com:p22233
Acceso en línea:https://fsjd.fundanetsuite.com/Publicaciones/ProdCientif/PublicacionFrw.aspx?id=22233
Access Level:acceso abierto
Palabra clave:Machine translation
Word embedding
Deep learning
Medical informatics
Deep phenotyping
Human Phenotype Ontology
Descripción
Sumario:In the medical domain there exists a terminological gap between patients and caregivers and the healthcare professionals. This gap may hinder the success of the communication between healthcare consumers and professionals in the field, with negative emotional and clinical consequences. In this work, we build a machine learning-based tool for the automatic translation between the terminology used by laypeople and that of the Human Phenotype Ontology (HPO). HPO is a structured vocabulary of phenotypic abnormalities found in human disease. Our method uses a vector space to represent an HPO-specific embedding as the output space for a neural network model trained on vector representations of layperson versions and other textual descriptors of medical terms. We explored different output embeddings coupled to different neural network architectures for the machine translation stage. We compute a similarity measure to evaluate the ability of the model to assign an HPO term to a layperson input. The best-performing models resulted with a similarity higher than 0.7 for more than 80% of the terms, with a median between 0.98 and 1.