Enhancing Location Entity Recognition in Spanish Medical Texts by Leveraging Domain Language Models and Data Augmentation

This work focuses on the automatic recognition of location entities in Spanish clinical reports, using the MEDDOPLACE challenge (IberLEF 2023) as the experimental framework. We evaluated both general-domain pre-trained models and biomedical-specific models. Furthermore, we explored data augmentation...

Descripción completa

Detalles Bibliográficos
Autores: Garitano López de Uralde, Irati, Martínez Unanue, Raquel
Tipo de recurso: artículo
Fecha de publicación:2026
País:España
Institución:Universidad Nacional de Educación a Distancia
Repositorio:e-spacio (DSpace). Repositorio Institucional de la UNED
Idioma:inglés
OAI Identifier:oai:dnet:e-spacio(ds_::edf98fdb39b7c5070ec9e80525d9ba1f
Acceso en línea:https://hdl.handle.net/20.500.14468/32213
Access Level:acceso abierto
Palabra clave:1203.18 Sistemas de información, diseño y componentes
location entity recognition
data augmentation
medical domain
domain-specific language models
reconocimiento de entidades de lugar
aumento de datos
dominio médico
modelos de lenguaje de dominio espec´ıfico
Descripción
Sumario:This work focuses on the automatic recognition of location entities in Spanish clinical reports, using the MEDDOPLACE challenge (IberLEF 2023) as the experimental framework. We evaluated both general-domain pre-trained models and biomedical-specific models. Furthermore, we explored data augmentation techniques via back-translation and LLM-based paraphrase generation. Our results outperform previous state-of-the-art approaches, demonstrating the effectiveness of combining these data augmentation strategies with pre-trained clinical domain models.