Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine

The rapid progress in artificial intelligence, machine learning, and natural language processing has led to increasingly sophisticated large language models (LLMs) for use in healthcare. This study assesses the performance of two LLMs, the GPT-3.5 and GPT-4 models, in passing the MIR medical examina...

ver descrição completa

Detalhes bibliográficos
Autores: Guillén Grima, Francisco, Guillén Aguinaga, Sara, Guillén Aguinaga, Laura, Alas Brun, Rosa María, Onambele, Luc, Ortega-León, Wilfrido, Montejo, Rocío, Aguinaga Ontoso, Enrique, Barach, Paul, Aguinaga Ontoso, Inés
Formato: artículo
Estado:Versión publicada
Fecha de publicación:2023
País:España
Recursos:Universidad Pública de Navarra
Repositorio:Academica-e. Repositorio Institucional de la Universidad Pública de Navarra
OAI Identifier:oai:academica-e.unavarra.es:2454/48088
Acesso em linha:https://hdl.handle.net/2454/48088
Access Level:acceso abierto
Palavra-chave:Artificial intelligence
ChatGPT
GPT-3.5
GPT-4
Image
Large language model
Machine learning
Medical education
Patient safety
Quality of care
id ES_5df749e2c9eaa1c7db7c41935f912e7c
oai_identifier_str oai:academica-e.unavarra.es:2454/48088
network_acronym_str ES
network_name_str España
repository_id_str
spelling Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicineGuillén Grima, FranciscoGuillén Aguinaga, SaraGuillén Aguinaga, LauraAlas Brun, Rosa MaríaOnambele, LucOrtega-León, WilfridoMontejo, RocíoAguinaga Ontoso, EnriqueBarach, PaulAguinaga Ontoso, InésArtificial intelligenceChatGPTGPT-3.5GPT-4ImageLarge language modelMachine learningMedical educationPatient safetyQuality of careThe rapid progress in artificial intelligence, machine learning, and natural language processing has led to increasingly sophisticated large language models (LLMs) for use in healthcare. This study assesses the performance of two LLMs, the GPT-3.5 and GPT-4 models, in passing the MIR medical examination for access to medical specialist training in Spain. Our objectives included gauging the model’s overall performance, analyzing discrepancies across different medical specialties, discerning between theoretical and practical questions, estimating error proportions, and assessing the hypothetical severity of errors committed by a physician. Material and methods: We studied the 2022 Spanish MIR examination results after excluding those questions requiring image evaluations or having acknowledged errors. The remaining 182 questions were presented to the LLM GPT-4 and GPT-3.5 in Spanish and English. Logistic regression models analyzed the relationships between question length, sequence, and performance. We also analyzed the 23 questions with images, using GPT-4’s new image analysis capability. Results: GPT-4 outperformed GPT-3.5, scoring 86.81% in Spanish (p < 0.001). English translations had a slightly enhanced performance. GPT-4 scored 26.1% of the questions with images in English. The results were worse when the questions were in Spanish, 13.0%, although the differences were not statistically significant (p = 0.250). Among medical specialties, GPT-4 achieved a 100% correct response rate in several areas, and the Pharmacology, Critical Care, and Infectious Diseases specialties showed lower performance. The error analysis revealed that while a 13.2% error rate existed, the gravest categories, such as “error requiring intervention to sustain life” and “error resulting in death”, had a 0% rate. Conclusions: GPT-4 performs robustly on the Spanish MIR examination, with varying capabilities to discriminate knowledge across specialties. While the model’s high success rate is commendable, understanding the error severity is critical, especially when considering AI’s potential role in real-world medical practice and its implications for patient safety.MDPICiencias de la SaludOsasun Zientziak2023info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionapplication/pdfapplication/ziphttps://hdl.handle.net/2454/48088reponame:Academica-e. Repositorio Institucional de la Universidad Pública de Navarrainstname:Universidad Pública de NavarraInglés© 2023 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.https://creativecommons.org/licenses/by/4.0/info:eu-repo/semantics/openAccessoai:academica-e.unavarra.es:2454/480882026-06-17T12:41:47Z
dc.title.none.fl_str_mv Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
title Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
spellingShingle Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
Guillén Grima, Francisco
Artificial intelligence
ChatGPT
GPT-3.5
GPT-4
Image
Large language model
Machine learning
Medical education
Patient safety
Quality of care
title_short Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
title_full Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
title_fullStr Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
title_full_unstemmed Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
title_sort Evaluating the efficacy of ChatGPT in navigating the spanish medical residency entrance examination (MIR): promising horizons for AI in clinical medicine
dc.creator.none.fl_str_mv Guillén Grima, Francisco
Guillén Aguinaga, Sara
Guillén Aguinaga, Laura
Alas Brun, Rosa María
Onambele, Luc
Ortega-León, Wilfrido
Montejo, Rocío
Aguinaga Ontoso, Enrique
Barach, Paul
Aguinaga Ontoso, Inés
author Guillén Grima, Francisco
author_facet Guillén Grima, Francisco
Guillén Aguinaga, Sara
Guillén Aguinaga, Laura
Alas Brun, Rosa María
Onambele, Luc
Ortega-León, Wilfrido
Montejo, Rocío
Aguinaga Ontoso, Enrique
Barach, Paul
Aguinaga Ontoso, Inés
author_role author
author2 Guillén Aguinaga, Sara
Guillén Aguinaga, Laura
Alas Brun, Rosa María
Onambele, Luc
Ortega-León, Wilfrido
Montejo, Rocío
Aguinaga Ontoso, Enrique
Barach, Paul
Aguinaga Ontoso, Inés
author2_role author
author
author
author
author
author
author
author
author
dc.contributor.none.fl_str_mv Ciencias de la Salud
Osasun Zientziak
dc.subject.none.fl_str_mv Artificial intelligence
ChatGPT
GPT-3.5
GPT-4
Image
Large language model
Machine learning
Medical education
Patient safety
Quality of care
topic Artificial intelligence
ChatGPT
GPT-3.5
GPT-4
Image
Large language model
Machine learning
Medical education
Patient safety
Quality of care
description The rapid progress in artificial intelligence, machine learning, and natural language processing has led to increasingly sophisticated large language models (LLMs) for use in healthcare. This study assesses the performance of two LLMs, the GPT-3.5 and GPT-4 models, in passing the MIR medical examination for access to medical specialist training in Spain. Our objectives included gauging the model’s overall performance, analyzing discrepancies across different medical specialties, discerning between theoretical and practical questions, estimating error proportions, and assessing the hypothetical severity of errors committed by a physician. Material and methods: We studied the 2022 Spanish MIR examination results after excluding those questions requiring image evaluations or having acknowledged errors. The remaining 182 questions were presented to the LLM GPT-4 and GPT-3.5 in Spanish and English. Logistic regression models analyzed the relationships between question length, sequence, and performance. We also analyzed the 23 questions with images, using GPT-4’s new image analysis capability. Results: GPT-4 outperformed GPT-3.5, scoring 86.81% in Spanish (p < 0.001). English translations had a slightly enhanced performance. GPT-4 scored 26.1% of the questions with images in English. The results were worse when the questions were in Spanish, 13.0%, although the differences were not statistically significant (p = 0.250). Among medical specialties, GPT-4 achieved a 100% correct response rate in several areas, and the Pharmacology, Critical Care, and Infectious Diseases specialties showed lower performance. The error analysis revealed that while a 13.2% error rate existed, the gravest categories, such as “error requiring intervention to sustain life” and “error resulting in death”, had a 0% rate. Conclusions: GPT-4 performs robustly on the Spanish MIR examination, with varying capabilities to discriminate knowledge across specialties. While the model’s high success rate is commendable, understanding the error severity is critical, especially when considering AI’s potential role in real-world medical practice and its implications for patient safety.
publishDate 2023
dc.date.none.fl_str_mv 2023
dc.type.none.fl_str_mv info:eu-repo/semantics/article
info:eu-repo/semantics/publishedVersion
format article
status_str publishedVersion
dc.identifier.none.fl_str_mv https://hdl.handle.net/2454/48088
url https://hdl.handle.net/2454/48088
dc.language.none.fl_str_mv Inglés
language_invalid_str_mv Inglés
dc.rights.none.fl_str_mv https://creativecommons.org/licenses/by/4.0/
info:eu-repo/semantics/openAccess
rights_invalid_str_mv https://creativecommons.org/licenses/by/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
application/zip
dc.publisher.none.fl_str_mv MDPI
publisher.none.fl_str_mv MDPI
dc.source.none.fl_str_mv reponame:Academica-e. Repositorio Institucional de la Universidad Pública de Navarra
instname:Universidad Pública de Navarra
instname_str Universidad Pública de Navarra
reponame_str Academica-e. Repositorio Institucional de la Universidad Pública de Navarra
collection Academica-e. Repositorio Institucional de la Universidad Pública de Navarra
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869409068995575808
score 15,812429