Combining the power of large language models with finetuning based on strategically collected human ratings: A case study about age-of-acquisition estimates of Spanish words

This study examined the ability of a large language model, GPT-4o mini, to predict age of acquisition (AoA) for Spanish words, as compared to human ratings. We found a strong correlation (ρ=.75) between the model's AoA estimates and mean human ratings. This correlation was lower than the level...

Descripción completa

Detalles Bibliográficos
Autores: Sendín, Eneko, Conde, Javier, Reviriego, Pedro, Haro, Juan, Ferré, Pilar, Hinojosa, José A, Brysbaert, Marc
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2025
País:España
Institución:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/399562
Acceso en línea:http://hdl.handle.net/10261/399562
https://doi.org/10.20350/digitalCSIC/17563
Access Level:acceso abierto
Palabra clave:Behavioral neuroscience
Cognitive psychology
Experimental psychology
Descripción
Sumario:This study examined the ability of a large language model, GPT-4o mini, to predict age of acquisition (AoA) for Spanish words, as compared to human ratings. We found a strong correlation (ρ=.75) between the model's AoA estimates and mean human ratings. This correlation was lower than the level of agreement observed between individual human raters (ρ=.85), but we found that finetuning the model on a relatively small dataset of 2000 human AoA ratings has the potential to enhance the model's performance to a level comparable to human consensus. Consistent with theoretical expectations, our analyses confirmed that AoA estimates are meaningful only for words within an individual's vocabulary. Finally, we present a novel dataset of AoA estimates for 28,453 Spanish words likely known by adult speakers.