Analysing the Predictability of Language Model Performance

[EN] Can a language model predict for which questions another language model will answer successfully? We investigate the extent to which performance prediction is possible and dissect various factors that influence it. Our experimental setting fine-tunes DeBERTa models, which we call assessors, on...

Descripción completa

Detalles Bibliográficos
Autores: Schellaert, Wout Willy M., Martínez-Plumed, Fernando|||0000-0003-2902-6477, Hernández-Orallo, José|||0000-0001-9746-7632
Tipo de recurso: artículo
Fecha de publicación:2025
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/223356
Acceso en línea:https://riunet.upv.es/handle/10251/223356
Access Level:acceso abierto
Palabra clave:AI evaluation
Assessor Models
Large Language Models
Performance Prediction
Benchmark
Descripción
Sumario:[EN] Can a language model predict for which questions another language model will answer successfully? We investigate the extent to which performance prediction is possible and dissect various factors that influence it. Our experimental setting fine-tunes DeBERTa models, which we call assessors, on the evaluation results of generative language models with up to 128 billion parameters, which we refer to as subject systems. Our analysis spans more than 100 tasks from BIG-bench. We find that the assessors can match and even exceed the subjects' confidence in both refinement and calibration, anticipating failures at near perfect levels for some tasks. We also find that for performance prediction it can be beneficial to learn from the scores on multiple tasks or to learn from the scores of multiple subjects, but both depend on the task at hand. Lastly, we find that large and small subject systems are equally predictable, showing promise for the scalability of the predictability problem.