Dual Indicators to Analyse AI Benchmarks: Difficulty, Discrimination, Ability and Generality

[EN] With the purpose of better analyzing the result of artificial intelligence (AI) benchmarks, we present two indicators on the side of the AI problems, difficulty and discrimination, and two indicators on the side of the AI systems, ability and generality. The first three are adapted from psychom...

Descripción completa

Detalles Bibliográficos
Autores: Martínez-Plumed, Fernando|||0000-0003-2902-6477, Hernández-Orallo, José|||0000-0001-9746-7632
Tipo de recurso: artículo
Fecha de publicación:2020
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/169021
Acceso en línea:https://riunet.upv.es/handle/10251/169021
Access Level:acceso abierto
Palabra clave:Artificial intelligence
Games
Benchmark testing
Task analysis
Adaptation models
Guidelines
Indexes
Artificial intelligence (AI) benchmarks
AI evaluation
Generality
Item response theory (ITR)
LENGUAJES Y SISTEMAS INFORMATICOS
Descripción
Sumario:[EN] With the purpose of better analyzing the result of artificial intelligence (AI) benchmarks, we present two indicators on the side of the AI problems, difficulty and discrimination, and two indicators on the side of the AI systems, ability and generality. The first three are adapted from psychometric models in item response theory (IRT), whereas generality is defined as a new metric that evaluates whether an agent is consistently good at easy problems and bad at difficult ones. We illustrate how these key indicators give us more insight on the results of two popular benchmarks in AI, the Arcade Learning Environment (Atari 2600 games) and the General Video Game AI competition, and we include some guidelines to estimate and interpret these indicators for other AI benchmarks and competitions.