The LambdaGap Framework for Precision-Oriented Ranking

LambdaRank has proven effective for optimizing information retrieval metrics such as Normalized Discounted Cumulative Gain (NDCG). However, its application to Precision at document k (P@k) poses significant challenges because of the metric's unique definition, which heavily restricts the number...

Descripción completa

Detalles Bibliográficos
Autores: Adàlia, Ramon|||0009-0004-9458-1922, Sanjuan, Gemma|||0000-0002-1946-4345, Margalef, Tomàs|||0000-0001-6384-7389, Zamora, Ismael|||0000-0002-7700-0354
Tipo de recurso: artículo
Fecha de publicación:2025
País:España
Institución:Universitat Autònoma de Barcelona
Repositorio:Dipòsit Digital de Documents de la UAB
Idioma:inglés
OAI Identifier:oai:ddd.uab.cat:317648
Acceso en línea:https://ddd.uab.cat/record/317648
https://dx.doi.org/urn:doi:10.1145/3733235
Access Level:acceso abierto
Palabra clave:Learning to Rank
LambdaRank
Ranking Metric Optimization
Descripción
Sumario:LambdaRank has proven effective for optimizing information retrieval metrics such as Normalized Discounted Cumulative Gain (NDCG). However, its application to Precision at document k (P@k) poses significant challenges because of the metric's unique definition, which heavily restricts the number of effective training document pairs. This limitation diminishes the learning signal for relevant documents beyond the top k, potentially resulting in suboptimal performance. To overcome this, we propose LambdaGap, a ranking algorithm inspired by LambdaRank specifically tailored for optimizing P@k. LambdaGap replaces the pairwise weighting scheme in LambdaRank by one where pairs of documents within k positions in the ranking are masked out. We establish a theoretical link between LambdaGap and P@k by identifying the implicit metric optimized by the model. Furthermore, we introduce a new metric, Average Relevance Position beyond document k, which can be used in conjunction with LambdaRank to indirectly optimize for P@k. Our extensive experiments on publicly available datasets demonstrate the effectiveness of the proposed methods, yielding statistically significant improvements in P@k performance and highlighting their potential for more efficient training.