PIRO: Permutation-invariant relational network for multi-person 3D pose estimation

Recovering multi-person 3D poses from a single RGB image is an ill-conditioned problem due to the inherent 2D-3D depth ambiguity, inter-person occlusions, and body truncation. To tackle these issues, recent works have shown promising results by simultaneously reasoning for different individuals. How...

ver descrição completa

Detalhes bibliográficos
Autores: Ugrinovic, Nicolás, Ruiz Ovejero, Adrià, Agudo Martínez, Antonio, Sanfeliu, Alberto, Moreno-Noguer, Francesc
Tipo de documento: artigo
Estado:Versión aceptada para publicación
Data de publicação:2024
País:España
Recursos:Consejo Superior de Investigaciones Científicas (CSIC)
Repositório:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/388086
Acesso em linha:http://hdl.handle.net/10261/388086
Access Level:Acceso aberto
Palavra-chave:Human pose estimation
3D
Single-view
Descrição
Resumo:Recovering multi-person 3D poses from a single RGB image is an ill-conditioned problem due to the inherent 2D-3D depth ambiguity, inter-person occlusions, and body truncation. To tackle these issues, recent works have shown promising results by simultaneously reasoning for different individuals. However, in most cases this is done by only considering pairwise inter-person interactions or between pairs of body parts, thus hindering a holistic scene representation able to capture long-range interactions. Some approaches that jointly process all people in the scene require defining one of the individuals as a reference and a pre-defined person ordering or limiting the number of individuals thus being sensitive to these choice. In this paper, we overcome both these limitations, and we propose an approach for multi-person 3D pose estimation that captures longrange interactions independently of the input order. We build a residual-like permutation-invariant network that successfully refines potentially corrupted initial 3D poses estimated by off-the-shelf detectors. The residual function is learned via a Set Attention (Lee et al., 2019) mechanism. Despite of our model being relatively straightforward, a thorough evaluation demonstrates that our approach is able to boost the performance of the initially estimated 3D poses by large margins, achieving state-of-the-art results on two standardized benchmarks