Skeleton-based estimation of interaction readiness via spatio-temporal graph convolution

Estimating a person’s readiness to engage is a critical prerequisite for achieving natural interactions with machines. In this work, we explore a skeleton-based approach for estimating interaction readiness from human motion. A two-stream spatio-temporal graph convolutional network is applied as bac...

ver descrição completa

Detalhes bibliográficos
Autores: Yuan, Junze, Mohammed, Wael M., Ferre Pérez, Manuel, Pérez de la Lastra, José Manuel
Formato: artículo
Estado:Versión publicada
Fecha de publicación:2026
País:España
Recursos:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:dnet:digitalcsic_::7dff520fffe82dc1b813fe046af69ff5
Acesso em linha:http://hdl.handle.net/10261/429802
Access Level:acceso abierto
Palavra-chave:Graph convolutional network
Human-robot interaction
Proactive interaction
Interaction readiness estimation
Descrição
Resumo:Estimating a person’s readiness to engage is a critical prerequisite for achieving natural interactions with machines. In this work, we explore a skeleton-based approach for estimating interaction readiness from human motion. A two-stream spatio-temporal graph convolutional network is applied as backbone. We propose two design changes to the backbone: Local Dense Connection (LDC), which enhances the flow of multi-scale features, and Cross-Stream Attention (CSA) module, allowing it to effectively relate joint and bone features. Rather than classifying actions directly, a probabilistic aggregation strategy is introduced to generate a scalar measure of interaction readiness, which helps the model generalize better to real-world scenes. Experiment on the processed NTU-RGB+D 120 dataset demonstrates the proposed method achieves 82.52% top-1 accuracy, outperforming backbone model. Moreover, experiment on real-world data achieves ROC-AUC = 0.9687 in realistic conditions, indicating the robustness and generalization ability of the proposed method with lightweight parameters (8.30 M). While relatively lightweight, the method offers a practical solution for scenarios that require fast, interpretable estimation such as human-robot interaction settings.