Skeleton-based estimation of interaction readiness via spatio-temporal graph convolution
Estimating a person’s readiness to engage is a critical prerequisite for achieving natural interactions with machines. In this work, we explore a skeleton-based approach for estimating interaction readiness from human motion. A two-stream spatio-temporal graph convolutional network is applied as bac...
| Autores: | , , , |
|---|---|
| Formato: | artículo |
| Estado: | Versión publicada |
| Fecha de publicación: | 2026 |
| País: | España |
| Recursos: | Consejo Superior de Investigaciones Científicas (CSIC) |
| Repositorio: | DIGITAL.CSIC. Repositorio Institucional del CSIC |
| OAI Identifier: | oai:dnet:digitalcsic_::7dff520fffe82dc1b813fe046af69ff5 |
| Acesso em linha: | http://hdl.handle.net/10261/429802 |
| Access Level: | acceso abierto |
| Palavra-chave: | Graph convolutional network Human-robot interaction Proactive interaction Interaction readiness estimation |
| Resumo: | Estimating a person’s readiness to engage is a critical prerequisite for achieving natural interactions with machines. In this work, we explore a skeleton-based approach for estimating interaction readiness from human motion. A two-stream spatio-temporal graph convolutional network is applied as backbone. We propose two design changes to the backbone: Local Dense Connection (LDC), which enhances the flow of multi-scale features, and Cross-Stream Attention (CSA) module, allowing it to effectively relate joint and bone features. Rather than classifying actions directly, a probabilistic aggregation strategy is introduced to generate a scalar measure of interaction readiness, which helps the model generalize better to real-world scenes. Experiment on the processed NTU-RGB+D 120 dataset demonstrates the proposed method achieves 82.52% top-1 accuracy, outperforming backbone model. Moreover, experiment on real-world data achieves ROC-AUC = 0.9687 in realistic conditions, indicating the robustness and generalization ability of the proposed method with lightweight parameters (8.30 M). While relatively lightweight, the method offers a practical solution for scenarios that require fast, interpretable estimation such as human-robot interaction settings. |
|---|