Aggregating the temporal coherent descriptors in videos using multiple learning kernel for action recognition

Action recognition methods enable several intelligent machines to recognize human action in their daily life videos. Indeed, many action recognition methods give a noticeable misclassification rate due to the big variations within the videos of the same class, and the changes in viewpoint, scale and...

Descripción completa

Detalles Bibliográficos
Autores: Saleh, Adel, Abdel-Nasser, Mohamed, García García, Miguel Ángel, Puig, Domenec
Tipo de recurso: artículo
Fecha de publicación:2018
País:España
Institución:Universidad Autónoma de Madrid
Repositorio:Biblos-e Archivo. Repositorio Institucional de la UAM
Idioma:inglés
OAI Identifier:oai:repositorio.uam.es:10486/692482
Acceso en línea:http://hdl.handle.net/10486/692482
https://dx.doi.org/10.1016/j.patrec.2017.06.010
Access Level:acceso abierto
Palabra clave:Action recognition
Representation learning
Coherence analysis
Learning-to-rank
Multiple kernel learning
Telecomunicaciones
Descripción
Sumario:Action recognition methods enable several intelligent machines to recognize human action in their daily life videos. Indeed, many action recognition methods give a noticeable misclassification rate due to the big variations within the videos of the same class, and the changes in viewpoint, scale and background. In this paper, we propose a new video representations method that captures temporal evolution of the action happening in the whole video. We show that 1) combining the descriptors of improved dense trajectories with a multiple kernel learning technique can reduce the misclassification rate, and also 2) aggregating the coherent frames in each video may have a different impact on the recognition results. Our experimental results using HMDB51 and Hollywood datasets demonstrate that our method is on par with the state-of-the-art methods