Deep learning-based video analysis for visitor detection and tracking in protected areas

Collecting visitor data in protected areas (PA) is often costly, and traditional camera-based methods remain inefficient due to the substantial manual effort required for image processing. This study pioneers integrating camera traps (CT) with Deep Learning (DL) models employing object-detection-bas...

Descripción completa

Detalles Bibliográficos
Autores: Moreno, Hugo, Gómez, Adrià, Andújar, Dionisio
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2025
País:España
Institución:Consejo Superior de Investigaciones Científicas (CSIC)
Repositorio:DIGITAL.CSIC. Repositorio Institucional del CSIC
OAI Identifier:oai:digital.csic.es:10261/421037
Acceso en línea:http://hdl.handle.net/10261/421037
Access Level:acceso abierto
Palabra clave:Monitoring
Camera trap
Protected area management
Recreation trail
Descripción
Sumario:Collecting visitor data in protected areas (PA) is often costly, and traditional camera-based methods remain inefficient due to the substantial manual effort required for image processing. This study pioneers integrating camera traps (CT) with Deep Learning (DL) models employing object-detection-based Convolutional Neural Networks (CNN) to develop an offline visitor-counting algorithm tailored for video analysis rather than static images. Different CT models to record human activity through video were deployed in the Spanish National Parks of Sierra de las Nieves and Sierra de Guadarrama. Visitors in staggered positions and obscuring each other represent a typical video analysis issue. To address this limitation, the proposed algorithm deploys multiple virtual counting lines, i.e., crossing lines and internal identifiers for individual visitors. The algorithm achieved accurate tracking regardless of the video's length, effectively handling challenges associated with varying recording durations. Furthermore, the algorithm was able to be executed in a high- and medium-low-quality computing environment with no difficulties. Moreover, the study included variations in camera position, perspective, weather, time of day, video resolution, length, and format, evaluating the algorithm's robustness in real-world scenarios. The CNN models implemented through the algorithm achieved high accuracies, with the best results ranging from 98.48 % to 99.77 %. In addition, high performances were achieved regardless of the CNN employed, the frame rate, and the variability of the video set. Therefore, this study proposes a monitoring system that provides an affordable, fast, and accurate approach for PA and other recreational landscapes, reliably tracking visitors in remote areas lacking real-time network coverage.