Deep learning-based video analysis for visitor detection and tracking in protected areas
Collecting visitor data in protected areas (PA) is often costly, and traditional camera-based methods remain inefficient due to the substantial manual effort required for image processing. This study pioneers integrating camera traps (CT) with Deep Learning (DL) models employing object-detection-bas...
| Autores: | , , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión publicada |
| Fecha de publicación: | 2025 |
| País: | España |
| Institución: | Consejo Superior de Investigaciones Científicas (CSIC) |
| Repositorio: | DIGITAL.CSIC. Repositorio Institucional del CSIC |
| OAI Identifier: | oai:digital.csic.es:10261/421037 |
| Acceso en línea: | http://hdl.handle.net/10261/421037 |
| Access Level: | acceso abierto |
| Palabra clave: | Monitoring Camera trap Protected area management Recreation trail |
| Sumario: | Collecting visitor data in protected areas (PA) is often costly, and traditional camera-based methods remain inefficient due to the substantial manual effort required for image processing. This study pioneers integrating camera traps (CT) with Deep Learning (DL) models employing object-detection-based Convolutional Neural Networks (CNN) to develop an offline visitor-counting algorithm tailored for video analysis rather than static images. Different CT models to record human activity through video were deployed in the Spanish National Parks of Sierra de las Nieves and Sierra de Guadarrama. Visitors in staggered positions and obscuring each other represent a typical video analysis issue. To address this limitation, the proposed algorithm deploys multiple virtual counting lines, i.e., crossing lines and internal identifiers for individual visitors. The algorithm achieved accurate tracking regardless of the video's length, effectively handling challenges associated with varying recording durations. Furthermore, the algorithm was able to be executed in a high- and medium-low-quality computing environment with no difficulties. Moreover, the study included variations in camera position, perspective, weather, time of day, video resolution, length, and format, evaluating the algorithm's robustness in real-world scenarios. The CNN models implemented through the algorithm achieved high accuracies, with the best results ranging from 98.48 % to 99.77 %. In addition, high performances were achieved regardless of the CNN employed, the frame rate, and the variability of the video set. Therefore, this study proposes a monitoring system that provides an affordable, fast, and accurate approach for PA and other recreational landscapes, reliably tracking visitors in remote areas lacking real-time network coverage. |
|---|