Multistage strategy for ground point filtering on large-scale datasets

[EN] Ground point filtering on national-level datasets is a challenge due to the presence of multiple types of landscapes. This limitation does not simply affect to individual users, but it is in particular relevant for those national institutions in charge of providing national-level Light Detectio...

Full description

Bibliographic Details
Authors: Teijeiro Paredes, Diego, Amor López, Margarita, Buján Seoane, Sandra, Richter, Rico, Döllner, Jürgen
Format: article
Status:Published version
Publication Date:2024
Country:España
Institution:Ajuntament de Barcelona
Repository:BULERIA. Repositorio Institucional de la Universidad de León
OAI Identifier:oai:buleria.unileon.es:10612/24141
Online Access:https://link.springer.com/article/10.1007/s11227-024-06406-0
https://hdl.handle.net/10612/24141
Access Level:Open access
Keyword:Informática
Ingenierías
Topografía
LiDAR point clouds
Landscape identification
Ground filtering
Apache spark
Description
Summary:[EN] Ground point filtering on national-level datasets is a challenge due to the presence of multiple types of landscapes. This limitation does not simply affect to individual users, but it is in particular relevant for those national institutions in charge of providing national-level Light Detection and Ranging (LiDAR) point clouds. Each type of landscape is typically better filtered by different filtering algorithms or parameters; therefore, in order to get the best quality classification, the LiDAR point cloud should be divided by the landscape before running the filtering algorithms. Despite the fact that the manual segmentation and identification of the landscapes can be very time intensive, only few studies have addressed this issue. In this work, we present a multistage approach to automate the identification of the type of landscape using several metrics extracted from the LiDAR point cloud, matching the best filtering algorithms in each type of landscape. An additional contribution is presented, a parallel implementation for distributed memory systems, using Apache Spark, that can achieve up to 34× of speedup using 12 compute node