3D Bounding box detection from monocular images

Object detection is particularly important in robotic applications that require interaction with the environment. Although 2D object detection methods obtain accurate results, these are not enough to provide a complete description of the 3D scenario. Therefore, many models have recently showed promi...

Descripción completa

Detalles Bibliográficos
Autor: Catà Villà, Marcel
Tipo de recurso: tesis de maestría
Fecha de publicación:2019
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/133405
Acceso en línea:https://hdl.handle.net/2117/133405
Access Level:acceso abierto
Palabra clave:Machine learning
Computer vision
Autonomous vehicles
object detection
autonomous driving
Aprenentatge automàtic
Visió per ordinador
Vehicles autònoms
Àrees temàtiques de la UPC::Enginyeria de la telecomunicació
Descripción
Sumario:Object detection is particularly important in robotic applications that require interaction with the environment. Although 2D object detection methods obtain accurate results, these are not enough to provide a complete description of the 3D scenario. Therefore, many models have recently showed promising progresses in this challenging field [5, 22, 25, 30]. In this work, the goal is to predict 3D bounding boxes from single images without using temporal data nor any explicit depth estimation. We propose an approach for 3D monocular object detection based on Deep3DBox [20]. We aim to replace the geometric constraints taken into account to predict the 3D location of objects by a deep learning module. Moreover, we undertake a study on the different parameters for the modules that are used to predict dimensions and orientation of objects. We conduct experiments in order to search for the best hyperparameters of our model for KITTI [7] cars and we reported and compared our results on KITTI and the challenging NuScenes [2] benchmarks for cars and pedestrians with other state of the art methods. Therefore, we conclude that our approach performs on par with similar methods [22, 30] and improves Deep3DBox [20] results.