3D Bounding box detection from monocular images
Object detection is particularly important in robotic applications that require interaction with the environment. Although 2D object detection methods obtain accurate results, these are not enough to provide a complete description of the 3D scenario. Therefore, many models have recently showed promi...
| Autor: | |
|---|---|
| Tipo de recurso: | tesis de maestría |
| Fecha de publicación: | 2019 |
| País: | España |
| Institución: | Universitat Politècnica de Catalunya (UPC) |
| Repositorio: | UPCommons. Portal del coneixement obert de la UPC |
| Idioma: | inglés |
| OAI Identifier: | oai:upcommons.upc.edu:2117/133405 |
| Acceso en línea: | https://hdl.handle.net/2117/133405 |
| Access Level: | acceso abierto |
| Palabra clave: | Machine learning Computer vision Autonomous vehicles object detection autonomous driving Aprenentatge automàtic Visió per ordinador Vehicles autònoms Àrees temàtiques de la UPC::Enginyeria de la telecomunicació |
| Sumario: | Object detection is particularly important in robotic applications that require interaction with the environment. Although 2D object detection methods obtain accurate results, these are not enough to provide a complete description of the 3D scenario. Therefore, many models have recently showed promising progresses in this challenging field [5, 22, 25, 30]. In this work, the goal is to predict 3D bounding boxes from single images without using temporal data nor any explicit depth estimation. We propose an approach for 3D monocular object detection based on Deep3DBox [20]. We aim to replace the geometric constraints taken into account to predict the 3D location of objects by a deep learning module. Moreover, we undertake a study on the different parameters for the modules that are used to predict dimensions and orientation of objects. We conduct experiments in order to search for the best hyperparameters of our model for KITTI [7] cars and we reported and compared our results on KITTI and the challenging NuScenes [2] benchmarks for cars and pedestrians with other state of the art methods. Therefore, we conclude that our approach performs on par with similar methods [22, 30] and improves Deep3DBox [20] results. |
|---|