Combining stereo vision and deep learning techniques for object detection in the 3D world

The objective of this project is to develop a deep learning algorithm so that, together with the use of a stereo camera, it is capable of detecting a person and locating them in the 3D world. The person’s location in the x-y plane is obtained from the object detector model, which consists of a convo...

Descripción completa

Detalles Bibliográficos
Autor: Gimeno I Jovés, Andrea
Tipo de recurso: tesis de maestría
Fecha de publicación:2020
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/189684
Acceso en línea:https://hdl.handle.net/2117/189684
Access Level:acceso abierto
Palabra clave:Artificial intelligence
Aprenentatge automàtic
Intel·ligència artificial
Àrees temàtiques de la UPC::Informàtica
Descripción
Sumario:The objective of this project is to develop a deep learning algorithm so that, together with the use of a stereo camera, it is capable of detecting a person and locating them in the 3D world. The person’s location in the x-y plane is obtained from the object detector model, which consists of a convolutional neural network, specifically the U-Net, that outputs heat maps. On the other hand, the person’s location in terms of depth (z) is obtained from the depth map given by the ZED stereo camera. The document begins by presenting the techniques used today for object detection (using heat maps). This is followed by an explanation of the key theory behind neural networks; from the simplest neural networks to the convolutional neural networks. To finish with the theoretical part of the project, the hardware and software equipment used is presented. To develop and implement the deep learning algorithm, the first thing that is done is the dataset creation. In order to do that, different images have been selected and prepared to enter the network and train the model (using PyTorch) adapted to the needs of this task. Eight different combination of parameters have been used and eight different models have been obtained. Previously, the metric that will be used to evaluate and compare the different models obtained and choose the one that best suits this application, is defined. Once the final model is chosen, it is stored in the Jetson AGX Xavier and tested using ZED camera images. In this case, the model is verified to being accurate detecting people and the cases where the algorithm fails are identified. The next step of this project consists of applying stereo vision techniques to extract the distance at which the detected person is. A ROS node is created to communicate the ZED camera with the deep learning algorithm. Once the node is ready, it is executed to test the whole program in real time. The ZED color images are passed through the network to detect the person (x, y), and from the ZED depth map, the distance (z) is obtained. From the results obtained, both for the person detection and for the distance extraction, the existing errors in the designed algorithm are identified, and improvements are made by applying filters and code modifications. Thanks to the improvements applied to the results, a sufficient precise algorithm is obtained, capable of detecting a person within a distance range in real time