Aprendizaje por refuerzo relacional con acciones continuas
Reinforcement Learning (RL) is a commonly used technique for learning tasks in robotics. This is mainly because it allows agents, i.e., robots, to develop optimal control policies through trial and error interactions with the environment in which these robots perform and because it does not require...
| Author: | |
|---|---|
| Format: | master thesis |
| Status: | Versión aceptada para publicación |
| Publication Date: | 2009 |
| Country: | México |
| Institution: | Instituto Nacional de Astrofísica, Óptica y Electrónica |
| Repository: | Repositorio Institucional del INAOE |
| Language: | Spanish |
| OAI Identifier: | oai:inaoe.repositorioinstitucional.mx:1009/387 |
| Online Access: | http://inaoe.repositorioinstitucional.mx/jspui/handle/1009/387 |
| Access Level: | Open access |
| Keyword: | info:eu-repo/classification/Inteligencia artificial/Artificial intelligence info:eu-repo/classification/Análisis de regresión/Regression analysis info:eu-repo/classification/Álgebra Relacional/Relational algebra info:eu-repo/classification/cti/1 info:eu-repo/classification/cti/12 info:eu-repo/classification/cti/1203 |
| Summary: | Reinforcement Learning (RL) is a commonly used technique for learning tasks in robotics. This is mainly because it allows agents, i.e., robots, to develop optimal control policies through trial and error interactions with the environment in which these robots perform and because it does not require a previous model of such environment. However traditional RL algorithms require long training times which can be several hours, are unnable to re-use learned policies in similar domains or similar tasks and perform discrete actions. In large search spaces with thousands of states, the policy generation process takes some hours and besides, once a policy has been generated, if the goal or the environment changes, a new policy has to be generated in order to take into account such changes. Finally, discrete actions produce imprecise movements by the robot which can accumulate an error up to tens of degrees for turning actions and up to tens of centimeters for displacement actions. Besides, discrete actions produce slower paths than continuous actions since, with discrete actions, the robot needs to stop in order to turn in discrete angles increasing, every time it stops, the tasks' execution times. In this work, a two stage method to tackle these problems is presented. In the rst stage, the low level sensor information coming from the robot's sensors is transformed into a relational description based on rooms, corridors, doors, walls and obstacles to characterize states and actions, signicantly reducing the state space. Behavoural Cloning (BC), i.e., traces provided by the user, are used to learn in few iterations, a control policy, which, due to the relational representation, can be re-used in similar but dierent domains or environments. However, this policy uses discrete actions. In the second stage, Locally Weighted Regression (LWR) is used to transform the discrete actions policy into a continuous actions policy. The method was used to generate control policies for navigation and following tasks for simulated and real mobile robots with very promising results. The results show that the policies are learned after few iterations, can be used on dierent domains, perform smoother, faster and shorter paths than the original relational policies and the tasks' quality is similar to the traces provided by the user. |
|---|