Reinforcement learning applied to production planning and control

[EN] The objective of this paper is to examine the use and applications of reinforcement learning (RL) techniques in the production planning and control (PPC) field addressing the following PPC areas: facility resource planning, capacity planning, purchase and supply management, production schedulin...

Descripción completa

Detalles Bibliográficos
Autores: Esteso, Ana|||0000-0003-0379-8786, Peidro Payá, David|||0000-0001-8678-6881, Mula, Josefa|||0000-0002-8447-3387, Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876
Tipo de recurso: artículo
Fecha de publicación:2023
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/196934
Acceso en línea:https://riunet.upv.es/handle/10251/196934
Access Level:acceso abierto
Palabra clave:Artificial intelligence
Machine learning
Reinforcement learning
Deep reinforcement learning
Production planning and control
Industry 4.0
ORGANIZACION DE EMPRESAS
09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación
Descripción
Sumario:[EN] The objective of this paper is to examine the use and applications of reinforcement learning (RL) techniques in the production planning and control (PPC) field addressing the following PPC areas: facility resource planning, capacity planning, purchase and supply management, production scheduling and inventory management. The main RL characteristics, such as method, context, states, actions, reward and highlights, were analysed. The considered number of agents, applications and RL software tools, specifically, programming language, platforms, application programming interfaces and RL frameworks, among others, were identified, and 181 articles were sreviewed. The results showed that RL was applied mainly to production scheduling problems, followed by purchase and supply management. The most revised RL algorithms were model-free and single-agent and were applied to simplified PPC environments. Nevertheless, their results seem to be promising compared to traditional mathematical programming and heuristics/metaheuristics solution methods, and even more so when they incorporate uncertainty or non-linear properties. Finally, RL value-based approaches are the most widely used, specifically Q-learning and its variants and for deep RL, deep Q-networks. In recent years however, the most widely used approach has been the actor-critic method, such as the advantage actor critic, proximal policy optimisation, deep deterministic policy gradient and trust region policy optimisation.