Reinforcement learning applied to production planning and control

[EN] The objective of this paper is to examine the use and applications of reinforcement learning (RL) techniques in the production planning and control (PPC) field addressing the following PPC areas: facility resource planning, capacity planning, purchase and supply management, production schedulin...

Descripción completa

Detalles Bibliográficos
Autores: Esteso, Ana|||0000-0003-0379-8786, Peidro Payá, David|||0000-0001-8678-6881, Mula, Josefa|||0000-0002-8447-3387, Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876
Tipo de recurso: artículo
Fecha de publicación:2023
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/196934
Acceso en línea:https://riunet.upv.es/handle/10251/196934
Access Level:acceso abierto
Palabra clave:Artificial intelligence
Machine learning
Reinforcement learning
Deep reinforcement learning
Production planning and control
Industry 4.0
ORGANIZACION DE EMPRESAS
09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación
id ES_f4b0321cb1ec8a633e6f8cfb3daee064
oai_identifier_str oai:riunet.upv.es:10251/196934
network_acronym_str ES
network_name_str España
repository_id_str
dc.title.none.fl_str_mv Reinforcement learning applied to production planning and control
title Reinforcement learning applied to production planning and control
spellingShingle Reinforcement learning applied to production planning and control
Esteso, Ana|||0000-0003-0379-8786
Artificial intelligence
Machine learning
Reinforcement learning
Deep reinforcement learning
Production planning and control
Industry 4.0
ORGANIZACION DE EMPRESAS
09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación
title_short Reinforcement learning applied to production planning and control
title_full Reinforcement learning applied to production planning and control
title_fullStr Reinforcement learning applied to production planning and control
title_full_unstemmed Reinforcement learning applied to production planning and control
title_sort Reinforcement learning applied to production planning and control
dc.creator.none.fl_str_mv Esteso, Ana|||0000-0003-0379-8786
Peidro Payá, David|||0000-0001-8678-6881
Mula, Josefa|||0000-0002-8447-3387
Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876
author Esteso, Ana|||0000-0003-0379-8786
author_facet Esteso, Ana|||0000-0003-0379-8786
Peidro Payá, David|||0000-0001-8678-6881
Mula, Josefa|||0000-0002-8447-3387
Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876
author_role author
author2 Peidro Payá, David|||0000-0001-8678-6881
Mula, Josefa|||0000-0002-8447-3387
Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876
author2_role author
author
author
dc.contributor.none.fl_str_mv Departamento de Organización de Empresas
Centro de Investigación en Gestión e Ingeniería de Producción
Escuela Técnica Superior de Ingeniería Industrial
Escuela Politécnica Superior de Alcoy
Escuela de Doctorado
GENERALITAT VALENCIANA
AGENCIA ESTATAL DE INVESTIGACION
European Regional Development Fund
COMISION DE LAS COMUNIDADES EUROPEA
Repositorio Institucional de la Universitat Politècnica de València Riunet
dc.subject.none.fl_str_mv Artificial intelligence
Machine learning
Reinforcement learning
Deep reinforcement learning
Production planning and control
Industry 4.0
ORGANIZACION DE EMPRESAS
09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación
topic Artificial intelligence
Machine learning
Reinforcement learning
Deep reinforcement learning
Production planning and control
Industry 4.0
ORGANIZACION DE EMPRESAS
09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación
description [EN] The objective of this paper is to examine the use and applications of reinforcement learning (RL) techniques in the production planning and control (PPC) field addressing the following PPC areas: facility resource planning, capacity planning, purchase and supply management, production scheduling and inventory management. The main RL characteristics, such as method, context, states, actions, reward and highlights, were analysed. The considered number of agents, applications and RL software tools, specifically, programming language, platforms, application programming interfaces and RL frameworks, among others, were identified, and 181 articles were sreviewed. The results showed that RL was applied mainly to production scheduling problems, followed by purchase and supply management. The most revised RL algorithms were model-free and single-agent and were applied to simplified PPC environments. Nevertheless, their results seem to be promising compared to traditional mathematical programming and heuristics/metaheuristics solution methods, and even more so when they incorporate uncertainty or non-linear properties. Finally, RL value-based approaches are the most widely used, specifically Q-learning and its variants and for deep RL, deep Q-networks. In recent years however, the most widely used approach has been the actor-critic method, such as the advantage actor critic, proximal policy optimisation, deep deterministic policy gradient and trust region policy optimisation.
publishDate 2023
dc.date.none.fl_str_mv 2023
2023-08-18
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
VoR
http://purl.org/coar/version/c_970fb48d4fbd8a85
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://riunet.upv.es/handle/10251/196934
url https://riunet.upv.es/handle/10251/196934
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.relation.none.fl_str_mv Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 RTI2018-101344-B-I00 OPTIMIZACION DE TECNOLOGIAS DE PRODUCCION CERO-DEFECTOS HABILITADORAS PARA CADENAS DE SUMINISTRO 4.0
Generalitat Valenciana https://doi.org/10.13039/501100003359 PROMETEO%2F2021%2F065 Industrial Production and Logistics Optimization in Industry 4.0 (i4OPT)
Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 RTI2018-102020-B-I00 INTEGRACION DE LA TOMA DE DECISIONES DE LOS NIVELES TACTICO-OPERATIVO PARA LA MEJORA DE LA EFICIENCIA DEL SISTEMA DE PRODUCTIVO EN ENTORNOS INDUSTRIA 4.0
Generalitat Valenciana https://doi.org/10.13039/501100003359 CIGE%2F2021%2F159 Optimización de cadenas de suministro 5.0 resilientes, sostenibles y orientadas a personas mediante inteligencia híbrida
European Commission https://doi.org/10.13039/501100000780 H2020 825631
European Commission https://doi.org/10.13039/501100000780 H2020 958205
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reconocimiento - No comercial - Sin obra derivada (by-nc-nd)
http://creativecommons.org/licenses/by-nc-nd/4.0/
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
Reconocimiento - No comercial - Sin obra derivada (by-nc-nd)
http://creativecommons.org/licenses/by-nc-nd/4.0/
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.publisher.none.fl_str_mv Taylor & Francis
publisher.none.fl_str_mv Taylor & Francis
dc.source.none.fl_str_mv reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
instname:Universitat Politècnica de València (UPV)
instname_str Universitat Politècnica de València (UPV)
reponame_str RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
collection RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869424494882324480
spelling Reinforcement learning applied to production planning and controlEsteso, Ana|||0000-0003-0379-8786Peidro Payá, David|||0000-0001-8678-6881Mula, Josefa|||0000-0002-8447-3387Díaz-Madroñero Boluda, Francisco Manuel|||0000-0003-1693-2876Artificial intelligenceMachine learningReinforcement learningDeep reinforcement learningProduction planning and controlIndustry 4.0ORGANIZACION DE EMPRESAS09.- Desarrollar infraestructuras resilientes, promover la industrialización inclusiva y sostenible, y fomentar la innovación[EN] The objective of this paper is to examine the use and applications of reinforcement learning (RL) techniques in the production planning and control (PPC) field addressing the following PPC areas: facility resource planning, capacity planning, purchase and supply management, production scheduling and inventory management. The main RL characteristics, such as method, context, states, actions, reward and highlights, were analysed. The considered number of agents, applications and RL software tools, specifically, programming language, platforms, application programming interfaces and RL frameworks, among others, were identified, and 181 articles were sreviewed. The results showed that RL was applied mainly to production scheduling problems, followed by purchase and supply management. The most revised RL algorithms were model-free and single-agent and were applied to simplified PPC environments. Nevertheless, their results seem to be promising compared to traditional mathematical programming and heuristics/metaheuristics solution methods, and even more so when they incorporate uncertainty or non-linear properties. Finally, RL value-based approaches are the most widely used, specifically Q-learning and its variants and for deep RL, deep Q-networks. In recent years however, the most widely used approach has been the actor-critic method, such as the advantage actor critic, proximal policy optimisation, deep deterministic policy gradient and trust region policy optimisation.The funding for the research work that has led to the obtained results came from the following grants: CADS4.0 (Ref. RTI2018-101344-B-I00) and NIOTOME (Ref. RTI2018102020-B-I00), financed byMCIN/AEI/10.13039/501100011033 and 'ERDF A way of making DEurope'; the EU H2020 research and innovation programme with grant numbers 825631 'Zero-Defect Manufacturing Platform (ZDMP)' and 958205 'Industrial Data Services for Quality Control in SmartManufacturing (i4Q)'; 'Industrial Production and Logistics Optimization in Industry 4.0' (i4OPT) (Ref. PROMETEO/2021/065) and 'Resilient, Sustainable and PeopleOriented Supply Chain 5.0 Optimization Using Hybrid Intelligence' (RESPECT) (Ref. CIGE/2021/159) Projects were funded by the Generalitat Valenciana (Valencian Regional Government).Taylor & FrancisDepartamento de Organización de EmpresasCentro de Investigación en Gestión e Ingeniería de ProducciónEscuela Técnica Superior de Ingeniería IndustrialEscuela Politécnica Superior de AlcoyEscuela de DoctoradoGENERALITAT VALENCIANAAGENCIA ESTATAL DE INVESTIGACIONEuropean Regional Development FundCOMISION DE LAS COMUNIDADES EUROPEARepositorio Institucional de la Universitat Politècnica de València Riunet20232023-08-18journal articlehttp://purl.org/coar/resource_type/c_6501VoRhttp://purl.org/coar/version/c_970fb48d4fbd8a85info:eu-repo/semantics/articleapplication/pdfhttps://riunet.upv.es/handle/10251/196934reponame:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valénciainstname:Universitat Politècnica de València (UPV)InglésengAgencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 RTI2018-101344-B-I00 OPTIMIZACION DE TECNOLOGIAS DE PRODUCCION CERO-DEFECTOS HABILITADORAS PARA CADENAS DE SUMINISTRO 4.0Generalitat Valenciana https://doi.org/10.13039/501100003359 PROMETEO%2F2021%2F065 Industrial Production and Logistics Optimization in Industry 4.0 (i4OPT)Agencia Estatal de Investigación http://dx.doi.org/10.13039/501100011033 Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020 RTI2018-102020-B-I00 INTEGRACION DE LA TOMA DE DECISIONES DE LOS NIVELES TACTICO-OPERATIVO PARA LA MEJORA DE LA EFICIENCIA DEL SISTEMA DE PRODUCTIVO EN ENTORNOS INDUSTRIA 4.0Generalitat Valenciana https://doi.org/10.13039/501100003359 CIGE%2F2021%2F159 Optimización de cadenas de suministro 5.0 resilientes, sostenibles y orientadas a personas mediante inteligencia híbridaEuropean Commission https://doi.org/10.13039/501100000780 H2020 825631European Commission https://doi.org/10.13039/501100000780 H2020 958205open accesshttp://purl.org/coar/access_right/c_abf2Reconocimiento - No comercial - Sin obra derivada (by-nc-nd) http://creativecommons.org/licenses/by-nc-nd/4.0/info:eu-repo/semantics/openAccessoai:riunet.upv.es:10251/1969342026-06-13T07:49:27Z
score 15,301603