Reinforcement learning with probabilistic boolean network models of smart grid devices

The area of smart power grids needs to constantly improve its efficiency and resilience, to provide high quality electrical power in a resilient grid, while managing faults and avoiding failures. Achieving this requires high component reliability, adequate maintenance, and a studied failure occurren...

ver descrição completa

Detalhes bibliográficos
Autores: Rivera Torres, Pedro Juan, Gershenson García, Carlos, Sánchez Puig, María Fernanda, Kanaan Izquierdo, Samir|||0000-0002-6564-0557
Formato: artículo
Fecha de publicación:2022
País:España
Recursos:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2117/374168
Acesso em linha:https://hdl.handle.net/2117/374168
https://dx.doi.org/10.1155/2022/3652441
Access Level:acceso abierto
Palavra-chave:Smart power grids
Xarxes elèctriques intel·ligents
Àrees temàtiques de la UPC::Matemàtiques i estadística
Descrição
Resumo:The area of smart power grids needs to constantly improve its efficiency and resilience, to provide high quality electrical power in a resilient grid, while managing faults and avoiding failures. Achieving this requires high component reliability, adequate maintenance, and a studied failure occurrence. Correct system operation involves those activities and novel methodologies to detect, classify, and isolate faults and failures and model and simulate processes with predictive algorithms and analytics (using data analysis and asset condition to plan and perform activities). In this paper, we showcase the application of a complex-adaptive, self-organizing modeling method, and Probabilistic Boolean Networks (PBNs), as a way towards the understanding of the dynamics of smart grid devices, and to model and characterize their behavior. This work demonstrates that PBNs are equivalent to the standard Reinforcement Learning Cycle, in which the agent/model has an interaction with its environment and receives feedback from it in the form of a reward signal. Different reward structures were created to characterize preferred behavior. This information can be used to guide the PBN to avoid fault conditions and failures.