Automated ethical design of multi-agent reinforcement learning environments

This paper introduces the Approximate Multi-Agent Ethical Embedding Process, an algorithm to ethically design reinforcement learning environments where agents learn behaviours aligned with a moral value, while pursuing their own goals. Building on Multi-Objective and Deep Reinforcement Learning, it...

Descripción completa

Detalles Bibliográficos
Autores: Mayoral-Macau, Arnau, Rodriguez-Soto, Manel, Marchesini, Enrico, Lopez-Sanchez, Maite, Sánchez-Fibla, Martí, Farinelli, Alessandro, Rodriguez-Aguilar, Juan Antonio
Tipo de recurso: artículo
Estado:Versión publicada
Fecha de publicación:2025
País:España
Institución:Universitat Pompeu Fabra
Repositorio:Repositorio Digital de la UPF
OAI Identifier:oai:dnet:rdupf_______::b3d0de3f397dfd3193c8d38517418114
Acceso en línea:https://hdl.handle.net/10230/73324
http://dx.doi.org/10.3233/FAIA250587
Access Level:acceso abierto
Palabra clave:Multi-agent reinforcement learning
Value-alignment
Multi-objective
Reinforcement learning
Descripción
Sumario:This paper introduces the Approximate Multi-Agent Ethical Embedding Process, an algorithm to ethically design reinforcement learning environments where agents learn behaviours aligned with a moral value, while pursuing their own goals. Building on Multi-Objective and Deep Reinforcement Learning, it extends a previously theory-driven method limited to small-scale problems. The new approach is tested in a scaled-up, ethically augmented version of the gathering game, demonstrating its effectiveness in managing increased complexity.