Ethical Coordination of LLM Multi-Agent Systems

[EN] Embedding large language model (LLM) coordinators in production electronic systems, connected vehicles, multi-robot fabrics, IoT control loops, telecommunications orchestration, demands a pre-delivery filter stage that preserves ethical guarantees under adversarial influence at deployment scale...

Descripción completa

Detalles Bibliográficos
Autores: de Curtò-I Díaz, Joaquim, de Zarzà-I Cubero, Irene, Tavares De Araujo Cesariny Calafate, Carlos Miguel|||0000-0001-5729-3041
Tipo de recurso: artículo
Fecha de publicación:2026
País:España
Institución:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:dnet:riunet______::2fdec582f8bb3e6effd5db5cd4b0d635
Acceso en línea:https://riunet.upv.es/handle/10251/236111
Access Level:acceso abierto
Palabra clave:Large language models
Multi-agent systems
Ethical coordination
Trustworthy AI
Agent governance
Agentic AI
Descripción
Sumario:[EN] Embedding large language model (LLM) coordinators in production electronic systems, connected vehicles, multi-robot fabrics, IoT control loops, telecommunications orchestration, demands a pre-delivery filter stage that preserves ethical guarantees under adversarial influence at deployment scale. We present a constitutional governance layer that filters compiled influence policies before they reach a heterogeneous population of grounded LLM agents whose hybrid decision model combines a game-theoretic base probability with an LLM-evaluated narrative shift attenuated by per-agent resistance. Four experiments on a Barabási¿Albert scale-free network of 30 agents powered by Llama-3.3-70B-Instruct show that the filter holds an Ethical Cooperation Score (ECS) of 0.176 (multi-seed mean 0.163, 95% confidence interval (CI) [0.150, 0.174]) against an unconstrained baseline of ECS = 0, enforced by a hard integrity gate (1.000 vs. 0.000). We surface an autonomy paradox in which unconstrained agents resist manipulation more forcefully (0.856 vs. 0.728) yet collapse to ECS = 0, establishing that system-level integrity cannot be delegated to agent-level defence. The advantage is monotonic in resistance (+0.174 to +0.183), seed-stable (Cliff¿s ¿ = 1.0, complete separation), topology- and backbone-invariant across five contemporary LLMs, robust to alternative ECS formulations, and reproduces at N = 100. Against constitutional artificial intelligence (CAI) critique-revise and LlamaGuard-style safety-classifier baselines, the framework matches the integrity floor and adds a measurable margin on the secondary risk surface (BURST timing, composite manipulation risk). The filter runs at 0.78 µs/call (¿ 1.3 × 106 decisions/s/core), supporting always-on deployment as a stateless, modelagnostic component of LLM agent pipelines in adversarially contested electronic systems.