Behavior monitoring under uncertainty using Bayesian surprise and optimal action selection
The increasing trend towards delegating tasks to autonomous artificial agents in safety–critical sociotechnical systems makes monitoring an action selection policy of paramount importance. Agent behavior monitoring may profit from a stochastic specification of an optimal policy under uncertainty. A...
| Autores: | , |
|---|---|
| Tipo de recurso: | artículo |
| Estado: | Versión publicada |
| Fecha de publicación: | 2014 |
| País: | Argentina |
| Institución: | Consejo Nacional de Investigaciones Científicas y Técnicas |
| Repositorio: | CONICET Digital (CONICET) |
| Idioma: | inglés |
| OAI Identifier: | oai:ri.conicet.gov.ar:11336/22441 |
| Acceso en línea: | http://hdl.handle.net/11336/22441 |
| Access Level: | acceso abierto |
| Palabra clave: | Bayesian Surprise Artificial Pancreas Behavior Monitoring Optimal Action Selection Kullback–Leibler Divergence https://purl.org/becyt/ford/2.2 https://purl.org/becyt/ford/2 |
| Sumario: | The increasing trend towards delegating tasks to autonomous artificial agents in safety–critical sociotechnical systems makes monitoring an action selection policy of paramount importance. Agent behavior monitoring may profit from a stochastic specification of an optimal policy under uncertainty. A probabilistic monitoring approach is proposed to assess if an agent behavior (or policy) respects its specification. The desired policy is modeled by a prior distribution for state transitions in an optimally-controlled stochastic process. Bayesian surprise is defined as the Kullback–Leibler divergence between the state transition distribution for the observed behavior and the distribution for optimal action selection. To provide a sensitive on-line estimation of Bayesian surprise with small samples twin Gaussian processes are used. Timely detection of a deviant behavior or anomaly in an artificial pancreas highlights the sensitivity of Bayesian surprise to a meaningful discrepancy regarding the stochastic optimal policy when there exist excessive glycemic variability, sensor errors, controller ill-tuning and infusion pump malfunctioning. To reject outliers and leave out redundant information, on-line sparsification of data streams is proposed. |
|---|