GBEES-GPU: An efficient parallel GPU algorithm for high-dimensional nonlinear uncertainty propagation
[EN] Eulerian nonlinear uncertainty propagation methods often suffer from finite domain limitations and computational inefficiencies. A recent approach to this class of algorithm, Grid-based Bayesian Estimation Exploiting Sparsity, addresses the first challenge by dynamically allocating a discretize...
| Autores: | , , , |
|---|---|
| Tipo de documento: | artigo |
| Estado: | Versión enviada para evaluación y publicación |
| Data de publicação: | 2025 |
| País: | España |
| Recursos: | Universidad de León |
| Repositório: | BULERIA. Repositorio Institucional de la Universidad de León |
| OAI Identifier: | oai:buleria.unileon.es:10612/25734 |
| Acesso em linha: | https://www.sciencedirect.com/science/article/pii/S0010465525003212?via%3Dihub https://hdl.handle.net/10612/25734 https://doi.org/10.1016/j.cpc.2025.109819 |
| Access Level: | Acceso aberto |
| Palavra-chave: | Aeronáutica CUDA Eulerian uncertainty propagation Corner transport upwind Dynamic gridding 3301 Ingeniería y Tecnología Aeronáuticas |
| Resumo: | [EN] Eulerian nonlinear uncertainty propagation methods often suffer from finite domain limitations and computational inefficiencies. A recent approach to this class of algorithm, Grid-based Bayesian Estimation Exploiting Sparsity, addresses the first challenge by dynamically allocating a discretized grid in regions of phase space where probability is non-negligible. However, the design of the original algorithm causes the second challenge to persist in high-dimensional systems. This paper presents an architectural optimization of the algorithm for CPU implementation, followed by its adaptation to the CUDA framework for single GPU execution. The algorithm is validated for accuracy and convergence, with performance evaluated across distinct GPUs. Tests include propagating a three-dimensional probability distribution subject to the Lorenz ’63 model and a six-dimensional probability distribution subject to the Lorenz ’96 model. The results imply that the improvements made result in a speedup of over 1000 times compared to the original implementation. |
|---|