Long integer NTT execution on UPMEM-PIM for 128-bit secure fully homomorphic encryption

Fully Homomorphic Encryption (FHE) enables secure computations on encrypted data, hence becoming an appealing technology for privacy-preserving data processing. A core kernel in many cryptographic and FHE workloads is the Number Theoretic Transform (NTT). While NTT involves frequent non-contiguous d...

Descripción completa

Detalles Bibliográficos
Autores: Barik, Tathagata, Mehta, Priyam|||0000-0001-9201-1994, Pindado, Zaira, Gupta, Harshita|||0009-0009-6704-0473, Kabra, Mayank|||0009-0001-5578-6237, Sadrosadati, Mohammad, Mutlu, Onur, Peña Monferrer, Antonio José
Tipo de recurso: artículo
Fecha de publicación:2026
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:dnet:upcommonspor::7905f6ebc92a474982f9b5cd4d5a3c90
Acceso en línea:https://hdl.handle.net/2117/460634
https://dx.doi.org/10.1016/j.future.2026.108386
Access Level:acceso abierto
Palabra clave:Fully homomorphic encryption
Processing in memory
Number theoretic transform
Àrees temàtiques de la UPC::Informàtica::Seguretat informàtica::Criptografia
Descripción
Sumario:Fully Homomorphic Encryption (FHE) enables secure computations on encrypted data, hence becoming an appealing technology for privacy-preserving data processing. A core kernel in many cryptographic and FHE workloads is the Number Theoretic Transform (NTT). While NTT involves frequent non-contiguous data accesses, limiting overall performance, processing–in–memory (PIM) has the potential to address this limitation. PIM, performing computations close to the data, reduces the need for extensive data transfers between memory and compute units. However, the performance of current PIM solutions is limited by inherent factors related to the integration of processing capabilities within memory modules. In this article we analyze the performance trade-offs of NTT kernel designs along with optimized modular multiplication algorithms on PIM systems based on UPMEM hardware. Our results include significant performance improvements of up to 4.3× over baseline approaches on UPMEM-PIM, while preserving, for the first time in the literature, 128-bit security at high precision.