GPU-Accelerated Sparse FHE: Encrypted DNN Inference on AMD GPUs
arXiv·medium signal
D'Agata et al. adapt the most computationally intensive operation in DNNs — matrix multiplication — for fully homomorphic encrypted execution on AMD GPUs, achieving practical speedups for sparse encrypted inference. As privacy-preserving ML moves toward production, hardware acceleration of FHE operations is the bottleneck. This is the first documented adaptation of sparse FHE matrix operations to AMD's GPU architecture, expanding encrypted inference beyond NVIDIA-only toolchains.