Skip to main content Skip to main navigation

Publikation

Exploiting Sum-Product Networks to Offload Database Query Cardinality Estimation to FPGA-Based Smart Storage Devices

Lukas Weber; Yannick Lavan; Johannes Wehrstein; Torben Kalkhof; Carsten Heinz; Carsten Binnig; Andreas Koch
In: Gianluca Leone; Andrés Otero; Paola Busia; Paolo Meloni (Hrsg.). Applied Reconfigurable Computing. Architectures, Tools, and Applications - 22nd International Symposium, ARC 2026, Cagliari, Sardinia, Italy, April 8-10, 2026, Proceedings. International Symposium on Applied Reconfigurable Computing (ARC), Pages 275-291, Lecture Notes in Computer Science, Vol. 16514, Springer, 2026.

Zusammenfassung

Cardinality estimation is crucial for optimizing query perfor- mance in database systems. This study explores the application of Sum- Product Networks for estimating the cardinalities of database queries across various architectures. We analyze the capability of SPNs to han- dle different query types, highlighting their strengths and limitations. Building upon existing research, we have developed a framework that creates both offload and smart storage hardware accelerators tailored for cardinality estimation. Compared to prior work, these accelerators utilize simplified fixed-point arithmetic to enhance resource efficiency. Our framework enables the generation of variants optimized for latency and throughput adaptable to diverse system architectures. Our approach now supports marginal and range-based queries essential for accurate cardinality estimation by extending prior functionality. We applied our framework to generate accelerators and integrated them into two distinct architectures: a PCIe-based accelerator card for offload- ing tasks in large-scale general-purpose databases and a Near-Data Pro- cessing system within the COSMOS+ OpenSSD smart storage SSD. Our evaluation indicates that the PCIe-based architecture achieves over 165 million inferences per second, with latencies as low as 6.62 microseconds, and the Smart Storage device achieves latencies under 2 microseconds. Additionally, we analyzed the impact of our simplified fixed-point num- ber system on resource efficiency and error margins, enabling a different trade-off between the different resources in a typical FPGA to make the existing framework more adaptable to different environments and appli- cations.

Weitere Links