REVIEW 3 major objections 5 minor 76 references
FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read A 2.4%-area GPU unit that natively performs modulo-linear transformations could halve FHE bootstrapping latency and roughly double end-to-end encrypted workload speed.
desk verdict A plausible architecture with a credible area model, but the evaluation builds on a 32-bit datapath that doesn't fit the paper's own 60-bit CKKS moduli. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
FHECore is a 2D systolic array organized as a 16x8 grid of processing elements, each computing a·b mod q with 32-bit operands; a six-stage pipeline and Barrett reduction module sit in every PE. The array uses output-stationary dataflow, completing a 16x8x16 modulo matrix multiplication in 44 cycles, and shares register-file ports with the existing tensor-core units. The load-bearing idea is the modulo-linear formulation: the NTT becomes a product of twiddle-factor matrices and base conversion becomes a mixed-moduli matrix-matrix product, so one instruction (FHEC.16816.S32) can replace the decompose–reassemble–reduce sequence that tensor-core approaches require.
What would settle it
Run a cycle-accurate trace of the bootstrap kernel using the stated CKKS parameters (logQP=1743) and count how many FHEC instructions are actually required when the 32-bit datapath is used. If any operand needs limb decomposition or extra reduction steps, the dynamic instruction count will not shrink by 2.41x/1.96x and the 50% bootstrapping speedup will not reproduce.
Extended reading notes
Core claim
The central claim is that the CKKS bottleneck on GPUs is a datatype mismatch, not parallelism: NTT and base conversion together account for more than 70% of runtime and are currently implemented by splitting wide integers into INT8 chunks for tensor-core matrix multiply, then reassembling and reducing them. The paper shows both kernels can be formulated as modulo matrix multiplications, so a 16x8 systolic array of processing elements that each compute a·b mod q on 32-bit operands, with Barrett reduction embedded, can execute them as a single instruction. In the paper's trace-driven simulation, this new instruction—issued from the same register-file ports as tensor-core operations—cuts dynami
Load-bearing premise
The design assumes every residue value in the evaluated CKKS workloads fits in the 32-bit modulo multiply-accumulate datapath, so no limb decomposition or intermediate reduction is needed; if the actual moduli are 55–60 bits, the single-instruction mapping collapses.
Editorial extensions
If this is right
- Bootstrapping latency drops by roughly half in simulation (from 314.67 ms to 163.90 ms), which directly extends the usable depth of CKKS computations.
- A single programmer-level intrinsic (fhe_sync) that mirrors existing matrix-multiply intrinsics lets GPU FHE libraries adopt the unit without a new compiler or ISA.
- Because the same two kernels dominate BFV and BGV, and TFHE can be expressed with NTTs, the unit would likely accelerate other FHE schemes, not just CKKS.
- The 2.4% area overhead (and an estimated 1.5% on newer GPUs) keeps the design within reticle limits, making it a physically feasible drop-in SM extension.
- Non-FHE GPU workloads are unaffected because FHECore shares register ports with tensor cores and only activates when FHEC instructions are issued; the paper argues FHE and plaintext ML workloads are disjoint in practice.
Reading between the lines
- The CKKS parameters in the paper (logQP=1743, L=26) imply RNS moduli of about 55–60 bits, not 32 bits; the paper never states that its workloads use 32-bit moduli. If the modulo multiply-accumulate unit must handle larger residues via limb decomposition, the single-instruction mapping and the reported 2.41x/1.96x instruction-count reductions are unlikely to hold as reported.
- The evaluation inserts FHEC instructions into traces manually because the compiler backend is closed; the real-world speedup depends on the compiler generating these instructions without extra register pressure or scheduling overhead, which the paper does not demonstrate.
- A concrete next step would be to build a small FPGA prototype of the 16x8 FHECore and run NTT and base conversion with both 32-bit and 55-bit moduli, measuring whether the 44-cycle latency and instruction-count compression persist. This would separate the benefit of native modulo arithmetic from the benefit of the specific operand width.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. FHECore proposes a specialized systolic-array functional unit, integrated into each Streaming Multiprocessor of an NVIDIA A100, to accelerate CKKS FHE workloads. The paper's key insight is that Number Theoretic Transforms and base conversion are modulo-linear transformations, allowing both to run on a common modulo-multiply-accumulate hardware unit. The authors add a new FHEC instruction, evaluate it by editing SASS traces and replaying them in Accel-Sim, and report geometric-mean dynamic instruction-count reductions of 2.41x for CKKS primitives and 1.96x for end-to-end workloads, translating to speedups of 1.57x and 2.12x, respectively, with a 2.4% area overhead estimated from RTL synthesis.
Significance. If the central claim were valid, FHECore would be a notable contribution: it would show that a small, area-efficient addition to commodity GPUs could roughly double FHE throughput while preserving the GPU programming model. The paper has genuine strengths: the RTL synthesis flow with ASAP7 and SiliconCompiler is a concrete, reproducible area estimate; the comparison to GME's 26.6% area overhead is informative; and the mapping of NTT/base conversion to a common modulo-linear formulation is conceptually appealing. However, the evaluation rests on two load-bearing assumptions that are not supported by the manuscript: the 32-bit datapath described in §IV-C cannot handle the ~55-60-bit RNS moduli implied by the paper's own CKKS parameters (Table V), and the performance results come from manually edited traces with an assumed 44-cycle latency rather than a validated integration model. These issues undermine the central speedup and instruction-count claims.
major comments (3)
- [§IV-C, §V-A, Table V] The PE is specified as computing 'a·b mod q over 32-bit operands' (§IV-C), and §V-A asserts that the NTT can run 'without any decomposition or intermediate reduction steps.' However, the benchmarks in Table V use logQP=1743, L=26, and dnum=3, which for standard CKKS-RNS implies individual moduli of roughly 55-60 bits. A 32-bit multiplier cannot even form the product of two such residues, let alone reduce modulo a 60-bit prime in one step. The manuscript nowhere states that the workloads use 32-bit moduli nor describes multi-limb handling. Consequently, the one-FHEC-instruction mapping to 16 or 64 INT8 Tensor Core instructions collapses for the stated parameters, invalidating the instruction-count and speedup numbers.
- [§VI-A] The performance evaluation is based on manually inserting FHEC instructions into NVBit-collected traces and assigning them a 44-cycle latency, down from the 64 cycles used for Tensor Core instructions. This is not a faithful simulation of integrating a new functional unit: there is no validation of the register-file port sharing, warp-scheduler interaction, or issue constraints, and the latency reduction is an input assumption rather than a measured or modeled property of the complete SM. The resulting speedups are therefore partly definitional. The paper should at least provide a sensitivity analysis over FHEC latency and a detailed account of how FHEC operations interact with the existing pipeline before claiming end-to-end speedups.
- [§IV-B, §IV-F, Fig. 7] The paper claims both that 'no modifications are required in either the compiler stack or the instruction stream sequence' (§IV-B) and that the closed-source nvcc backend prevents direct SASS insertion, requiring manual trace editing (§IV-F). More importantly, the occupancy and IPC improvements in Figure 7 are reported without any microarchitectural mechanism beyond replacing instructions with shorter-latency FHEC ops. Since non-FHE workloads are never evaluated, the repeated claim that FHECore 'does not compromise general-purpose GPU performance' is unsupported. The authors should either remove that claim or provide evidence, such as running standard GPU benchmarks with the FHECore unit present but idle.
minor comments (5)
- [Abstract] Typo: 'inuring' should be 'incurring'.
- [Table VI and Abstract/Conclusion] The reported geometric-mean instruction-count reduction for end-to-end workloads is 1.96x, but computing the geometric mean from Table VI (Bootstrap 2.12x, LR 2.68x, ResNet 1.89x, BERT-Tiny 1.71x) gives approximately 2.07x. Please reconcile the numbers.
- [§IV-C] The notation 'a·b mod q over 32-bit operands' is ambiguous: specify whether q is a 32-bit modulus, whether the product is truncated before reduction, and whether signed or unsigned operands are supported. This matters for the Barrett reduction implementation.
- [§IV-F] The proposed PTX instruction format and the SASS 'FHEC.16816' instruction are described only at a high level. Clarify how the 16x8x16 tile maps to 32-bit operands in the register file, especially since WMMA fragments for INT8 have a different layout.
- [Table VII] Comparing latencies across different GPUs (RTX 4090, A100) and different libraries is informative but should be accompanied by a statement about architectural differences and clock speeds to avoid implying a head-to-head comparison.
Circularity Check
No significant circularity: the modulo-linear-transform formulation is self-contained, and the performance evaluation is an openly modeled simulation rather than a constructional equivalence.
full rationale
The core design argument—that NTT and base conversion can be expressed as modulo-linear transformations—is established by the paper's own equations (Eqs. 1–3) and does not depend on a self-citation or on the FHECore results. The hardware mapping is also independently grounded: the PE datapath and Barrett reduction are described in §IV-C, and the area/latency numbers come from RTL synthesis with ASAP7 (§VI-D). The main performance evaluation is a trace simulation in which FHEC instructions are manually inserted and assigned a 44-cycle latency (§VI-A). This is an acknowledged modeling limitation—not a circular derivation—because the instruction-count and speedup numbers are presented as consequences of the assumed ISA semantics, not as an independent empirical fit. The equivalence 'one FHECoreMMM invocation corresponds to 16 or 64 TensorCoreGEMM calls' (§V-A) is a design assumption whose correctness matters for the validity of the speedup estimates, but it is not a case of the paper defining the result in terms of the input or fitting a parameter and then renaming it a prediction. Self-citations such as GME [68] and FIDESlib [5] are used for context and baselines, not as the load-bearing justification for the central claim. The potential mismatch between the 32-bit PE and the 60-bit CKKS moduli (Table V) is a correctness risk, not a circularity. Overall, the derivation chain is not circular; the main caveat is that the simulated performance rests on the manual instruction-substitution methodology, which the paper discloses.
Assumptions & free parameters
free parameters (3)
- FHEC instruction latency in Accel-Sim =
44 cycles
- Systolic array dimensions =
16×8
- PE pipeline depth =
6 cycles
assumptions (4)
- domain assumption CKKS RNS residues fit in 32-bit integers.
- standard math Barrett reduction with precomputed constant mu correctly computes a·b mod q for the moduli used.
- domain assumption Accel-Sim with SPECIALIZED_UNIT_3_OP and manually edited traces faithfully models FHECore integration.
- standard math The 4-step NTT decomposition can be expressed as matrix multiplications with twiddle factors.
invented entities (2)
-
FHECore functional unit
-
FHEC instruction
Cite this review
Pith. "Pith review of FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption." pith.science (2026). https://pith.science/paper/Y5FIEDXW
@misc{pith2026260222229,
author = {Pith},
title = {Pith review of: FHECore: Rethinking GPU Microarchitecture for Fully Homomorphic Encryption},
year = {2026},
howpublished = {\url{https://pith.science/paper/Y5FIEDXW}},
note = {Machine review of arXiv:2602.22229}
}
abstract
Fully Homomorphic Encryption (FHE) enables computation directly on encrypted data but incurs massive computational and memory overheads, often exceeding plaintext execution by several orders of magnitude. While custom ASIC accelerators can mitigate these costs, their long time-to-market and the rapid evolution of FHE algorithms threaten their long-term relevance. GPUs, by contrast, offer scalability, programmability, and widespread availability, making them an attractive platform for FHE. However, modern GPUs are increasingly specialized for machine learning workloads, emphasizing low-precision datatypes (e.g., INT$8$, FP$8$) that are fundamentally mismatched to the wide-precision modulo arithmetic required by FHE. Essentially, while GPUs offer ample parallelism, their functional units, like Tensor Cores, are not suited for wide-integer modulo arithmetic required by FHE schemes such as CKKS. Despite this constraint, researchers have attempted to map FHE primitives on Tensor Cores by segmenting wide integers into low-precision (INT$8$) chunks. To overcome these bottlenecks, we propose FHECore, a specialized functional unit integrated directly into the GPU's Streaming Multiprocessor. Our design is motivated by a key insight: the two dominant contributors to latency$-$Number Theoretic Transform and Base Conversion$-$can be formulated as modulo-linear transformations. This allows them to be mapped on a common hardware unit that natively supports wide-precision modulo-multiply-accumulate operations. Our simulations demonstrate that FHECore reduces dynamic instruction count by a geometric mean of $2.41\times$ for CKKS primitives and $1.96\times$ for end-to-end workloads. These reductions translate to performance speedups of $1.57\times$ and $2.12\times$, respectively$-$including a $50\%$ reduction in bootstrapping latency$-$all while inuring a modest $2.4\%$ area overhead.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Heap: A fully homomorphic encryption accelerator with parallelized bootstrapping,
R. Agrawal, A. Chandrakasan, and A. Joshi, “Heap: A fully homomorphic encryption accelerator with parallelized bootstrapping,” in Proceedings of the 51st Annual International Symposium on Computer Architecture, ser. ISCA ’24. IEEE Press, 2025, p. 756–769. [Online]. Available: https://doi.org/10.1109/ISCA59077.2024.00060
arXiv 2025
-
[2]
Mad: Memory-aware design techniques for accelerating fully homomorphic encryption,
R. Agrawal, L. De Castro, C. Juvekar, A. Chandrakasan, V . Vaikuntanathan, and A. Joshi, “Mad: Memory-aware design techniques for accelerating fully homomorphic encryption,” in Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 685–697. [On...
arXiv 2023
-
[3]
Fab: An fpga-based accelerator for bootstrappable fully homomorphic encryption,
R. Agrawal, L. de Castro, G. Yang, C. Juvekar, R. Yazicigil, A. Chandrakasan, V . Vaikuntanathan, and A. Joshi, “Fab: An fpga-based accelerator for bootstrappable fully homomorphic encryption,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2023, pp. 882–895
2023
-
[4]
Agrawal and A
R. Agrawal and A. Joshi,On Architecting Fully Homomorphic Encryption-based Computing Systems / by Rashmi Agrawal, Ajay Joshi., 1st ed., ser. Synthesis Lectures on Computer Architecture. Cham: Springer International Publishing, 2023
2023
-
[5]
Fideslib: A fully-fledged open-source fhe library for efficient ckks on gpus,
C. Agull ´o-Domingo, O. Vera-L ´opez, S. Guzelhan, L. Daksha, A. E. Jerari, K. Shivdikar, R. Agrawal, D. Kaeli, A. Joshi, and J. L. Abell ´an, “Fideslib: A fully-fledged open-source fhe library for efficient ckks on gpus,” in2025 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2025, pp. 1–3
2025
-
[6]
REED: Chiplet-based accelerator for fully homomorphic encryption,
A. Aikata, A. C. Mert, S. Kwon, M. Deryabin, and S. S. Roy, “REED: Chiplet-based accelerator for fully homomorphic encryption,” Cryptology ePrint Archive, Paper 2023/1190, 2023. [Online]. Available: https://eprint.iacr.org/2023/1190
2023
-
[7]
Openfhe: Open-source fully homomorphic encryption library,
A. Al Badawi, J. Bates, F. Bergamaschi, D. B. Cousins, S. Erabelli, N. Genise, S. Halevi, H. Hunt, A. Kim, Y . Lee, Z. Liu, D. Micciancio, I. Quah, Y . Polyakov, S. R.V ., K. Rohloff, J. Saylor, D. Suponitsky, M. Triplett, V . Vaikuntanathan, and V . Zucca, “Openfhe: Open-source fully homomorphic encryption library,” inProceedings of the 10th Workshop on ...
arXiv 2022
-
[8]
Multi-gpu design and performance evaluation of homomorphic encryption on gpu clusters,
A. Al Badawi, B. Veeravalli, J. Lin, N. Xiao, M. Kazuaki, and A. Khin Mi Mi, “Multi-gpu design and performance evaluation of homomorphic encryption on gpu clusters,”IEEE Transactions on Parallel and Distributed Systems, vol. 32, no. 2, pp. 379–391, 2021
2021
Show all 76 references
-
[9]
Homomorphic encryption standard,
M. Albrecht, M. Chase, H. Chen, J. Ding, S. Goldwasser, S. Gorbunov, S. Halevi, J. Hoffstein, K. Laine, K. Lauter, S. Lokam, D. Micciancio, D. Moody, T. Morrison, A. Sahai, and V . Vaikuntanathan, “Homomorphic encryption standard,” Cryptology ePrint Archive, Paper 2019/939, 20...
2019
-
[10]
Ffts in external of hierarchical memory,
D. H. Bailey, “Ffts in external of hierarchical memory,” inProceedings of the 1989 ACM/IEEE Conference on Supercomputing, ser. Supercomputing ’89. New York, NY , USA: Association for Computing Machinery, 1989, p. 234–242. [Online]. Available: https://doi.org/10.1145/76263.76288
1989
-
[11]
Sapphire: A configurable crypto-processor for post-quantum lattice-based protocols (extended version),
U. Banerjee, T. S. Ukyab, and A. P. Chandrakasan, “Sapphire: A configurable crypto-processor for post-quantum lattice-based protocols (extended version),” Cryptology ePrint Archive, Paper 2019/1140, 2019. [Online]. Available: https://eprint.iacr.org/2019/1140
2019
-
[12]
He3db: An efficient and elastic encrypted database via arithmetic-and-logic fully homomorphic encryption,
S. Bian, Z. Zhang, H. Pan, R. Mao, Z. Zhao, Y . Jin, and Z. Guan, “He3db: An efficient and elastic encrypted database via arithmetic-and-logic fully homomorphic encryption,” inProceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’23. ...
2023
-
[13]
Fully homomorphic encryption without bootstrapping,
Z. Brakerski, C. Gentry, and V . Vaikuntanathan, “Fully homomorphic encryption without bootstrapping,” Cryptology ePrint Archive, Paper 2011/277, 2011. [Online]. Available: https://eprint.iacr.org/2011/277
2011
-
[14]
Secure image processing using lwe based homomorphic encryption,
R. Challa, G. VijayaKumari, and S. B, “Secure image processing using lwe based homomorphic encryption,” in2015 IEEE International Conference on Electrical, Computer and Communication Technologies (ICECCT), 2015, pp. 1–6
2015
-
[15]
Authenticable data analytics over encrypted data in the cloud,
L. Chen, Y . Mu, L. Zeng, F. Rezaeibagha, and R. H. Deng, “Authenticable data analytics over encrypted data in the cloud,”IEEE Transactions on Information Forensics and Security, vol. 18, pp. 1800–1813, 2023
2023
-
[16]
Eyeriss: a spatial architecture for energy-efficient dataflow for convolutional neural networks,
Y .-H. Chen, J. Emer, and V . Sze, “Eyeriss: a spatial architecture for energy-efficient dataflow for convolutional neural networks,” in Proceedings of the 43rd International Symposium on Computer Architecture, ser. ISCA ’16. IEEE Press, 2016, p. 367–379. [Online]. Available: ...
2016 doi
-
[17]
Bootstrapping for approximate homomorphic encryption,
J. H. Cheon, K. Han, A. Kim, M. Kim, and Y . Song, “Bootstrapping for approximate homomorphic encryption,” inAdvances in Cryptology– EUROCRYPT 2018: 37th Annual International Conference on the Theory and Applications of Cryptographic Techniques, Tel Aviv, Israel, April 29-May ...
2018
-
[18]
Homomorphic encryption for arithmetic of approximate numbers,
J. H. Cheon, A. Kim, M. Kim, and Y . Song, “Homomorphic encryption for arithmetic of approximate numbers,” Cryptology ePrint Archive, Paper 2016/421, 2016. [Online]. Available: https://eprint.iacr.org/2016/421
2016
-
[19]
TFHE: Fast fully homomorphic encryption over the torus,
I. Chillotti, N. Gama, M. Georgieva, and M. Izabach `ene, “TFHE: Fast fully homomorphic encryption over the torus,” Cryptology ePrint Archive, Paper 2018/421, 2018. [Online]. Available: https://eprint.iacr.org/2018/421
2018
-
[20]
Cheddar: A swift fully homomorphic encryption library designed for gpu architectures,
W. Choi, J. Kim, and J. H. Ahn, “Cheddar: A swift fully homomorphic encryption library designed for gpu architectures,” inProceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (ASPLOS ’26). Pitts...
-
[21]
A computational model for tensor core units,
R. Chowdhury, F. Silvestri, and F. Vella, “A computational model for tensor core units,” inProceedings of the 32nd ACM Symposium on Parallelism in Algorithms and Architectures, ser. SPAA ’20. New York, NY , USA: Association for Computing Machinery, 2020, p. 519–521. [Online]. ...
2020
-
[22]
cuHE: A homomorphic encryption accelerator library,
W. Dai and B. Sunar, “cuHE: A homomorphic encryption accelerator library,” Cryptology ePrint Archive, Paper 2015/818, 2015. [Online]. Available: https://eprint.iacr.org/2015/818
2015
-
[23]
Accelerating reduction and scan using tensor core units,
A. Dakkak, C. Li, J. Xiong, I. Gelado, and W.-m. Hwu, “Accelerating reduction and scan using tensor core units,” inProceedings of the ACM International Conference on Supercomputing, 2019, pp. 46–57
2019
-
[24]
Toward a universal cryptographic accelerator,
S. Devadas and D. Sanchez, “Toward a universal cryptographic accelerator,”Computer, vol. 58, no. 1, pp. 105–108, 2025
2025
-
[25]
Leveraging the error resilience of machine-learning applications for designing highly energy efficient accelerators,
Z. Du, K. V . Palem, L. Avinash, O. Temam, Y . Chen, and C. Wu, “Leveraging the error resilience of machine-learning applications for designing highly energy efficient accelerators,”2014 19th Asia and South Pacific Design Automation Conference (ASP-DAC), pp. 201–206, 2014. [On...
2014
-
[26]
Osiris: A systolic approach to accelerating fully homomorphic encryption,
A. Ebel and B. Reagen, “Osiris: A systolic approach to accelerating fully homomorphic encryption,” 2024. [Online]. Available: https://arxiv.org/abs/2408.09593
2024 arXiv
-
[27]
Warpdrive: Gpu-based fully homomorphic encryption acceleration leveraging tensor and cuda cores,
G. Fan, M. Zhang, F. Zheng, S. Fan, T. Zhou, X. Deng, W. Tang, L. Kong, Y . Song, and S. Yan, “Warpdrive: Gpu-based fully homomorphic encryption acceleration leveraging tensor and cuda cores,” in2025 IEEE International Symposium on High Performance Computer Architecture (HPCA)...
2025
-
[28]
Somewhat practical fully homomorphic encryption,
J. Fan and F. Vercauteren, “Somewhat practical fully homomorphic encryption,” Cryptology ePrint Archive, Paper 2012/144, 2012. [Online]. Available: https://eprint.iacr.org/2012/144
2012
-
[29]
TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU ,
S. Fan, Z. Wang, W. Xu, R. Hou, D. Meng, and M. Zhang, “ TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU ,” in2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Los Alamitos, CA, USA: IEEE Computer Society, Mar. 2023, p...
2023
-
[30]
A configurable cloud-scale dnn processor for real-time ai,
J. Fowers, K. Ovtcharov, M. Papamichael, T. Massengill, M. Liu, D. Lo, S. Alkalay, M. Haselman, L. Adams, M. Ghandi, S. Heil, P. Patel, A. Sapek, G. Weisz, L. Woods, S. Lanka, S. K. Reinhardt, A. M. Caulfield, E. S. Chung, and D. Burger, “A configurable cloud-scale dnn process...
2018
-
[31]
Memfhe: End-to-end computing with fully homomorphic encryption in memory,
S. Gupta, R. Cammarota, and T. ˇSimuni´c, “Memfhe: End-to-end computing with fully homomorphic encryption in memory,”ACM Trans. Embed. Comput. Syst., vol. 23, no. 2, Mar. 2024. [Online]. Available: https://doi.org/10.1145/3569955
2024 doi
-
[32]
The big chip: Challenge, model and architecture,
Y . Han, H. Xu, M. Lu, H. Wang, J. Huang, Y . Wang, Y . Wang, F. Min, Q. Liu, M. Liu, and N. Sun, “The big chip: Challenge, model and architecture,”Fundamental Research, vol. 4, no. 6, pp. 1431–1441, 2024. [Online]. Available: https://www.sciencedirect.com/science/article/pii/...
2024
-
[33]
One server for the price of two: simple and fast single-server private information retrieval,
A. Henzinger, M. M. Hong, H. Corrigan-Gibbs, S. Meiklejohn, and V . Vaikuntanathan, “One server for the price of two: simple and fast single-server private information retrieval,” inProceedings of the 32nd USENIX Conference on Security Symposium, ser. SEC ’23. USA: USENIX Asso...
2023
-
[34]
Optimizing homomorphic encryption based secure image analytics,
N. Jain, K. Nandakumar, N. Ratha, S. Pankanti, and U. Kumar, “Optimizing homomorphic encryption based secure image analytics,” in2021 IEEE 23rd International Workshop on Multimedia Signal Processing (MMSP), 2021, pp. 1–6
2021
-
[35]
Cinnamon: A framework for scale-out encrypted ai,
S. Jayashankar, E. Chen, T. Tang, W. Zheng, and D. Skarlatos, “Cinnamon: A framework for scale-out encrypted ai,” inProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ser. ASPLOS ’25. New Yor...
2025
-
[36]
Secure outsourced matrix computation and application to neural networks,
X. Jiang, M. Kim, K. Lauter, and Y . Song, “Secure outsourced matrix computation and application to neural networks,” inProceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. Association for Computing Machinery, 2018, p. 1209–1222. [Online]. Ava...
2018
-
[37]
Neo: Towards efficient fully homomorphic encryption acceleration using tensor core,
D. Jiao, X. Deng, Z. Wang, S. Fan, Y . Chen, D. Meng, R. Hou, and M. Zhang, “Neo: Towards efficient fully homomorphic encryption acceleration using tensor core,” inProceedings of the 52nd Annual International Symposium on Computer Architecture, ser. ISCA ’25. New York, NY , US...
2025
-
[38]
Tinybert: Distilling bert for natural language understanding,
X. Jiao, Y . Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu, “Tinybert: Distilling bert for natural language understanding,” arXiv preprint arXiv:1909.10351, 2019
1909 arXiv
-
[39]
Over 100x faster bootstrapping in fully homomorphic encryption through memory-centric optimization with gpus,
W. Jung, S. Kim, J. H. Ahn, J. H. Cheon, and Y . Lee, “Over 100x faster bootstrapping in fully homomorphic encryption through memory-centric optimization with gpus,”IACR Transactions on Cryptographic Hardware and Embedded Systems, vol. 2021, no. 4, p. 114–148, Aug. 2021. [Onli...
2021
-
[40]
Accel-sim: an extensible simulation framework for validated gpu modeling,
M. Khairy, Z. Shen, T. M. Aamodt, and T. G. Rogers, “Accel-sim: an extensible simulation framework for validated gpu modeling,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture, ser. ISCA ’20. IEEE Press, 2020, p. 473–486. [Online]. A...
2020
-
[41]
Sharp: A short-word hierarchical accelerator for robust and practical fully homomorphic encryption,
J. Kim, S. Kim, J. Choi, J. Park, D. Kim, and J. H. Ahn, “Sharp: A short-word hierarchical accelerator for robust and practical fully homomorphic encryption,” inProceedings of the 50th Annual International Symposium on Computer Architecture, ser. ISCA ’23. New York, NY , USA: ...
2023
-
[42]
Ark: Fully homomorphic encryption accelerator with runtime data generation and inter-operation key reuse,
J. Kim, G. Lee, S. Kim, G. Sohn, M. Rhu, J. Kim, and J. H. Ahn, “Ark: Fully homomorphic encryption accelerator with runtime data generation and inter-operation key reuse,” in2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2022, pp. 1237–1254
2022
-
[43]
Anaheim: Architecture and algorithms for processing fully homomorphic encryption in memory,
J. Kim, S. Yun, H. Ji, W. Choi, S. Kim, and J. H. Ahn, “Anaheim: Architecture and algorithms for processing fully homomorphic encryption in memory,”2025 IEEE International Symposium on High Performance Computer Architecture (HPCA), pp. 1158–1173, 2025. [Online]. Available: htt...
2025
-
[44]
Accelerating number theoretic transformations for bootstrappable homomorphic encryption on gpus,
S. Kim, W. Jung, J. Park, and J. H. Ahn, “Accelerating number theoretic transformations for bootstrappable homomorphic encryption on gpus,” in2020 IEEE International Symposium on Workload Characterization (IISWC), 2020, pp. 264–275
2020
-
[45]
Cifher: A chiplet-based fhe accelerator with a resizable structure,
S. Kim, J. Kim, J. Choi, and J. H. Ahn, “Cifher: A chiplet-based fhe accelerator with a resizable structure,” in2024 International Symposium on Secure and Private Execution Environment Design (SEED), 2024, pp. 119–130
2024
-
[46]
Bts: an accelerator for bootstrappable fully homomorphic encryption,
S. Kim, J. Kim, M. J. Kim, W. Jung, J. Kim, M. Rhu, and J. H. Ahn, “Bts: an accelerator for bootstrappable fully homomorphic encryption,” inProceedings of the 49th Annual International Symposium on Computer Architecture, ser. ISCA ’22. New York, NY , USA: Association for Compu...
2022
-
[47]
Mnist handwritten digit database,
Y . LeCun, C. Cortes, and C. Burges, “Mnist handwritten digit database,” ATT Labs [Online]. Available: http://yann.lecun.com/exdb/mnist, vol. 2, 2010
2010
-
[48]
Modular multiplication without trial division,
P. L. Montgomery, “Modular multiplication without trial division,” in Proceedings of the 5th Symposium on the Arithmetic of Finite Fields, ser. LNCS, vol. 72. Springer, 1985, pp. 261–271
1985
-
[49]
Onionpir: Response efficient single-server pir,
M. H. Mughees, H. Chen, and L. Ren, “Onionpir: Response efficient single-server pir,” inProceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’21. New York, NY , USA: Association for Computing Machinery, 2021, p. 2292–2306. [Online]. A...
2021
-
[50]
NuFHE: A gpu implementation of fully homomorphic encryption on the torus,
NuCypher, “NuFHE: A gpu implementation of fully homomorphic encryption on the torus,” https://github.com/nucypher/nufhe, 2019, version 0.0.3, accessed: 2025-11-16
2019
-
[51]
Nvidia nvlink high-speed interconnect: Application performance brief,
NVIDIA Corporation, “Nvidia nvlink high-speed interconnect: Application performance brief,” https://info.nvidianews.com/rs/nvidia/ images/NVIDIA%20NVLink%20High-Speed%20Interconnect% 20Application%20Performance%20Brief.pdf, NVIDIA, Tech. Rep., 2017, accessed: 2025-07-17
2017
-
[52]
Nvidia tesla v100 gpu architecture whitepaper,
——, “Nvidia tesla v100 gpu architecture whitepaper,” NVIDIA Corpo- ration, Santa Clara, CA, USA, White Paper WP-08608-001 v1.1, Dec. 2017, also published as “Tesla V100 GPU Architecture: The World’s Most Advanced Data Center GPU”. [Online]. Available: https://images.nvidia. co...
2017
-
[53]
(2020) Nvidia a100 tensor core gpu architecture
——. (2020) Nvidia a100 tensor core gpu architecture. https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia- ampere-architecture-whitepaper.pdf
2020
-
[54]
(2022) Nvidia hopper architecture in-depth
——. (2022) Nvidia hopper architecture in-depth. https://developer.nvidia. com/blog/nvidia-hopper-architecture-in-depth/. Accessed: 2025-07-17
2022
-
[55]
(2024) Nvidia blackwell architecture
——. (2024) Nvidia blackwell architecture. https://resources.nvidia.com/ en-us-blackwell-architecture
2024
-
[56]
——,PTX ISA Version 8.8, March 2024, https://docs.nvidia.com/cuda/ pdf/ptx isa 8.8.pdf
2024
-
[57]
Tensor cores: Versatility for HPC & AI,
——, “Tensor cores: Versatility for HPC & AI,” https: //www.nvidia.com/en-us/data-center/tensor-cores/, NVIDIA Corporation, 2026, accessed: 2026-02-09
2026
-
[58]
A distributed approach to silicon compilation: Invited,
A. Olofsson, W. Ransohoff, and N. Moroze, “A distributed approach to silicon compilation: Invited,” inProceedings of the 59th ACM/IEEE Design Automation Conference, 2022, p. 1343–1346
2022
-
[59]
Scnn: An accelerator for compressed-sparse convolutional neural networks,
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” SIGARCH Comput. Archit. News, vol. 45, no. 2, p. 27–40, Jun. 2017. [Online]. Availabl...
2017
-
[60]
Modeling deep learning accelerator enabled gpus,
M. A. Raihan, N. Goli, and T. M. Aamodt, “Modeling deep learning accelerator enabled gpus,” in2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2019, pp. 79–92
2019
-
[61]
Heax: An architecture for computing on encrypted data,
M. S. Riazi, K. Laine, B. Pelton, and W. Dai, “Heax: An architecture for computing on encrypted data,” inProceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems, ser. ASPLOS ’20. New York, NY , USA: Asso...
2020
-
[62]
Encrypted image classification with low mem- ory footprint using fully homomorphic encryption,
L. Rovida and A. Leporati, “Encrypted image classification with low mem- ory footprint using fully homomorphic encryption,”International Journal of Neural Systems, vol. 34, no. 05, p. 2450025, 2024, pMID: 38516871. [Online]. Available: https://doi.org/10.1142/S0129065724500254
2024 doi
-
[63]
A systematic methodology for characterizing scalability of dnn accelerators using scale-sim,
A. Samajdar, J. M. Joseph, Y . Zhu, P. Whatmough, M. Mattina, and T. Kr- ishna, “A systematic methodology for characterizing scalability of dnn accelerators using scale-sim,” in2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), 2020, pp. 58–68
2020
-
[64]
F1: A fast and programmable accelerator for fully homomorphic encryption,
N. Samardzic, A. Feldmann, A. Krastev, S. Devadas, R. Dreslinski, C. Peikert, and D. Sanchez, “F1: A fast and programmable accelerator for fully homomorphic encryption,” inMICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO ’21. New York, NY...
2021
-
[65]
Craterlake: a hardware accelerator for efficient unbounded computation on encrypted data,
N. Samardzic, A. Feldmann, A. Krastev, N. Manohar, N. Genise, S. Devadas, K. Eldefrawy, C. Peikert, and D. Sanchez, “Craterlake: a hardware accelerator for efficient unbounded computation on encrypted data,” inProceedings of the 49th Annual International Symposium on Computer ...
2022
-
[66]
Microsoft SEAL (release 4.1),
“Microsoft SEAL (release 4.1),” https://github.com/Microsoft/SEAL, Jan. 2023, microsoft Research, Redmond, W A
2023
-
[68]
Gme: Gpu-based microarchitectural extensions to accelerate homomorphic encryption,
K. Shivdikar, Y . Bao, R. Agrawal, M. Shen, G. Jonatan, E. Mora, A. Ingare, N. Livesay, J. L. Abell ´AN, J. Kim, A. Joshi, and D. Kaeli, “Gme: Gpu-based microarchitectural extensions to accelerate homomorphic encryption,” inProceedings of the 56th Annual IEEE/ACM International...
2023
-
[69]
A computational introduction to number theory and algebra,
V . Shoup, “A computational introduction to number theory and algebra,” New York University, Tech. Rep., 2005, see Section on modular reduction with precomputed reciprocals
2005
-
[70]
Deep homeomorphic data encryption for privacy preserving machine learning,
V . Terziyan, B. Bilokon, and M. Gavriushenko, “Deep homeomorphic data encryption for privacy preserving machine learning,”Procedia Computer Science, vol. 232, pp. 2201–2212, 2024, 5th International Conference on Industry 4.0 and Smart Manufacturing (ISM 2023). [Online]. Avail...
2024
-
[71]
Leveraging asic ai chips for homomorphic encryption,
J. Tong, T. Huang, L. de Castro, A. Itagi, J. Dang, A. Golder, A. Ali, J. Jiang, Arvind, G. E. Suh, and T. Krishna, “Leveraging asic ai chips for homomorphic encryption,” 2025. [Online]. Available: https://arxiv.org/abs/2501.07047
2025 arXiv
-
[72]
Zkprophet: Understanding performance of zero-knowledge proofs on gpus,
T. Verma, Y . Yuan, N. Talati, and T. Austin, “Zkprophet: Understanding performance of zero-knowledge proofs on gpus,” 2025. [Online]. Available: https://arxiv.org/abs/2509.22684
2025
-
[73]
Nvbit: A dynamic binary instrumentation framework for nvidia gpus,
O. Villa, M. Stephenson, D. Nellans, and S. W. Keckler, “Nvbit: A dynamic binary instrumentation framework for nvidia gpus,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture, ser. MICRO-52. New York, NY , USA: Association for Computing Ma...
2019
-
[74]
Standard cell library design and optimization methodology for asap7 pdk: (invited paper),
X. Xu, N. Shah, A. Evans, S. Sinha, B. Cline, and G. Yeric, “Standard cell library design and optimization methodology for asap7 pdk: (invited paper),” in2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD), 2017, pp. 999–1004
2017
-
[75]
Phantom: A cuda-accelerated word-wise homomorphic encryption library,
H. Yang, S. Shen, W. Dai, L. Zhou, Z. Liu, and Y . Zhao, “Phantom: A cuda-accelerated word-wise homomorphic encryption library,”IEEE Transactions on Dependable and Secure Computing, vol. 21, no. 5, pp. 4895–4906, 2024
2024
-
[76]
An efficient image homomorphic encryption scheme with small ciphertext expansion,
P. Zheng and J. Huang, “An efficient image homomorphic encryption scheme with small ciphertext expansion,” inProceedings of the 21st ACM International Conference on Multimedia, ser. MM ’13. New York, NY , USA: Association for Computing Machinery, 2013, p. 803–812. [Online]. Av...
2013
-
[77]
HEonGPU: a GPU-based fully homomorphic encryption library 1.0,
A. S ¸ah ¨Ozcan and E. Sava s ¸, “HEonGPU: a GPU-based fully homomorphic encryption library 1.0,” Cryptology ePrint Archive, Paper 2024/1543, 2024. [Online]. Available: https://eprint.iacr.org/2024/1543
2024
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.