REVIEW 3 major objections 6 minor 80 references
EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
T0 review · 3 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read EFFACT claims that a compact FHE accelerator with 27 MB of SRAM can approach the throughput of resource-unconstrained designs while using a fraction of the area and power.
desk verdict Credible FHE accelerator with real RTL evidence, but the headline 1.22x/1.46x/1.48x margins all rest on an unvalidated linear-scaling extrapolation from a 12.5 MHz/64-lane prototype. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a compiler-directed partial streaming dataflow: single-consumer temporaries are merged into the instruction that uses them and are delivered from DRAM through a FIFO directly into the functional unit, bypassing SRAM staging. A second mechanism is circuit-level function-unit reuse, in which the NTT butterfly units are reconfigured to run as multiply-accumulate units, so the same modular multipliers serve both NTT and ordinary MULT/ADD work. A fine-grained NTT unit that shares one modular multiplier and adder across all pipeline stages, and a double-Montgomery representation that merges iNTT post-scaling and base-conversion constants, further reduce area without a proportional throughput loss.
What would settle it
Measure the 256-lane FPGA at its advertised 300 MHz clock after relieving the reported routing congestion; if the observed bootstrapping latency and DRAM bandwidth utilization are not close to the values linearly extrapolated from the 12.5 MHz prototype, the headline speedup and efficiency ratios shrink.
Extended reading notes
Core claim
EFFACT's central claim is that a cost-sensitive FHE accelerator can be both small and fast if its resources match the actual instruction mix of real workloads. Profiling bootstrapping, HELR, and ResNet-20 at the residue-polynomial level shows that NTT instructions are a small fraction of total instructions and that most modular MULT/ADD instructions cannot overlap with NTT. EFFACT therefore removes dedicated base-conversion units, uses a fine-grained NTT unit, reuses the butterfly data paths as MAC units, and adds a compiler pass that streams single-use operands from DRAM straight to functional units. On this basis, the paper reports a 27 MB SRAM ASIC that runs fully-packed bootstrapping in 0.0548 ms amortized time and claims performance-per-area and per-watt gains of at least 1.46× and 1.48× over prior ASIC accelerators, with an FPGA version claiming a 1.22× geometric-mean speedup over prior FPGAs.
Load-bearing premise
The reported FPGA and ASIC numbers are scaled from a physical prototype that runs at 12.5 MHz with 64 lanes, so the central claim assumes that throughput scales linearly with frequency and lane count all the way to 300 MHz with 256 lanes and to the 1024-lane ASIC.
Editorial extensions
If this is right
- If the reported scaling is sound, 27 MB of SRAM and about two thousand multipliers place bootstrapping throughput within a small multiple of designs with more than 280 MB of SRAM.
- The compiler's streaming and scheduling pass reduces bootstrapping DRAM transfers by roughly 40 percent relative to the memory-aware baseline, so the efficiency gains are credited largely to software rather than extra hardware.
- Because the ISA and compiler backend are scheme-generic, the same hardware accelerates CKKS, BGV, and BFV workloads, and the paper demonstrates this with a BGV database-lookup workload.
- Removing dedicated base-conversion hardware and reusing NTT units as MAC units is claimed to preserve performance while cutting computing area, since the profiled workloads keep most MULT/ADD work serialized behind NTT chains.
Reading between the lines
- The paper leaves implicit that its 27 MB SRAM choice is a knee in the design space: the sensitivity study shows EFFACT-54 and EFFACT-108 scale almost linearly on HELR and ResNet, so a memory-heavy or throughput-critical deployment could rationally trade area for another 2–3×.
- The instruction-mix analysis is portable: any ring-based FHE accelerator with the same serial NTT/BConv pattern could adopt the fine-grained NTT plus MAC-reuse scheme even if it keeps a larger SRAM.
- A direct test the paper does not run is whether the scheme-generic ISA extends efficiently to bit-oriented TFHE workloads; the automorphism unit's shift mode is sketched but not measured, so that is the natural next benchmark.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents EFFACT, a full-stack FHE acceleration platform comprising a vector ISA, a compiler backend with static scheduling and streaming optimizations, and RTL implementations targeting both FPGA and ASIC. The authors argue that by rebalancing computing resources, using a 27 MB on-chip SRAM, and introducing streaming memory access plus circuit-level function-unit reuse, a cost-sensitive accelerator can approach the throughput of much larger designs. The experimental sections report an FPGA version running on a VCU128 board and an ASIC version synthesized in 28 nm, with claims that FPGA-EFFACT outperforms state-of-the-art FPGA accelerators by 1.22x geometric mean and that ASIC-EFFACT achieves at least 1.46x performance per area and 1.48x performance per Watt versus prior ASICs.
Significance. If the headline numbers hold, EFFACT would be a substantial step toward practical, cost-effective FHE acceleration: it demonstrates that a 27 MB SRAM, ~2K multipliers, and a 500 MHz clock can deliver bootstrapping and ML workloads at efficiency comparable to much larger and faster designs. The paper's strengths include a complete RTL implementation, functional verification against the Lattigo software library on the FPGA, an actual end-to-end FPGA evaluation, a compiler with automatic scheduling and streaming, and a design-space analysis that goes beyond simple resource scaling. The novelty of the streaming memory access and the NTT-as-MAC reuse scheme is plausible and well motivated by the instruction-mix analysis. However, the central efficiency claims rest on unmeasured extrapolation from the 12.5 MHz/64-lane FPGA prototype to the 300 MHz/256-lane and 500 MHz/1024-lane configurations, which is a load-bearing weakness that must be resolved before the results can be accepted as stated.
major comments (3)
- [§V.C and Table VII] The headline performance and efficiency numbers are not measured at the advertised configurations. Section V.C states that the FPGA 'runs only at 12.5 MHz with 64 lanes' and that the authors 'scale the performance of the 12.5 MHz with 64 lanes version to our target 300 MHz with 256 lanes FPGA-EFFACT and 1024 lanes ASIC-EFFACT.' Every decisive figure, including the 0.566 us / 0.0548 us bootstrapping times, the HELR times, and the Figures 9 and 10 efficiency ratios, therefore assumes throughput scales linearly with frequency and lane count. The paper provides no measurement at any intermediate point to validate this assumption. The reported routing congestion level 7, the deliberately lowered HBM bandwidth through an asynchronous FIFO, and the observation that bootstrapping is memory-bound jointly make linear scaling unlikely to hold exactly. Because the claimed margins over SOTA are modest (1.22x, 1.46x, 1.48x), even a 20–30% shortfall from linear scaling could erase them. The authors should either measure the scaled configurations end-to-end or provide a validated performance model anchored to the measured 12.5 MHz runtime, with explicit accounting for memory bandwidth scaling and congestion effects.
- [§VI.C] The cycle-accurate simulator used for the scalability study (EFFACT-54/108/162 and Figure 10) and for the DRAM-transfer analysis (Figure 11) is not calibrated against the 12.5 MHz/64-lane FPGA implementation. The simulator is the only evidence that performance scales as resources are added, and it is also used to attribute the DRAM-transfer reduction to the streaming optimization. Without a reported comparison of simulated versus measured cycles for at least one benchmark on the actual FPGA, the simulator's treatment of memory stalls, NTT pipeline conflicts, and arbiter contention is unverified. I request a calibration plot (simulated vs. measured cycles for bootstrapping or HELR on the 64-lane prototype) and, ideally, a validation of the simulator's scaling predictions against a second measured configuration.
- [§IV.D.5 and Figure 4] The 27 MB SRAM capacity is selected from the paper's own design-space exploration ('we choose 27MB as a trade-off') and then used in the evaluated ASIC configuration and in the comparisons against prior designs that use much larger SRAMs. This is a legitimate design methodology, but it makes the efficiency claims sensitive to the chosen operating point on the SRAM-vs-performance curve. The paper should clarify whether 27 MB is a conclusion of the analysis or an input, and report how the headline area- and power-efficiency ratios change if the second turning point (54 MB) is used instead. This would strengthen the robustness of the '1.46x/1.48x' claims against the criticism that the comparison is tuned to a single favorable point.
minor comments (6)
- [§VI.B] The speedup factors for bootstrapping (13.49x, 4743.79x, 0.82x, 0.31x, 0.26x, and 4.93x for GPU, F1, BTS, CraterLake, ARK, and MAD) do not align unambiguously with the entries in Table VII: for example, the GPU column ('Over 100x') lists 0.270 us, which gives 4.93x relative to ASIC-EFFACT's 0.0548 us, not 13.49x. Please verify the ordering and the values.
- [§VI.B] The text refers to 'SAHRP-EFFACT'; this appears to be a typo for 'SHARP-EFFACT'.
- [Figure 4] The figure caption contains stray '(a)' and '(b)(a)' labels that do not match the intended subfigure structure; please clean up the caption and the panel labels.
- [§IV.B.1] The sentence 'we have excluded code optimization from our evaluation' is unclear, since the sensitivity study in Figure 11 appears to attribute part of the runtime improvement to 'global streaming and memory opt' and 'full EFFACT', both of which involve compiler passes. Please clarify what exactly is excluded and how the 12.9% instruction reduction figure is used.
- [§VI.D] The TFHE bootstrapping result (0.576 ms) is presented without comparison to prior TFHE accelerators or to the authors' own CKKS results; please provide context or soften the claim of 'excellent acceleration capabilities' for boolean schemes.
- [Table VII] The column labeled 'Over 100x' is used interchangeably as 'GPU [30]' in the text; please unify the notation so the reader can map the table to the references.
Circularity Check
No circularity: headline margins come from an unvalidated linear-scaling extrapolation of a measured prototype, which is a correctness risk rather than a self-referential derivation.
full rationale
No circularity found. EFFACT's central efficiency claims rest on a measured 12.5 MHz/64-lane FPGA prototype whose runtime is scaled to the 300 MHz/256-lane FPGA and 1024-lane ASIC configurations: 'Worth mentioning that the system in FPGA runs only at 12.5 MHz with 64 lanes, although we successfully synthesized it at 300 MHz with 256 lanes on Vivado. The major bottleneck lies in the routing congestion in which the congestion level reaches 7. The HBM bandwidth is also lowered through an asynchronous FIFO to ensure that we can correctly scale the performance of the 12.5 MHz with 64 lanes version to our target 300 MHz with 256 lanes FPGA-EFFACT and 1024 lanes ASIC-EFFACT.' This is an unvalidated linear-scaling assumption and therefore a measurement-validity risk, but it is not a circular reduction: the anchor runtime is measured, and the target values are not fitted parameters or hidden unknowns. The 27 MB SRAM size is selected from the paper's own design-space exploration ('We choose 27MB as a trade-off between performance, cost, and efficiency'), which is a mild internal-tuning loop, but the paper does not then present that choice as a predicted result; the comparisons in Table VII and Figures 9-10 are against externally published accelerators (F1, BTS, CraterLake, ARK, MAD, FAB, Poseidon), and functional correctness is checked against Lattigo. The paper also states that compiler code optimization was excluded from evaluation because prior designs relied on manual optimizations, avoiding a self-scoring loop. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted constants hidden in the performance model. The linear-scaling extrapolation should be weighed as a correctness concern, not as circularity.
Assumptions & free parameters
free parameters (4)
- On-chip SRAM capacity =
27 MB
- ASIC lane count =
1024 lanes
- ASIC clock frequency =
500 MHz
- FPGA target frequency and lane count =
300 MHz, 256 lanes
assumptions (4)
- domain assumption Performance scales linearly with clock frequency and lane count in the target configurations.
- domain assumption The bootstrapping, HELR, ResNet-20, and DBLookup workloads are representative for the FHE operation mix used to allocate resources.
- domain assumption Technology scaling of prior ASIC area and power using references [51], [72], [73] is accurate enough for the efficiency comparisons.
- domain assumption Verification against Lattigo establishes functional correctness of the RTL, so the scaled runtime reflects the intended computation.
Cite this review
Pith. "Pith review of EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform." pith.science (2026). https://pith.science/paper/NKRHA3X6
@misc{pith2026250415817,
author = {Pith},
title = {Pith review of: EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform},
year = {2026},
howpublished = {\url{https://pith.science/paper/NKRHA3X6}},
note = {Machine review of arXiv:2504.15817}
}
read the original abstract
Fully Homomorphic Encryption (FHE) is a set of powerful cryptographic schemes that allows computation to be performed directly on encrypted data with an unlimited depth. Despite FHE's promising in privacy-preserving computing, yet in most FHE schemes, ciphertext generally blows up thousands of times compared to the original message, and the massive amount of data load from off-chip memory for bootstrapping and privacy-preserving machine learning applications (such as HELR, ResNet-20), both degrade the performance of FHE-based computation. Several hardware designs have been proposed to address this issue, however, most of them require enormous resources and power. An acceleration platform with easy programmability, high efficiency, and low overhead is a prerequisite for practical application. This paper proposes EFFACT, a highly efficient full-stack FHE acceleration platform with a compiler that provides comprehensive optimizations and vector-friendly hardware. We start by examining the computational overhead across different real-world benchmarks to highlight the potential benefits of reallocating computing resources for efficiency enhancement. Then we make a design space exploration to find an optimal SRAM size with high utilization and low cost. On the other hand, EFFACT features a novel optimization named streaming memory access which is proposed to enable high throughput with limited SRAMs. Regarding the software-side optimization, we also propose a circuit-level function unit reuse scheme, to substantially reduce the computing resources without performance degradation. Moreover, we design novel NTT and automorphism units that are suitable for a cost-sensitive and highly efficient architecture, leading to low area. For generality, EFFACT is also equipped with an ISA and a compiler backend that can support several FHE schemes like CKKS, BGV, and BFV.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Efficient backtracking instruction schedulers,
S. Abraham, W. Meleis, and I. Baev, “Efficient backtracking instruction schedulers,” in Proceedings 2000 International Conference on Parallel Architectures and Compilation Techniques (Cat. No.PR00622), pp. 301– 308, ISSN: 1089-795X
work page 2000
-
[2]
Mad: Memory-aware design techniques for accelerating fully homomorphic encryption,
R. Agrawal, L. d. Castro, C. Juvekar, A. Chandrakasan, V . Vaikun- tanathan, and A. Joshi, “Mad: Memory-aware design techniques for accelerating fully homomorphic encryption,” in 2023 56th IEEE/ACM International Symposium on Microarchitecture (MICRO), 2023, pp. 685– 697
work page 2023
-
[3]
Heap: A fully homomorphic encryption accelerator with parallelized bootstrapping,
R. Agrawal, A. Chandrakasan, and A. Joshi, “Heap: A fully homomorphic encryption accelerator with parallelized bootstrapping,” 2024, https://bu-icsg.github.io/ publications/2024/fhe parallelized bootstrapping isca 2024.pdf. [On- line]. Available: https://bu-icsg.github.io/publications/2024/fhe parallelized bootstrapping isca 2024.pdf
work page 2024
-
[4]
FAB: An FPGA-based Accelerator for Bootstrappable Fully Homomorphic Encryption
R. Agrawal, L. de Castro, G. Yang, C. Juvekar, R. Yazicigil, A. Chandrakasan, V . Vaikuntanathan, and A. Joshi, “FAB: An FPGA- based accelerator for bootstrappable fully homomorphic encryption,” version: 1. [Online]. Available: http://arxiv.org/abs/2207.11872
-
[5]
E2E near-standard and practical authenticated transciphering,
E. Aharoni, N. Drucker, G. Ezov, E. Kushnir, H. Shaul, and O. Soceanu, “E2E near-standard and practical authenticated transciphering,” Cryptology ePrint Archive, Paper 2023/1040, 2023. [Online]. Available: https://eprint.iacr.org/2023/1040
work page 2023
-
[6]
Program analysis and specialization for the c programming language,
L. O. Andersen and P. Lee, “Program analysis and specialization for the c programming language,” 2005
work page 2005
-
[7]
Low-cost and area-efficient fpga implementations of lattice-based cryptography,
A. Aysu, C. Patterson, and P. Schaumont, “Low-cost and area-efficient fpga implementations of lattice-based cryptography,” in 2013 IEEE International Symposium on Hardware-Oriented Security and Trust (HOST), 2013, pp. 81–86
work page 2013
-
[8]
Openfhe: Open-source fully homomorphic encryption library,
A. A. Badawi, J. Bates, F. Bergamaschi, D. B. Cousins, S. Erabelli, N. Genise, S. Halevi, H. Hunt, A. Kim, Y . Lee, Z. Liu, D. Micciancio, I. Quah, Y . Polyakov, S. R.V ., K. Rohloff, J. Saylor, D. Suponitsky, M. Triplett, V . Vaikuntanathan, and V . Zucca, “Openfhe: Open-source fully homomorphic encryption library,” Cryptology ePrint Archive, Paper 2022/...
2022
Show all 80 references
-
[9]
Ffts in external or hierarchical memory,
D. H. Bailey, “Ffts in external or hierarchical memory,” in Super- computing ’89:Proceedings of the 1989 ACM/IEEE Conference on Supercomputing, 1989, pp. 234–242
1989
-
[10]
A full RNS variant of FV like somewhat homomorphic encryption schemes,
J.-C. Bajard, J. Eynard, M. A. Hasan, and V . Zucca, “A full RNS variant of FV like somewhat homomorphic encryption schemes,” in Selected Areas in Cryptography – SAC 2016 , R. Avanzi and H. Heys, Eds. Springer International Publishing, vol. 10532, pp. 423–442, series Title: Le...
2016 doi
-
[11]
2.3 an energy-efficient configurable lattice cryptography processor for the quantum-secure internet of things,
U. Banerjee, A. Pathak, and A. P. Chandrakasan, “2.3 an energy-efficient configurable lattice cryptography processor for the quantum-secure internet of things,” in 2019 IEEE International Solid-State Circuits Conference - (ISSCC) , 2019, pp. 46–48
2019
-
[12]
Intel hexl: Accelerating homomorphic encryption with intel avx512-ifma52,
F. Boemer, S. Kim, G. Seifu, F. D.M. de Souza, and V . Gopal, “Intel hexl: Accelerating homomorphic encryption with intel avx512-ifma52,” in Proceedings of the 9th on Workshop on Encrypted Computing & Applied Homomorphic Cryptography, ser. W AHC ’21. New York, NY , USA: Associ...
2021
-
[13]
Efficient bootstrapping for approximate homomorphic encryption with non-sparse keys,
J.-P. Bossuat, C. Mouchet, J. Troncoso-Pastoriza, and J.-P. Hubaux, “Efficient bootstrapping for approximate homomorphic encryption with non-sparse keys,” Cryptology ePrint Archive, Paper 2020/1203, 2020, https://eprint.iacr.org/2020/1203. [Online]. Available: https://eprint.i...
2020
-
[14]
(leveled) fully homo- morphic encryption without bootstrapping,
Z. Brakerski, C. Gentry, and V . Vaikuntanathan, “(leveled) fully homo- morphic encryption without bootstrapping,” p. 35
-
[15]
Effective partial redundancy elimination,
P. Briggs and K. D. Cooper, “Effective partial redundancy elimination,” ACM SIGPLAN Notices , vol. 29, no. 6, pp. 159–170, 1994
1994
-
[16]
Simple encrypted arithmetic library v2.3.0,
H. Chen, K. Han, Z. Huang, A. Jalali, and K. Laine, “Simple encrypted arithmetic library v2.3.0,” p. 35
-
[17]
A full RNS variant of approximate homomorphic encryption,
J. H. Cheon, K. Han, A. Kim, M. Kim, and Y . Song, “A full RNS variant of approximate homomorphic encryption,” in Selected Areas in Cryptography – SAC 2018 , C. Cid and M. J. Jacobson, Eds. Springer International Publishing, vol. 11349, pp. 347–368, series Title: Lecture Notes...
2018 doi
-
[18]
Homomorphic encryption for arithmetic of approximate numbers
J. H. Cheon, A. Kim, M. Kim, and Y . Song, “Homomorphic encryption for arithmetic of approximate numbers.” [Online]. Available: https://eprint.iacr.org/undefined/undefined
-
[19]
Tfhe: Fast fully homomorphic encryption over the torus,
I. Chillotti, N. Gama, M. Georgieva, and M. Izabach `ene, “Tfhe: Fast fully homomorphic encryption over the torus,” Cryptology ePrint Archive, Paper 2018/421, 2018, https://eprint.iacr.org/2018/421. [Online]. Available: https://eprint.iacr.org/2018/421
2018
-
[20]
TFHE: Fast Fully Homomorphic Encryption Over the Torus,
I. Chillotti, N. Gama, M. Georgieva, and M. Izabach `ene, “TFHE: Fast Fully Homomorphic Encryption Over the Torus,” Journal of Cryptology, vol. 33, no. 1, pp. 34–91, Jan. 2020. [Online]. Available: https://doi.org/10.1007/s00145-019-09319-x
2020 doi
-
[21]
Somewhat practical fully homomorphic encryption,
J. Fan and F. Vercauteren, “Somewhat practical fully homomorphic encryption,” Cryptology ePrint Archive, Paper 2012/144, 2012, https://eprint.iacr.org/2012/144. [Online]. Available: https://eprint.iacr. org/2012/144
2012
-
[22]
Privacy-preserving semi-parallel logistic regression training with Fully Homomorphic Encryption,
Georgieva, G. Nicolas, Mariya, C. Sergiu, Troncoso-Pastoriza, and J. Ramon, “Privacy-preserving semi-parallel logistic regression training with Fully Homomorphic Encryption,” 2019, report Number: 101. [Online]. Available: https://eprint.iacr.org/2019/101
2019
-
[23]
An improved RNS variant of the BFV homomorphic encryption scheme,
S. Halevi, Y . Polyakov, and V . Shoup, “An improved RNS variant of the BFV homomorphic encryption scheme,” in Topics in Cryptology – CT-RSA 2019, M. Matsui, Ed. Springer International Publishing, vol. 11405, pp. 83–105, series Title: Lecture Notes in Computer Science. [Online...
-
[24]
Design and implementation of HElib: a homomorphic encryption library,
S. Halevi and V . Shoup, “Design and implementation of HElib: a homomorphic encryption library,” Cryptology ePrint Archive, Paper 2020/1481, 2020, https://eprint.iacr.org/2020/1481. [Online]. Available: https://eprint.iacr.org/2020/1481
2020
-
[25]
Logistic Regression on Homomorphic Encrypted Data at Scale,
K. Han, S. Hong, J. H. Cheon, and D. Park, “Logistic Regression on Homomorphic Encrypted Data at Scale,” AAAI, vol. 33, pp. 9466–9471, Jul. 2019. [Online]. Available: https://www.aaai.org/ojs/ index.php/AAAI/article/view/5000
2019
-
[26]
Better bootstrapping for approximate homomorphic encryption,
K. Han and D. Ki, “Better bootstrapping for approximate homomorphic encryption,” Cryptology ePrint Archive, Paper 2019/688, 2019, https://eprint.iacr.org/2019/688. [Online]. Available: https://eprint.iacr. org/2019/688
2019
-
[27]
Hbm2e and gddr6: Memory solutions for ai,
R. Inc, “Hbm2e and gddr6: Memory solutions for ai,” 2020, https://go.rambus.com/hbm2e-gddr6-memory-solutions-for-ai. [Online]. Available: https://go.rambus.com/hbm2e-gddr6-memory-solutions-for- ai
2020
-
[28]
Threshold fully homomor- phic encryption,
A. Jain, P. M. R. Rasmussen, and A. Sahai, “Threshold fully homomor- phic encryption,” p. 40
-
[29]
Matcha: A fast and energy- efficient accelerator for fully homomorphic encryption over the torus,
L. Jiang, Q. Lou, and N. Joshi, “Matcha: A fast and energy- efficient accelerator for fully homomorphic encryption over the torus,” in Proceedings of the 59th ACM/IEEE Design Automation Conference, ser. DAC ’22. New York, NY , USA: Association for Computing Machinery, 2022, p....
2022
-
[30]
Over 100x faster bootstrapping in fully homomorphic encryption through memory- centric optimization with GPUs,
W. Jung, S. Kim, J. H. Ahn, J. H. Cheon, and Y . Lee, “Over 100x faster bootstrapping in fully homomorphic encryption through memory- centric optimization with GPUs,” Cryptology ePrint Archive, Paper 2021/508, 2021, https://eprint.iacr.org/2021/508. [Online]. Available: https:...
2021
-
[31]
Accelerating fully homomorphic encryption through architecture-centric analysis and optimization,
W. Jung, E. Lee, S. Kim, J. Kim, N. Kim, K. Lee, C. Min, J. H. Cheon, and J. H. Ahn, “Accelerating fully homomorphic encryption through architecture-centric analysis and optimization,” IEEE Access, vol. 9, pp. 98 772–98 789, 2021
2021
-
[32]
Partial redundancy elimination in SSA form,
R. Kennedy, S. Chan, S.-M. Liu, R. Lo, P. Tu, and F. Chow, “Partial redundancy elimination in SSA form,” vol. 21, no. 3, pp. 627–676. [Online]. Available: https://dl.acm.org/doi/10.1145/319301.319348
-
[34]
ARK: Fully homomorphic encryption accelerator with runtime data generation and inter-operation key reuse
J. Kim, G. Lee, S. Kim, G. Sohn, J. Kim, M. Rhu, and J. H. Ahn, “ARK: Fully homomorphic encryption accelerator with runtime data generation and inter-operation key reuse.” [Online]. Available: http://arxiv.org/abs/2205.00922
-
[35]
BTS: An accelerator for bootstrappable fully homomorphic encryption,
S. Kim, J. Kim, M. J. Kim, W. Jung, M. Rhu, J. Kim, and J. H. Ahn, “BTS: An accelerator for bootstrappable fully homomorphic encryption,” in Proceedings of the 49th Annual International Symposium on Computer Architecture , pp. 711–725. [Online]. Available: http://arxiv.org/abs...
-
[36]
Lazy code motion,
J. Knoop, O. R ¨uthing, and B. Steffen, “Lazy code motion,” vol. 27, no. 7, pp. 224–234. [Online]. Available: https://dl.acm.org/doi/10.1145/ 143103.143136
-
[37]
Aha: An agile approach to the design of coarse-grained reconfigurable accelerators and compilers,
K. Koul, J. Melchert, K. Sreedhar, L. Truong, G. Nyengele, K. Zhang, Q. Liu, J. Setter, P.-H. Chen, Y . Mei, M. Strange, R. Daly, C. Donovick, A. Carsello, T. Kong, K. Feng, D. Huff, A. Nayak, R. Setaluri, J. Thomas, N. Bhagdikar, D. Durst, Z. Myers, N. Tsiskaridze, S. Richard...
-
[38]
Automatic domain-specific soc design for autonomous unmanned aerial vehicles,
S. Krishnan, Z. Wan, K. Bhardwaj, P. Whatmough, A. Faust, S. Neuman, G.-Y . Wei, D. Brooks, and V . J. Reddi, “Automatic domain-specific soc design for autonomous unmanned aerial vehicles,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2022, pp. 300–317
2022
-
[39]
Archgym: An open-source gymnasium for machine learning assisted architecture design,
S. Krishnan, A. Yazdanbakhsh, S. Prakash, J. Jabbour, I. Uchendu, S. Ghosh, B. Boroujerdian, D. Richins, D. Tripathy, A. Faust, and V . Janapa Reddi, “Archgym: An open-source gymnasium for machine learning assisted architecture design,” in Proceedings of the 50th Annual Intern...
-
[40]
PALISADE lattice cryptography library,
Y . Lab., “PALISADE lattice cryptography library,” https://github.com/ yamanalab/PALISADE, 2021
2021
-
[41]
Available: https://doi.org/10.1145/3579371.3589049
[Online]. Available: https://doi.org/10.1145/3579371.3589049
-
[42]
Low-complexity deep convolutional neural networks on fully homomorphic encryption using multiplexed parallel convolutions,
E. Lee, J.-W. Lee, J. Lee, Y .-S. Kim, Y . Kim, J.-S. No, and W. Choi, “Low-complexity deep convolutional neural networks on fully homomorphic encryption using multiplexed parallel convolutions,” Cryptology ePrint Archive, Paper 2021/1688, 2021, https://eprint.iacr. org/2021/1...
2021
-
[43]
Llvm: a compilation framework for lifelong program analysis & transformation,
C. Lattner and V . Adve, “Llvm: a compilation framework for lifelong program analysis & transformation,” in International Symposium on Code Generation and Optimization, 2004. CGO 2004. , 2004, pp. 75–86
2004
-
[44]
Performance- aware scale analysis with reserve for homomorphic encryption,
Y . Lee, S. Cheon, D. Kim, D. Lee, and H. Kim, “Performance- aware scale analysis with reserve for homomorphic encryption,” in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 , ser. ASPLOS ...
2024
-
[45]
Privacy-preserving machine learning with fully homomorphic encryption for deep neural network,
J.-W. Lee, H. Kang, Y . Lee, W. Choi, J. Eom, M. Deryabin, E. Lee, J. Lee, D. Yoo, Y .-S. Kim, and J.-S. No, “Privacy-preserving machine learning with fully homomorphic encryption for deep neural network,” IEEE Access, vol. 10, pp. 30 039–30 054, 2022
2022
-
[46]
On- device training under 256kb memory,
J. Lin, L. Zhu, W.-M. Chen, W.-C. Wang, C. Gan, and S. Han, “On- device training under 256kb memory,” in Proceedings of the 36th International Conference on Neural Information Processing Systems, ser. NIPS ’22. Red Hook, NY , USA: Curran Associates Inc., 2024
2024
-
[47]
Semi- parallel Logistic Regression for GW AS on Encrypted Data,
Li, S. Yongsoo, Baiyu, K. Miran, Micciancio, and Daniele, “Semi- parallel Logistic Regression for GW AS on Encrypted Data,” 2019, report Number: 294. [Online]. Available: https://eprint.iacr.org/2019/294
2019
-
[48]
Scale-out processors,
P. Lotfi-Kamran, B. Grot, M. Ferdman, S. V olos, O. Kocberber, J. Picorel, A. Adileh, D. Jevdjic, S. Idgunji, E. Ozer, and B. Falsafi, “Scale-out processors,” SIGARCH Comput. Archit. News , vol. 40, no. 3, p. 500–511, jun 2012. [Online]. Available: https: //doi.org/10.1145/236...
2012
-
[49]
Overgen: Improving fpga usability through domain-specific overlay generation,
S. Liu, J. Weng, D. Kupsh, A. Sohrabizadeh, Z. Wang, L. Guo, J. Liu, M. Zhulin, R. Mani, L. Zhang, J. Cong, and T. Nowatzki, “Overgen: Improving fpga usability through domain-specific overlay generation,” in 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICR...
2022
-
[50]
CoFHEE: A co-processor for fully homomorphic encryption execution,
M. Nabeel, D. Soni, M. Ashraf, M. A. Gebremichael, H. Gamil, E. Chielle, R. Karri, M. Sanduleanu, and M. Maniatakos, “CoFHEE: A co-processor for fully homomorphic encryption execution,” version:
-
[51]
Lattigo: a multiparty homomorphic encryption library in go,
C. Mouchet, J.-P. Bossuat, J. Troncoso-Pastoriza, and J.-P. Hubaux, “Lattigo: a multiparty homomorphic encryption library in go,” p. 6
-
[52]
Ciflow: Dataflow analysis and optimization of key switching for homomorphic encryption,
N. Neda, A. Ebel, B. Reynwar, and B. Reagen, “Ciflow: Dataflow analysis and optimization of key switching for homomorphic encryption,” 2024. [Online]. Available: https://arxiv.org/abs/2311.01598
2024 arXiv
-
[53]
Available: http://arxiv.org/abs/2204.08742
[Online]. Available: http://arxiv.org/abs/2204.08742
-
[54]
A 7nm cmos technology platform for mobile and high performance compute application,
S. Narasimha, B. Jagannathan, A. Ogino, D. Jaeger, B. Greene, C. Sheraw, K. Zhao, B. Haran, U. Kwon, A. K. M. Mahalingam, B. Kan- nan, B. Morganfeld, J. Dechene, C. Radens, A. Tessier, A. Hassan, H. Narisetty, I. Ahsan, M. Aminpur, C. An, M. Aquilino, A. Arya, R. Augur, N. Bal...
2017
-
[55]
Modular multiplication without trial division,
M. Peter, L., “Modular multiplication without trial division,” Mathematics of Computation , vol. 44, pp. 519–521, 1985. [Online]. Available: https://api.semanticscholar.org/CorpusID:119574413
1985
-
[56]
Stream-dataflow acceleration,
T. Nowatzki, V . Gangadhar, N. Ardalani, and K. Sankaralingam, “Stream-dataflow acceleration,” in Proceedings of the 44th Annual International Symposium on Computer Architecture, 2017, pp. 416–429
2017
-
[57]
Heaan.mlir: An optimizing compiler for fast ring-based homomorphic encryption,
S. Park, W. Song, S. Nam, H. Kim, J. Shin, and J. Lee, “Heaan.mlir: An optimizing compiler for fast ring-based homomorphic encryption,” Proc. ACM Program. Lang. , vol. 7, no. PLDI, jun 2023. [Online]. Available: https://doi.org/10.1145/3591228
2023 doi
-
[58]
CAeSaR: Unified cluster-assignment scheduling and communication reuse for clustered VLIW processors,
V . Porpodas and M. Cintra, “CAeSaR: Unified cluster-assignment scheduling and communication reuse for clustered VLIW processors,” in 2013 International Conference on Compilers, Architecture and Synthesis for Embedded Systems (CASES) , pp. 1–10
2013
-
[59]
Linear scan register allocation,
M. Poletto and V . Sarkar, “Linear scan register allocation,” vol. 21, no. 5, pp. 895–913. [Online]. Available: https://dl.acm.org/doi/10.1145/ 330249.330250
-
[60]
Towards efficient arithmetic for lattice-based cryptography on reconfigurable hardware,
T. P ¨oppelmann and T. G ¨uneysu, “Towards efficient arithmetic for lattice-based cryptography on reconfigurable hardware,” in Progress in Cryptology – LATINCRYPT 2012, A. Hevia and G. Neven, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2012, pp. 139–158
2012
-
[61]
Heax: An architecture for computing on encrypted data,
M. S. Riazi, K. Laine, B. Pelton, and W. Dai, “Heax: An architecture for computing on encrypted data,” Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems , 2019. [Online]. Available: https://api.sem...
2019
-
[62]
Cheetah: Optimizing and accelerating homomorphic encryption for private inference,
B. Reagen, W.-S. Choi, Y . Ko, V . T. Lee, H.-H. S. Lee, G.-Y . Wei, and D. Brooks, “Cheetah: Optimizing and accelerating homomorphic encryption for private inference,” in2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA), pp. 26–39, ISSN: 2378-203X
-
[63]
The learning with errors problem (invited survey),
O. Regev, “The learning with errors problem (invited survey),” in 2010 IEEE 25th Annual Conference on Computational Complexity , 2010, pp. 191–204
2010
-
[64]
FPGA-based high-performance parallel architecture for homomorphic computing on encrypted data,
S. Sinha Roy, F. Turan, K. Jarvinen, F. Vercauteren, and I. Verbauwhede, “FPGA-based high-performance parallel architecture for homomorphic computing on encrypted data,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, pp. 387–
2019
-
[65]
F1: A fast and programmable accelerator for fully homomorphic encryption,
N. Samardzic, A. Feldmann, A. Krastev, S. Devadas, R. Dreslinski, C. Peikert, and D. Sanchez, “F1: A fast and programmable accelerator for fully homomorphic encryption,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture . ACM, pp. 238–252. [Online]...
-
[66]
CraterLake: a hardware accelerator for efficient unbounded computation on encrypted data,
N. Samardzic, A. Feldmann, A. Krastev, N. Manohar, N. Genise, S. Devadas, K. Eldefrawy, C. Peikert, and D. Sanchez, “CraterLake: a hardware accelerator for efficient unbounded computation on encrypted data,” in Proceedings of the 49th Annual International Symposium on Computer...
-
[67]
Aurora: Automated refinement of coarse-grained reconfigurable accelerators,
C. Tan, C. Xie, A. Li, K. J. Barker, and A. Tumeo, “Aurora: Automated refinement of coarse-grained reconfigurable accelerators,” in 2021 De- sign, Automation & Test in Europe Conference & Exhibition (DATE) , 2021, pp. 1388–1393
2021
-
[68]
Quality and speed in linear-scan register allocation,
O. Traub, G. Holloway, and M. D. Smith, “Quality and speed in linear-scan register allocation,” vol. 33, no. 5, pp. 142–151. [Online]. Available: https://dl.acm.org/doi/10.1145/277652.277714
-
[69]
Spatial memory streaming,
S. Somogyi, T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos, “Spatial memory streaming,” ACM SIGARCH Computer Architecture News, vol. 34, no. 2, pp. 252–263, 2006
2006
-
[70]
Dominator-path scheduling: a global scheduling method,
P. H. Sweany and S. J. Beaty, “Dominator-path scheduling: a global scheduling method,” vol. 23, no. 1, pp. 260–263. [Online]. Available: https://dl.acm.org/doi/10.1145/144965.145824
-
[71]
A highly manufacturable 28nm cmos low power platform technology with fully functional 64mb sram using dual/tripe gate oxide process,
S.-Y . Wu, J. Liaw, C. Lin, M. Chiang, C. Yang, J. Cheng, M. Tsai, M. Liu, P. Wu, C. Chang, L. Hu, C. Lin, H. Chen, S. Chang, S. Wang, P. Tong, Y . Hsieh, K. Pan, C. Hsieh, C. Chen, C. Yao, C. Chen, T. Lee, C. Chang, H. Lin, S. Chen, J. Shieh, M. Tsai, S. Jang, K. Chen, Y . Ku...
2009
-
[72]
A 16nm finfet cmos technology for mobile soc and computing applications,
S.-Y . Wu, C. Y . Lin, M. C. Chiang, J. J. Liaw, J. Y . Cheng, S. H. Yang, M. Liang, T. Miyashita, C. H. Tsai, B. C. Hsu, H. Y . Chen, T. Yamamoto, S. Y . Chang, V . S. Chang, C. H. Chang, J. H. Chen, H. F. Chen, K. C. Ting, Y . K. Wu, K. H. Pan, R. F. Tsui, C. H. Yao, P. R. C...
2013
-
[73]
Dsagen: Synthesizing programmable spatial accelerators,
J. Weng, S. Liu, V . Dadu, Z. Wang, P. Shah, and T. Nowatzki, “Dsagen: Synthesizing programmable spatial accelerators,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 268–281
2020
-
[74]
Linear scan register allocation on SSA form,
C. Wimmer and M. Franz, “Linear scan register allocation on SSA form,” in Proceedings of the 8th annual IEEE/ACM international symposium on Code generation and optimization . ACM, pp. 170–179. [Online]. Available: https://dl.acm.org/doi/10.1145/1772954.1772979
-
[75]
Poseidon: Practical homomorphic encryption accelerator,
Y . Yang, H. Zhang, S. Fan, H. Lu, M. Zhang, and X. Li, “Poseidon: Practical homomorphic encryption accelerator,” 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) , pp. 870–881, 2023. [Online]. Available: https: //api.semanticscholar.org/Corpu...
2023
-
[76]
Sok: Fully homomorphic encryption accelerators,
J. Zhang, X. Cheng, L. Yang, J. Hu, X. Liu, and K. Chen, “Sok: Fully homomorphic encryption accelerators,” ACM Computing Surveys,
-
[77]
A 7nm cmos platform technology featuring 4th generation finfet transistors with a 0.027um2 high density 6-t sram cell for mobile soc applications,
S.-Y . Wu, C. Lin, M. Chiang, J. Liaw, J. Cheng, S. Yang, C. Tsai, P. Chen, T. Miyashita, C. Chang, V . Chang, K. Pan, J. Chen, Y . Mor, K. Lai, C. Liang, H. Chen, S. Chang, C. Lin, C. Hsieh, R. Tsui, C. Yao, C. Chen, R. Chen, C. Lee, H. Lin, C. Chang, K. Chen, M. Tsai, K. Che...
2016
-
[78]
Hasco: Towards agile hardware and software co-design for tensor computation,
Q. Xiao, S. Zheng, B. Wu, P. Xu, X. Qian, and Y . Liang, “Hasco: Towards agile hardware and software co-design for tensor computation,” in 2021 ACM/IEEE 48th Annual International Symposium on Computer Architecture (ISCA), 2021, pp. 1055–1068
2021
-
[398]
Available: https://ieeexplore.ieee.org/document/8675244/
[Online]. Available: https://ieeexplore.ieee.org/document/8675244/
-
[2022]
Available: https://api.semanticscholar.org/CorpusID: 254247068
[Online]. Available: https://api.semanticscholar.org/CorpusID: 254247068
- [2023]
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.