A BFP NPU microarchitecture using row/column blocking and per-path protections achieves near-DMR reliability at 3.55% geometric mean performance overhead and under 2% hardware cost.
Designing Efficient LLM Accelerators for Edge Devices,
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
Sparse FHE matrix multiplication on AMD GPUs via FIDESlib achieves 3x CPU speedup and shifts complexity from cubic to semi-linear.
SECDA-DSE uses LLMs with RAG and chain-of-thought to generate three FPGA accelerator designs that synthesize and run on hardware, extending prior SECDA work.
citing papers explorer
-
From Characterization to Microarchitecture: Designing an Elegant and Reliable BFP-Based NPU
A BFP NPU microarchitecture using row/column blocking and per-path protections achieves near-DMR reliability at 3.55% geometric mean performance overhead and under 2% hardware cost.
-
GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs
Sparse FHE matrix multiplication on AMD GPUs via FIDESlib achieves 3x CPU speedup and shifts complexity from cubic to semi-linear.
-
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA
SECDA-DSE uses LLMs with RAG and chain-of-thought to generate three FPGA accelerator designs that synthesize and run on hardware, extending prior SECDA work.