On a commercial compute-in-SRAM device, three data-movement optimizations enable RAG retrieval at GPU-level latency with 54x-118x lower energy, using simulated HBM for off-chip bandwidth.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Characterizing and Optimizing Realistic Workloads on a Commercial Compute-in-SRAM Device
On a commercial compute-in-SRAM device, three data-movement optimizations enable RAG retrieval at GPU-level latency with 54x-118x lower energy, using simulated HBM for off-chip bandwidth.