PALUTE is a new PIM accelerator using in-DRAM LUTs on M3D DRAM that reports 1264 TPS at 0.16 W with 12.8x energy efficiency gains over CHIME for quantized edge LLM inference.
Data movement is all you need: A case study on optimizing transformers
4 Pith papers cite this work, alongside 24 external citations. Polarity classification is still indexing.
years
2026 4representative citing papers
IO-aware GPU kernels for SpMM convolutions, degree-aware reductions, and fused attention layers deliver median speedups of 1.6-2.6x (up to 10x) and memory reductions up to 76x over DGL/PyG baselines on realistic graphs.
OFU is a hardware-counter metric that approximates application MFU to within 2 percentage points after tile correction and shows r=0.78 correlation on 608 production jobs.
Coordinated microarchitectural fixes to Ara close a large fraction of the sustained-throughput gap to an ideal multi-lane chaining model, yielding 1.33× geometric-mean speedup.
citing papers explorer
-
PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference
PALUTE is a new PIM accelerator using in-DRAM LUTs on M3D DRAM that reports 1264 TPS at 0.16 W with 12.8x energy efficiency gains over CHIME for quantized edge LLM inference.
-
On Efficient Scaling of GNNs via IO-Aware Layers Implementations
IO-aware GPU kernels for SpMM convolutions, degree-aware reductions, and fused attention layers deliver median speedups of 1.6-2.6x (up to 10x) and memory reductions up to 76x over DGL/PyG baselines on realistic graphs.
-
Instant GPU Efficiency Visibility at Fleet Scale
OFU is a hardware-counter metric that approximates application MFU to within 2 percentage points after tile correction and shows r=0.78 correlation on 608 production jobs.
-
Microarchitectural Co-Optimization for Sustained Throughput of RISC-V Multi-Lane Chaining Vector Processors
Coordinated microarchitectural fixes to Ara close a large fraction of the sustained-throughput gap to an ideal multi-lane chaining model, yielding 1.33× geometric-mean speedup.