A pruning method and a digital compute-in-memory accelerator that jointly support layer-wise flexible N:M sparsity, improving LLM perplexity and zero-shot accuracy over fixed N:M baselines while cutting simulated inference latency and energy.
15.3 a 65nm 3t dynamic analog ram- based computing-in-memory macro and cnn accelerator with retention enhancement, adaptive analog sparsity and 44tops/w system energy efficiency,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
A pruning method and a digital compute-in-memory accelerator that jointly support layer-wise flexible N:M sparsity, improving LLM perplexity and zero-shot accuracy over fixed N:M baselines while cutting simulated inference latency and energy.