A dual-view memory system with offline joint mapping optimization and runtime accessor-aware scheduling achieves up to 2.32x higher LLM inference throughput than prior unified NPU-PIM memory designs in simulation.
cublas docs,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AR 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM
A dual-view memory system with offline joint mapping optimization and runtime accessor-aware scheduling achieves up to 2.32x higher LLM inference throughput than prior unified NPU-PIM memory designs in simulation.