Combining SimPoint basic-block vectors with memory-access-frequency vectors lifts projected performance accuracy for 523.xalancbmk_r from 80% to 98% on a 192-core AmpereOne SoC.
NPS: A Framework for Accurate Program Sampling Using Graph Neural Network
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
With the end of Moore's Law, there is a growing demand for rapid architectural innovations in modern processors, such as RISC-V custom extensions, to continue performance scaling. Program sampling is a crucial step in microprocessor design, as it selects representative simulation points for workload simulation. While SimPoint has been the de-facto approach for decades, its limited expressiveness with Basic Block Vector (BBV) requires time-consuming human tuning, often taking months, which impedes fast innovation and agile hardware development. This paper introduces Neural Program Sampling (NPS), a novel framework that learns execution embeddings using dynamic snapshots of a Graph Neural Network. NPS deploys AssemblyNet for embedding generation, leveraging an application's code structures and runtime states. AssemblyNet serves as NPS's graph model and neural architecture, capturing a program's behavior in aspects such as data computation, code path, and data flow. AssemblyNet is trained with a data prefetch task that predicts consecutive memory addresses. In the experiments, NPS outperforms SimPoint by up to 63%, reducing the average error by 38%. Additionally, NPS demonstrates strong robustness with increased accuracy, reducing the expensive accuracy tuning overhead. Furthermore, NPS shows higher accuracy and generality than the state-of-the-art GNN approach in code behavior learning, enabling the generation of high-quality execution embeddings.
citation-role summary
citation-polarity summary
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Memory Access Vectors: Improving Sampling Fidelity for CPU Performance Simulations
Combining SimPoint basic-block vectors with memory-access-frequency vectors lifts projected performance accuracy for 523.xalancbmk_r from 80% to 98% on a 192-core AmpereOne SoC.