A trace of QWEN-0.5B CPU inference shows 98.06% of addresses are accessed once per generated token, while simulated prefetchers cut the L2 cache miss rate from 69.9% to 30.8%.
Understanding Performance Implications of LLM Inference on CPUs ,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
A trace of QWEN-0.5B CPU inference shows 98.06% of addresses are accessed once per generated token, while simulated prefetchers cut the L2 cache miss rate from 69.9% to 30.8%.