A trace of QWEN-0.5B CPU inference shows 98.06% of addresses are accessed once per generated token, while simulated prefetchers cut the L2 cache miss rate from 69.9% to 30.8%.
Gpt-4 technical report,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Memory Access Characterization of Large Language Models in CPU Environment and its Potential Impacts
A trace of QWEN-0.5B CPU inference shows 98.06% of addresses are accessed once per generated token, while simulated prefetchers cut the L2 cache miss rate from 69.9% to 30.8%.