Multi-strided memory access, created by loop unrolling over non-contiguous axes, improves hardware prefetcher utilization and speeds up dense memory-bound kernels by up to 2.18x over single-strided code and 2.99x over Intel MKL.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.PF 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Strided Access Patterns to Boost Hardware Prefetching
Multi-strided memory access, created by loop unrolling over non-contiguous axes, improves hardware prefetcher utilization and speeds up dense memory-bound kernels by up to 2.18x over single-strided code and 2.99x over Intel MKL.