Pith. sign in

REVIEW 1 cited by

Characterizing Optimizations to Memory Access Patterns using Architecture-Independent Program Features

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.06064 v1 pith:GS7PP3MF submitted 2020-03-12 cs.DC

classification cs.DC
keywords memoryopenclaccessaiwcmetricpatternsarchitectureslocality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High-performance computing developers are faced with the challenge of optimizing the performance of OpenCL workloads on diverse architectures. The Architecture-Independent Workload Characterization (AIWC) tool is a plugin for the Oclgrind OpenCL simulator that gathers metrics of OpenCL programs that can be used to understand and predict program performance on an arbitrary given hardware architecture. However, AIWC metrics are not always easily interpreted and do not reflect some important memory access patterns affecting efficiency across architectures. We propose a new metric of parallel spatial locality -- the closeness of memory accesses simultaneously issued by OpenCL work-items (threads). We implement the parallel spatial locality metric in the AIWC framework, and analyse gathered results on matrix multiply and the Extended OpenDwarfs OpenCL benchmarks. The differences in the observed parallel spatial locality metric across implementations of matrix multiply reflect the optimizations performed. The new metric can be used to distinguish between the OpenDwarfs benchmarks based on the memory access patterns affecting their performance on various architectures. The improvements suggested to AIWC will help HPC developers better understand memory access patterns of complex codes and guide optimization of codes for arbitrary hardware targets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Memory Access Vectors: Improving Sampling Fidelity for CPU Performance Simulations

    cs.AR 2025-06 conditional novelty 6.0 of 10

    Combining SimPoint basic-block vectors with memory-access-frequency vectors lifts projected performance accuracy for 523.xalancbmk_r from 80% to 98% on a 192-core AmpereOne SoC.

Pith tools