LEO is a cross-vendor GPU stall root-cause analyzer that traces stalled instructions backward through register and synchronization dependencies, achieving 1.73–1.82x geometric-mean speedups across 21 workloads on NVIDIA, AMD, and Intel GPUs.
Title resolution pending
3 Pith papers cite this work, alongside 66 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.DC 3years
2026 3roles
background 1polarities
background 1representative citing papers
OpenMP port of gPLUTO achieves comparable performance to OpenACC on NVIDIA but is 3x slower at application level and up to 10x at kernel level on AMD MI250X, driven by strided memory accesses, latency bounds, and C++ abstraction overheads.
New hardware-usage-based similarity metrics can identify matching computational kernels between proxy applications and performance suites on both CPU and GPU systems.
citing papers explorer
-
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
LEO is a cross-vendor GPU stall root-cause analyzer that traces stalled instructions backward through register and synchronization dependencies, achieving 1.73–1.82x geometric-mean speedups across 21 workloads on NVIDIA, AMD, and Intel GPUs.
-
On the Limits of Performance Portability in Directive-Based GPU Programming
OpenMP port of gPLUTO achieves comparable performance to OpenACC on NVIDIA but is 3x slower at application level and up to 10x at kernel level on AMD MI250X, driven by strided memory accesses, latency bounds, and C++ abstraction overheads.
-
On Similarity of Computational Kernels in our Codes and Proxies
New hardware-usage-based similarity metrics can identify matching computational kernels between proxy applications and performance suites on both CPU and GPU systems.