A single 256-core shared-memory cluster outperforms configurations of many smaller clusters by up to 2x for memory-bound kernels and 24% for compute-bound kernels, with a soft barrier adding further speedups.
NVIDIA Ampere GA102 GPU architecture,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimizing Scalable Multi-Cluster Architectures for Next-Generation Wireless Sensing and Communication
A single 256-core shared-memory cluster outperforms configurations of many smaller clusters by up to 2x for memory-bound kernels and 24% for compute-bound kernels, with a soft barrier adding further speedups.