First distributed performance-portable NUFFT scales to 1024 GPUs on heterogeneous systems and supports large particle-in-Fourier plasma simulations.
Richards, and Laxmikant V
7 Pith papers cite this work, alongside 6 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 7roles
background 2polarities
background 2representative citing papers
PackSELL packs delta-encoded indices and values into single words with tunable bit allocation, delivering up to 1.63x faster FP16 SpMV and FP32-accurate performance exceeding FP16 cuSPARSE while reducing memory traffic.
Portable Ewald summation algorithms for Stokes flow achieve ~8M particles/sec on H200 GPU with a novel P2G kernel providing 16x speedup and good multi-GPU scaling.
GNN-DRL cloud schedulers for DAG workflows degrade under topology shifts because structural mismatches disrupt message passing and policy generalization.
Extends ItoyoriFBC with promise-future synchronization via MPI one-sided communication for dynamic dependencies in AMT runtimes, shown with HLU achieving 15.6x speedup on 16 nodes.
Defines differentiable weak distance on SE(3) for surface measures via Sobolev norms and shows local optimization with trust-region methods and NUFFT gradients.
A survey categorizing vendor mechanisms and user-level libraries for GPU-centric communication within and across nodes, with discussion of benefits, challenges, and open questions.
citing papers explorer
-
A Performance-Portable, Massively Parallel Distributed Nonuniform FFT
First distributed performance-portable NUFFT scales to 1024 GPUs on heterogeneous systems and supports large particle-in-Fourier plasma simulations.
-
PackSELL: A Sparse Matrix Format for Precision-Agnostic High-Performance SpMV
PackSELL packs delta-encoded indices and values into single words with tunable bit allocation, delivering up to 1.63x faster FP16 SpMV and FP32-accurate performance exceeding FP16 cuSPARSE while reducing memory traffic.
-
A performance portable fast Ewald summation for Stokes flow
Portable Ewald summation algorithms for Stokes flow achieve ~8M particles/sec on H200 GPU with a novel P2G kernel providing 16x speedup and good multi-GPU scaling.
-
On the Role of DAG topology in Energy-Aware Cloud Scheduling : A GNN-Based Deep Reinforcement Learning Approach
GNN-DRL cloud schedulers for DAG workflows degrade under topology shifts because structural mismatches disrupt message passing and policy generalization.
-
Promise-Future Synchronization for Cluster Asynchronous Many-Task Runtimes via MPI One-Sided Communication
Extends ItoyoriFBC with promise-future synchronization via MPI one-sided communication for dynamic dependencies in AMT runtimes, shown with HLU achieving 15.6x speedup on 16 nodes.
-
Local optimization of weak distance between compact surfaces on special Euclidean group
Defines differentiable weak distance on SE(3) for surface measures via Sobolev norms and shows local optimization with trust-region methods and NUFFT gradients.
-
The Landscape of GPU-Centric Communication
A survey categorizing vendor mechanisms and user-level libraries for GPU-centric communication within and across nodes, with discussion of benefits, challenges, and open questions.