A domain-decomposition scheme sizes subdomains to GPU shared memory to remove synchronization from sparse triangular solves, reporting 10.7x and 3.2x speedups for triangular solves and ILU0-BiCGSTAB on the AMD MI210.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.PF 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Mapping Sparse Triangular Solves to GPUs via Fine-grained Domain Decomposition
A domain-decomposition scheme sizes subdomains to GPU shared memory to remove synchronization from sparse triangular solves, reporting 10.7x and 3.2x speedups for triangular solves and ILU0-BiCGSTAB on the AMD MI210.