A compiler pipeline (flatten, crush, graph-match to 2:4) retargets NVIDIA sparse tensor cores to scientific stencil computation, reporting average 3.1x speedups over the previous best stencil-on-tensor-core system.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation
A compiler pipeline (flatten, crush, graph-match to 2:4) retargets NVIDIA sparse tensor cores to scientific stencil computation, reporting average 3.1x speedups over the previous best stencil-on-tensor-core system.