FlashSparse uses the identity A×B=(B^T×A^T)^T to reduce sparse matrix multiplication's nonzero-vector granularity from 16×1 to 8×1 on tensor cores, reporting SOTA speedups.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
FlashSparse uses the identity A×B=(B^T×A^T)^T to reduce sparse matrix multiplication's nonzero-vector granularity from 16×1 to 8×1 on tensor cores, reporting SOTA speedups.