REVIEW 3 major objections 5 minor 47 references
A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Enumerate-and-sparse-coarsen, a compiler transformation that groups GPU threads around sparsity patterns, accelerates sparse matrix-matrix multiplication by 1.76x over cuSPARSE and 2.28x over cuBLAS on an A100.
desk verdict A novel GPU sparse-matmul transformation with a plausible mechanism, but the paper's own speedup numbers contradict one another and the autotuning protocol overfits, so the headline result is not yet well-defined. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the enumerate-and-sparse-coarsen lowering, applied as a source-to-source transformation from Python to CUDA. Enumeration unrolls the row loop of A and groups conditionals into enumerated blocks, one per possible nonzero pattern of a UFi-row tile, then maps thread blocks to those patterns so that each block executes a conditional-free, fixed-size workload. Sparse coarsening maps 32 threads across columns of B, unrolls the k loop by UFk, and assigns multiple rows of C per thread, so the same elements of B are reused across several FMA operations inside a thread. The load-bearing mechanism is this in-thread reuse of B: it raises FMA/cycle and relieves the load-store unit, while the enumeration phase provides the balanced, divergence-free distribution that makes coarsening safe. A final data-transformation step stores the nonzeros of A in ANNZ, an order matching the schedule, keeping all memory accesses coalesced.
What would settle it
Run the single best schedule from the paper's Table 2, without retuning, on held-out pruned neural-network matrices from outside the benchmark collection and on a different GPU generation; if the geometric mean speedup over cuSPARSE and cuBLAS drops to roughly 1x or the method loses on most matrices, the central acceleration claim would be shown not to generalize.
Extended reading notes
Core claim
The central claim is that the main obstacle to fast SPMM on GPUs is not only load imbalance but underused register and cache reuse of the dense matrix when multiplying a sparse matrix A with a dense matrix B. The transformation unrolls rows of A into enumerated blocks, each corresponding to a distinct nonzero pattern in a tile, and maps thread blocks to pattern-specific tiles so that conditionals disappear from thread blocks and work is balanced. It then coarsens each thread to compute several FMAs per loaded B value across multiple rows of A and columns of B, raising FMA/cycle and load-store utilization. Finally it rewrites A into a compressed array ordered by the schedule so all accesses to A, B, and C are coalesced, with atomics at tile boundaries for correctness. Across 3,608 matrices from pruned ResNet50 and Transformer models, the paper reports geometric mean speedups of 1.76x over cuSPARSE and 2.28x over cuBLAS on an A100.
Load-bearing premise
The reported speedups rest on a profiler-based tuner picking four schedule parameters from sweeps over the benchmark matrices; if those choices are overfit to that collection, the gains may not transfer to new models, new column counts, or other GPUs.
Editorial extensions
If this is right
- Unstructured sparse neural networks can be accelerated on commodity GPUs without relying on block-structured sparsity or tensor cores, as long as the number of columns of B is modest.
- Sparse compiler transformations can target register reuse as a first-class goal alongside load balance and coalescing, rather than treating thread-level optimization as secondary.
- A per-architecture tuning step for UFi, UFk, WarpTile, and ThreadBlockSize is needed to realize the reported gains; the paper estimates that a single fixed schedule would lose almost 10% of the Table 1 performance.
- The generated compressed storage is more compact than CSR for matrices with sparsity roughly 50-80%, so the transformation can save memory as well as time in that range.
- Speedups grow as sparsity increases, making the method most attractive for aggressively pruned transformer and ResNet layers.
Reading between the lines
- If the schedule parameters are as architecture- and workload-sensitive as the tuning section suggests, the practical value of the method will depend on how well a schedule trained on one benchmark collection transfers to new models and future GPU generations; a learned cost model could replace per-matrix profiling.
- The same enumeration-and-coarsening recipe could be extended to sparse attention patterns, where B is itself sparse or has irregular column counts, since the paper's evaluation is limited to dense B.
- For sparsities above roughly 98%, the paper's own storage measurements show CSR becoming more compact, so a production system would likely choose between the two formats per layer rather than using one everywhere.
- Combining the transformation with tensor-core execution for the dense side of the product, or with coarse-grained structured pruning, might extend the speedups to larger bCols where tensor cores currently win.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a compiler transformation, enumerate-and-sparse-coarsen, for sparse matrix-matrix multiplication (SPMM) on GPUs. The transformation enumerates sparsity patterns of the sparse matrix A through unrolling and thread-block mapping, then applies thread coarsening with parameters UFi, UFk, WarpTile, and ThreadBlockSize to increase register reuse, improve load balance, and reduce thread divergence. The generated CUDA code is evaluated on matrices from the Deep Learning Matrix Collection (DLMC) on an NVIDIA A100 GPU, comparing against cuBLAS (dense), cuBLAS with tensor cores, and cuSPARSE (CSR). The paper reports geometric mean speedups ranging from about 1.4x to 2.3x depending on the baseline and bCol size, and includes an ablation study, a sparsity sweep, storage-size analysis, and a compile-time analysis.
Significance. If the claims hold, the idea of combining enumeration of sparsity patterns with register-level thread coarsening is a plausible and potentially useful technique for improving GPU inference of unstructured sparse neural networks. The ablation in Section 4.3.1 shows monotonic performance gains from adding enumeration and coarsening, which is encouraging. The paper also includes a sparsity sweep and a storage-efficiency analysis. However, the empirical claims as written are internally inconsistent, and the tuning methodology appears to confound schedule selection with performance evaluation. Because the central contribution is a quantitative speedup claim, these issues must be resolved before the results can be relied upon.
major comments (3)
- [Abstract, Section 1, Section 4.2, Section 6, Table 1] The reported aggregate speedup numbers are mutually inconsistent. The abstract claims 1.84x and 2.27x against cuBLAS and cuSPARSE, respectively; Section 1 claims 1.47x and 1.7x; Section 4.2 claims 1.76x and 2.28x over cuSparse and cuBlas, respectively; and Section 6 claims 1.76x and 2.28x over cuBLAS and cuSparse, respectively. The per-bCol geometric means in Table 1 combine to approximately 1.47x vs cuBLAS and 1.74x vs cuSparse. No single set of runs can support all of these statements. The authors must choose one consistent set of numbers and specify exactly which schedule configuration, which bCol range, and which baselines each number refers to.
- [Section 3.4 and Section 4.3.2] The evaluation protocol does not separate schedule selection from performance measurement. The tuning sweep in Section 3.4 is performed on the same DLMC matrices used for the evaluation in Section 4, and Section 4.3.2 states that using one fixed schedule would lose almost 10% of the performance reported in Table 1. This implies that the headline speedups incorporate schedules fitted to the evaluation dataset. To make the central claim reproducible and generalizable, the authors should either report the performance of a single fixed schedule as the main result or use a held-out set of matrices that does not participate in tuning.
- [Table 1 and Table 2] The relationship between Table 1 and Table 2 is not explained. Table 1 reports per-bCol geometric mean speedups (e.g., 1.85/1.78/1.60 vs cuSparse and 1.40/1.48/1.54 vs cuBLAS), while Table 2 reports the speedup 'achieved with that schedule' (1.73/1.69/1.62 vs cuSPARSE and 1.32/1.38/1.55 vs cuBLAS). The text does not say whether Table 1 uses the per-matrix best schedule, the per-bCol best schedule, or some other protocol, and the aggregate figures quoted in Sections 4.2 and 6 do not match either table. The evaluation protocol needs to be defined precisely so that every reported speedup can be derived from the presented data.
minor comments (5)
- [Listing 6, Section 3.3.2] Listing 6 line 19 uses the variable 'c10', but no such accumulator is declared; the intended variable is likely 'c01'. The same typo appears in Listing 7 line 21.
- [Listing 3 and Listing 4, Section 3.2.1] The k-loop is written as 'for (int k = 0; j < K; k++)' in Listings 3 and 4, with the loop bound using 'j' instead of 'k'.
- [Section 4.3.1 and Table 1] The text states that the transformation outperforms baselines in '60% to 100%' of matrices, but Table 1 reports 59.62% for the cuBLASTC baseline at bCol=128.
- [Figure 1 caption] The caption says 'All stores are atomic and not shown in the code,' but Listings 6 and 7 explicitly show AtomicAdd operations.
- [Abstract] The abstract contains a grammatical error: 'across a columns of matrix B (bCols)' should read 'across the number of columns of matrix B (bCols)'.
Circularity Check
Headline speedups are produced by schedules tuned on the same DLMC matrices used for evaluation, so the quantitative claim is partly fitted rather than independently predicted.
-
fitted input called prediction
[Section 3.4 (Scheduler & Tuner), Section 4.2 (Performance Evaluation), Section 4.3.2 (Cost model)]
"This figure is generated by running a sweep of values for all unique dimension and the 5 sparsity ranges in DLMC, which makes 60 matrices. ... Table 2 shows the best configuration and the speedup achieved with that schedule for all unstructured DLMC matrices. As can be seen, if only one schedule were to be chosen, we would lose almost 10% in performance compared to the results presented in Table 1. ... Enumerate-and-sparse-coarsen achieves a geometric mean speedup of 1.76 and 2.28 over cuSparse and cuBlas, respectively, for these DLMC matrices."
The schedule parameters UFi, UFk, WarpTile, and ThreadBlockSize are selected by profiling the DLMC dataset (Section 3.4), and the evaluation in Section 4.2 and Table 1 measures speedups on the same DLMC matrices using the best per-bCol schedule (as confirmed by Section 4.3.2's statement that a single fixed schedule would lose about 10%). The reported geometric-mean speedup is therefore not an independent, parameter-free property of the transformation: it is an oracle result in which the schedule is chosen to maximize performance on the evaluation set. The quantitative claim is partly constituted by the tuner's fit to the test data rather than a prediction validated on held-out workloads.
full rationale
The transformation's mechanism—enumeration plus sparse coarsening to improve register and cache reuse—is described concretely with code listings and compared stage-by-stage (Base, Base+Enumeration, Base+Enumeration+Coarsening) against cuBLAS, so the central algorithmic idea does not reduce to a self-citation or to a definitional identity. Nor is any load-bearing uniqueness theorem or ansatz imported solely from the authors' prior work. The circularity concern is confined to the empirical claim: the profiler-based scheduler (Section 3.4) selects schedules using DLMC, and Section 4.2 reports the speedups on DLMC with those tuned schedules. Section 4.3.2's admission that one fixed schedule would cost about 10% of the Table 1 performance confirms that the reported gains include per-dataset schedule selection. This fits the fitted-input-called-prediction pattern and warrants a 6, not a higher score, because the underlying transformation can be evaluated independently and the paper discloses the tuning step. The arithmetic inconsistency among the abstract (1.84x-2.27x), Section 4.2 / conclusion (1.76x and 2.28x), and Table 1 (~1.47x and ~1.74x) is a correctness/reproducibility defect, not a circularity defect, and is not counted in this score.
Assumptions & free parameters
free parameters (4)
- UFi (unroll factor of i) =
e.g., 4 for bCol 32, 3 for bCol 64/128 in Table 2; varies per matrix
- UFk (unroll factor of k) =
e.g., 7 for bCol 32/64, 8 for bCol 128 in Table 2
- WarpTile =
1, 2, 2 for bCol 32, 64, 128 in Table 2
- ThreadBlockSize =
32, 32, 64 for bCol 32, 64, 128 in Table 2
assumptions (4)
- domain assumption The input matrix A is available in dense form or can be materialized, since dataTransformer iterates over dense A with O(M*K) cost.
- domain assumption Element-wise sparse matrices in DLMC have a diverse range of patterns with nearly uniform frequency, so creating thread blocks for all enumerated patterns balances load.
- domain assumption Floating-point atomic add is acceptable output semantics for inference, so bitwise reproducibility is not required.
- domain assumption Neural network weights are static during inference, allowing the compressed format to be built once and reused.
Cite this review
Pith. "Pith review of A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs." pith.science (2026). https://pith.science/paper/5HXWCBYO
@misc{pith2026250615174,
author = {Pith},
title = {Pith review of: A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs},
year = {2026},
howpublished = {\url{https://pith.science/paper/5HXWCBYO}},
note = {Machine review of arXiv:2506.15174}
}
abstract
Sparse data structures are commonly used in neural networks to reduce the memory footprint. These data structures are compact but cause irregularities such as random memory accesses, which prevent efficient use of the memory hierarchy. GPUs are a common platform for machine learning practitioners, but running compact data structures on these devices often leads to slow-downs due to inefficient use of computing and memory resources. This paper proposes a new compiler transformation, enumerate-and-sparse-coarsen, that accelerates sparse matrix-matrix multiplication (SPMM) on GPU devices. The transformation increases data reuse in registers and caches while creating more balanced workloads for GPU computing resources. The transformation is tested on sparse neural networks in convolutional and transformer models. On an A100 GPU and across a columns of matrix B (bCols) in $ A \times B = C$ from range of 32 to 128, the transformation yields a geometric mean speedup of 1.84$\times$ to 2.27$\times$ compared to cuBLAS and cuSPARSE baselines, respectively.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[17]
Ahan Gupta, Yueming Yuan, Devansh Jain, Yuhao Ge, David Aponte, Yanqi Zhou, and Charith Mendis. 2025. SPLAT: A Framework for Optimised GPU Code-Generation for SParse reguLar ATtention. Proc. ACM Program. Lang. 9, OOPSLA1, Article 138 (April 2025), 29 pages. https://doi.org/10.1145/3720503
doi:10.1145/3720503 2025
-
[1]
Willow Ahrens, Teodoro Fields Collin, Radha Patel, Kyle Deeds, Changwan Hong, and Saman Amarasinghe. 2024. Finch: Sparse and Structured Array Programming with Control Flow. arXiv preprint arXiv:2404.16730 (2024)
arXiv 2024
-
[2]
Prithayan Barua, Jun Shirako, and Vivek Sarkar. 2018. Cost-driven thread coars- ening for GPU kernels. In Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques. 1–14
work page 2018
-
[3]
Patrick Carribault, Albert Cohen, and William Jalby. 2005. Deep Jam: Conver- sion of coarse-grain parallelism to instruction-level and vector parallelism for irregular applications. In 14th International Conference on Parallel Architectures and Compilation Techniques (PACT’05). IEEE, 291–300
work page 2005
-
[4]
2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al. 2018.{TVM}: An automated{End-to-End} optimizing compiler for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 578–594
2018
-
[5]
Kazem Cheshmi. 2022. Transforming Sparse Matrix Computations . Ph. D. Disser- tation. University of Toronto (Canada)
work page 2022
-
[6]
Kazem Cheshmi, Zachary Cetinic, and Maryam Mehri Dehnavi. 2022. Vectorizing sparse matrix computations with partially-strided codelets. InSC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–15
work page 2022
-
[7]
Kazem Cheshmi, Shoaib Kamil, Michelle Mills Strout, and Maryam Mehri Dehnavi. 2017. Sympiler: transforming sparse matrix codes by decoupling sym- bolic analysis. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . 1–13
work page 2017
Show all 47 references
-
[8]
Mohammad Mahdi Salehi Dezfuli and Kazem Cheshmi. 2024. Improving Locality in Sparse and Dense Matrix Multiplications. arXiv preprint arXiv:2407.00243 (2024)
2024 arXiv
-
[9]
Adhitha Dias, Kirshanthan Sundararajah, Charitha Saumya, and Milind Kulka- rni. 2022. SparseLNR: accelerating sparse tensor computations using loop nest restructuring. In Proceedings of the 36th ACM International Conference on Super- computing. 1–14
2022
-
[10]
Qiang Fu, Thomas B Rolinger, and H Howie Huang. 2024. JITSPMM: Just-in-Time Instruction Generation for Accelerated Sparse Matrix-Matrix Multiplication. In 2024 IEEE/ACM International Symposium on Code Generation and Optimization (CGO). IEEE, 448–459
2024
-
[11]
Trevor Gale, Erich Elsen, and Sara Hooker. 2019. The state of sparsity in deep neural networks. arXiv preprint arXiv:1902.09574 (2019)
2019 arXiv
-
[12]
Trevor Gale, Matei Zaharia, Cliff Young, and Erich Elsen. 2020. Sparse gpu kernels for deep learning. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–14
2020
-
[13]
Mahdi Ghorbani, Emilien Bauer, Tobias Grosser, and Amir Shaikhha. 2024. Com- pressing Structured Tensor Algebra. arXiv preprint arXiv:2407.13726 (2024)
2024 arXiv
-
[14]
Scott Gray, Alec Radford, and Diederik P Kingma. 2017. Gpu kernels for block- sparse weights. arXiv preprint arXiv:1711.09224 3, 2 (2017), 2
2017 arXiv
-
[15]
Joseph L Greathouse and Mayank Daga. 2014. Efficient sparse matrix-vector multiplication on GPUs using the CSR storage format. InSC’14: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 769–780
2014
-
[16]
Yue Guan, Changming Yu, Yangjie Zhou, Jingwen Leng, Chao Li, and Minyi Guo. 2024. Fractal: Joint Multi-Level Sparse Pattern Tuning of Accuracy and Performance for DNN Pruning. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Lang...
2024
-
[18]
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste
-
[19]
Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, Süreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, Ümit V Çatalyürek, Srini- vasan Parthasarathy, and P Sadayappan. 2018. Efficient sparse-matrix multi- vector product on gpus. In Proceedings of the 27th I...
2018
-
[20]
Changwan Hong, Aravind Sukumaran-Rajam, Israt Nisa, Kunal Singh, and P Sadayappan. 2019. Adaptive sparse tiling for sparse matrix multiplication. InPro- ceedings of the 24th Symposium on Principles and Practice of Parallel Programming . 300–314
2019
-
[21]
Marcos Horro, Louis-Noël Pouchet, Gabriel Rodríguez, and Juan Touriño. 2022. Custom High-Performance Vector Code Generation for Data-Specific Sparse Computations. In Proceedings of the International Conference on Parallel Architec- tures and Compilation Techniques. 160–171
2022
-
[22]
Yuka Ikarashi, Gilbert Louis Bernstein, Alex Reinking, Hasan Genc, and Jonathan Ragan-Kelley. 2022. Exocompilation for productive programming of hardware accelerators. In Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementa...
2022
-
[23]
Durk P Kingma, Tim Salimans, and Max Welling. 2015. Variational dropout and the local reparameterization trick. Advances in neural information processing systems 28 (2015)
2015
-
[24]
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amaras- inghe. 2017. The tensor algebra compiler.Proceedings of the ACM on Programming Languages 1, OOPSLA (2017), 1–29
2017
-
[25]
Junqing Lin, Honghe Zhang, Xiaolong Shi, Jingwei Sun, Xianzhi Yu, Jun Yao, and Guangzhong Sun. 2023. EC-SpMM: Efficient Compilation of SpMM Kernel on GPUs. In Proceedings of the 52nd International Conference on Parallel Processing . 21–30
2023
-
[26]
Yiqian Liu, Noushin Azami, Corbin Walters, and Martin Burtscher. 2022. The Indigo Program-Verification Microbenchmark Suite of Irregular Parallel Code Patterns. In 2022 IEEE International Symposium on Performance Analysis of Sys- tems and Software (ISPASS). IEEE, 24–34
2022
-
[27]
Christos Louizos, Max Welling, and Diederik P Kingma. 2017. Learning sparse neural networks through 𝐿_0 regularization. arXiv preprint arXiv:1712.01312 (2017)
2017 arXiv
-
[28]
Alberto Magni, Christophe Dubach, and Michael O’Boyle. 2014. Automatic optimization of thread-coarsening for graphics processors. In Proceedings of the 23rd international conference on Parallel architectures and compilation . 455–466
2014
-
[29]
Duane Merrill and Michael Garland. 2016. Merge-based parallel sparse matrix- vector multiplication. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 678–689
2016
-
[30]
Mahdi Soltan Mohammadi, Kazem Cheshmi, Maryam Mehri Dehnavi, Anand Venkat, Tomofumi Yuki, and Michelle Mills Strout. 2019. Extending index-array properties for data dependence analysis. In Languages and Compilers for Parallel Computing: 31st International Workshop, LCPC 2018, ...
2019
-
[31]
Erdal Mutlu, Ruiqin Tian, Bin Ren, Sriram Krishnamoorthy, Roberto Gioiosa, Jacques Pienaar, and Gokcen" Kestor. 2022. COMET: A Domain-Specific Compila- tion of High-Performance Computational Chemistry. InLanguages and Compilers for Parallel Computing, Barbara Chapman and José ...
2022
-
[32]
Maxim Naumov, L Chien, Philippe Vandermersch, and Ujval Kapasi. 2010. Cus- parse library. In GPU Technology Conference, Vol. 12
2010
-
[33]
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013. Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. Acm Sigplan Notices 48, 6 (2013), 519–530
2013
-
[34]
Mohammad Mahdi Salehi and Kazem Cheshmi. 2025. Loop Fusion in Matrix Multiplications with Sparse Dependence. In International Conference on Super- computing (ICS ’25) (ICS ’25) . Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3524059.3532386
2025
-
[35]
Nicolai Stawinoga and Tony Field. 2018. Predictable thread coarsening. ACM Transactions on Architecture and Code Optimization (TACO) 15, 2 (2018), 1–26
2018
-
[36]
Michelle Mills Strout, Mary Hall, and Catherine Olschanowsky. 2018. The sparse polyhedral framework: Composing compiler-generated inspector-executor code. Proc. IEEE 106, 11 (2018), 1921–1934. A Novel Compiler Transformation for Fast Sparse Matrix Multiplication in GPUs
2018
-
[37]
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018. Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions. arXiv preprint arXiv:180...
2018 arXiv
-
[38]
Vasily Volkov. 2010. Better performance at lower occupancy. InProceedings of the GPU technology conference, GTC , Vol. 10. San Jose, CA, 16
2010
-
[39]
Lucas Wilkinson, Kazem Cheshmi, and Maryam Mehri Dehnavi. 2023. Register Tiling for Unstructured Sparsity in Neural Network Inference. Proceedings of the ACM on Programming Languages 7, PLDI (2023), 1995–2020
2023
-
[40]
Jaeyeon Won, Charith Mendis, Joel S Emer, and Saman Amarasinghe. 2023. WACO: learning workload-aware co-optimization of the format and schedule of a sparse tensor program. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Language...
2023
-
[41]
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023. Flash-llm: Enabling cost- effective and highly-efficient large generative model inference with unstructured sparsity. arXiv preprint arXiv:2309.10285 (2023)
2023 arXiv
-
[42]
Carl Yang, Aydın Buluç, and John D. Owens. 2018. Design Principles for Sparse Matrix Multiplication on the GPU. In Euro-Par 2018: Parallel Processing: 24th International Conference on Parallel and Distributed Computing, Turin, Italy, Au- gust 27 - 31, 2018, Proceedings (Turin,...
2018 doi
-
[43]
Zihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen, and Luis Ceze. 2023. Sparsetir: Composable abstractions for sparse compilation in deep learning. In Proceedings of the 28th ACM International Conference on Architectural Support for Program- ming Languages and Operating Systems, ...
2023
-
[44]
Xin You, Changxi Liu, Hailong Yang, Pengbo Wang, Zhongzhi Luan, and De- pei Qian. 2022. Vectorizing spmv by exploiting dynamic regular patterns. In Proceedings of the 51st International Conference on Parallel Processing . 1–12
2022
-
[45]
Ningxin Zheng, Bin Lin, Quanlu Zhang, Lingxiao Ma, Yuqing Yang, Fan Yang, Yang Wang, Mao Yang, and Lidong Zhou. 2022. {SparTA}:{Deep-Learning} Model Sparsity via{Tensor-with-Sparsity-Attribute}. In16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . 213–232
2022
-
[46]
Michael Zhu and Suyog Gupta. 2017. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878 (2017). Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009
2017 arXiv
-
[2021]
Journal of Machine Learning Research 22, 241 (2021), 1–124
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks. Journal of Machine Learning Research 22, 241 (2021), 1–124
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.