REVIEW 3 major objections 4 minor 53 references
SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read SparStencil claims that scientific stencil computations can be reshaped into the 2:4 sparse format that AI tensor cores require, and that this yields up to 7.1x speedup over state-of-the-art stencil frameworks.
desk verdict Genuinely new and useful, but the optimality theorem has a gap and there is no artifact; the 3.1x average is a claim to verify, not a result to trust. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are the k-staircase matrix property and the Permutation Invariant Transformation (PIT). A matrix is k-staircase when its nonzeros in column c appear only in rows r through r+k-1, and the crushed matrix is self-similar: both the block-level structure and each nonzero block satisfy this property. This yields the non-conflict theorem that columns at least k hops apart never share a nonzero row, which turns 2:4 conversion into a conflict-free matching problem on a column conflict graph. PIT is the shared permutation of A's columns and B's rows that keeps the product A×B unchanged while reordering nonzeros; the hierarchical two-level matching algorithm finds a permutation that satisfies 2:4 alignment with minimal zero-column padding, with the Blossom algorithm as a fallback for patterns that violate the staircase assumption.
What would settle it
Run the pipeline on a stencil whose boundary handling or layout parameters break the diagonal-offset property, then measure the ratio of inserted zero columns to nonzero columns; if that ratio becomes large enough to erase the throughput advantage predicted by the analytical model, the central claim fails.
Extended reading notes
Core claim
The paper's central claim is that stencil computations, whose clustered sparsity has been treated as an obstacle to tensor-core acceleration, can be restructured into the hardware-required 2:4 sparsity pattern through a sequence of semantic-preserving transformations. The flatten-and-crush pipeline eliminates redundant data accesses and exposes a self-similar staircase sparsity structure; the Permutation Invariant Transformation then permutes columns of the matrix and rows of its partner so that a graph-theoretic matching algorithm can group nonzeros into conflict-free 1:2 or 0:2 pairs. The authors prove that, under the k-staircase property, their hierarchical two-level matching is valid and inserts the minimum number of zero columns, and they report that the resulting sparse MMA kernels consistently outperform prior dense tensor-core stencil methods on an A100 GPU.
Load-bearing premise
The whole speedup rests on the crushed matrix keeping its self-similar staircase shape after the flatten-and-crush step; if real kernels break that shape, the 2:4 conversion needs more zero columns and the reported gains shrink.
Editorial extensions
If this is right
- Stencil kernels compiled through this pipeline can exceed dense tensor-core stencil methods on A100 hardware, with an average 3.1x speedup and a peak of 7.1x over the state-of-the-art baseline.
- The transformation overhead, including metadata, lookup tables, and restructuring, is mostly amortized over the run, staying below roughly 10% of sustained runtime for most kernels.
- Automated layout search over the morphing parameters (r1, r2) can discover high-compute-density configurations that avoid hand-tuned schedules.
- Even on dense FP64 tensor cores that lack hardware sparsity support, the layout morphing and search still yield speedups of 1.11x to 7.13x over prior methods, suggesting the approach is not limited to FP16 sparse hardware.
- Sparse tensor cores deliver roughly 2x the throughput of dense tensor cores for the converted 2:4 matrices, and this mechanism is what drives the reported throughput of up to 156.7 GStencil/s.
Reading between the lines
- If the k-staircase guarantee generalizes beyond the idealized construction, the same flatten-and-crush plus graph-matching recipe could apply to other structured scientific operators with clustered sparsity, such as convolutions or FFT-based kernels, not only stencils.
- The reported sensitivity at small problem sizes, where PIT overhead can produce slowdowns, suggests a crossover point below which a system should fall back to dense execution; exposing that threshold as a compile-time decision is a natural extension.
- Because the conversion relies only on a shared permutation, the method should port to any accelerator with 2:4-style structured sparsity, and future sparse tensor cores with FP64 support would likely amplify the measured gains.
- The existence of the Blossom fallback means correctness does not strictly depend on the staircase assumption; a testable extension is to measure how often real stencil boundaries and layout parameters actually trigger the fallback and how much padding grows when they do.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SparStencil proposes an automated pipeline for mapping stencil computations onto NVIDIA sparse Tensor Cores. It first flattens the stencil into a matrix product and 'crushes' duplicated entries to produce a staircase-sparse matrix A' parameterized by layout choices (r1, r2); it then converts A' to a 2:4-compatible matrix via a permutation, formulated as a minimum zero-column matching problem on a conflict graph and solved by a hierarchical two-level matching algorithm (Algorithm 1); finally, it searches layout parameters with an analytical model and generates CUDA kernels using lookup tables and double buffering. The evaluation on an A100 reports up to 7.1x and on average 3.1x speedup over ConvStencil on 79 kernels, with preprocessing overhead mostly below 10% of runtime.
Significance. If the reported results are correct, SparStencil is a meaningful first demonstration that sparse AI tensor cores can accelerate a non-AI scientific workload class; the graph-theoretic formulation of 2:4 conversion and the automatic code generation are attractive contributions. The PIT identity is standard but used cleanly, and the evaluation breadth (79 kernels, six baselines, overhead breakdown, FP64 extension) is a strength. However, the paper does not ship an artifact, and the optimality proof for the core conversion algorithm has a concrete counterexample, so the central low-padding claim is not yet established.
major comments (3)
- [Section 3.2, Theorem 2] The minimality proof analyzes only a single subgraph G_i and never reasons about the global matching M1 together with M2. This is not a mere proof gap: Algorithm 1 is not optimal on the manuscript's own definitions. Take m=3 blocks of size g=5 with k=3. Algorithm 1 sets s1=max(floor(3/2),3)=3, so no M1 pairs; within each block M2 creates pairs (0,3), (1,4) and one zero column, for three zero columns total. The full conflict graph (columns 0-14, edges between columns at distance less than k) admits a valid matching with seven real-real pairs and one zero column, e.g. (0,12), (1,4), (2,5), (3,6), (7,10), (8,11), (9,13), because every pair has column distance at least 3. Hence p=1 < 3, contradicting Theorem 2(ii). Since the claimed minimal padding is load-bearing for the low overhead and speedup, the theorem and Algorithm 1 need to be revised, or the paper must use the Blossom fallback and report its padding for every evaluated kernel.
- [Section 3.2, Definition 4 and following paragraph] The self-similar k-staircase property is asserted for the crushed matrix A', but the manuscript only illustrates it on a 3x3 stencil (Figures 3-5). The text itself says Algorithm 1's guarantee holds 'under the k-staircase assumption' and falls back to Blossom otherwise. No theorem establishes that arbitrary stencil shapes, boundary handling, and the explored (r1,r2) values produce a k-staircase A'; alternatively, no measured padding statistics are reported for the 79 kernels. Without one of these, the central claim that 2:4 conversion is cheap enough to preserve the speedup is not established. Please add a proof of the property for the supported class, or report per-kernel padding ratios and their effect on MMA count.
- [Section 3.3, Eqs. (6)-(11)] The layout exploration selects (r1,r2) by minimizing an analytical model, but the paper gives no validation of the model. Equations 9 and 10 assume simple formulas for MMA count and shared-memory traffic; no comparison with measured execution times is shown, and no ablation demonstrates that the selected layout coincides with the measured optimum. Since model error would silently select suboptimal layouts and affect the end-to-end speedup, please add predicted-versus-measured parity plots and sensitivity analysis over the explored (r1,r2) grid.
minor comments (4)
- [Table 2] The Problem Size column appears to have too many factors for the stated dimensionality (Heat-2D lists three factors and Heat-3D lists four); clarify which factors are spatial grid sizes and which are time iterations.
- [Section 4.2 and Figure 8] The first use of 'LUT' and 'TS' should be defined in the caption or text; the text says LUT peaks near 30% but the figure is not reproduced with readable labels.
- [Figures 6 and 10] The figures use small fonts and dense legends; please increase readability and provide numerical values or a table for the 79-kernel aggregate.
- [Footnote 1] The footnote announcing acceptance at SC'25 seems out of place in a journal submission; if this is a journal extension, the relationship to the conference version should be stated in the introduction rather than a footnote.
Circularity Check
No significant circularity: the speedup claims are measured against external baselines, and the transformation pipeline is self-contained linear algebra and graph matching.
full rationale
The central claims are empirical: the 7.1x/3.1x speedups reported in Figure 10 are measured against cuDNN, DRStencil, and ConvStencil, not derived from the paper's assumptions, so the primary result is not circular. The transformation correctness chain (PIT identity Eq. 5; conflict graph Definitions 1-3; Theorems 1-2) is an internal proof based on the explicitly stated k-staircase assumption; it does not assume the conclusion. The only self-citation is the analytical model taken from the authors' prior ConvStencil [9] (Eqs. 6-11) to select (r1, r2). This is a minor self-citation, but it is not load-bearing in a circular sense: the model's equations are restated in the paper, the model is not fitted to SparStencil's own measured outputs, and the claimed end-to-end speedups are validated against external benchmarks rather than derived from the model. The acknowledged limitation that the k-staircase property is not proven for all stencils and the Blossom fallback for arbitrary patterns is a correctness/robustness caveat, not a circularity. No step reduces, by construction or by citation, to its own input, so no circularity is found.
Assumptions & free parameters
free parameters (1)
- (r1, r2) layout morphing parameters =
kernel-dependent; chosen by exhaustive search over S via Eq 11 (Section 3.3)
assumptions (4)
- ad hoc to paper The crushed matrix A' has the self-similar k-staircase property (Definition 4) for all supported kernels and layout choices
- domain assumption Sparse TCU hardware correctly processes sub-2:4 patterns by treating explicit zeros as nonzeros, so padded 2:4 groups preserve the result
- domain assumption The analytical performance model (Eqs 6 to 10) ranks layout candidates correctly enough for the exhaustive search in Eq 11 to find near-optimal (r1, r2)
- standard math Permutation invariance of the reduction over k in GEMM (PIT, Eq 5)
Cite this review
Pith. "Pith review of SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation." pith.science (2026). https://pith.science/paper/I6MQJLWG
@misc{pith2026250622969,
author = {Pith},
title = {Pith review of: SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation},
year = {2026},
howpublished = {\url{https://pith.science/paper/I6MQJLWG}},
note = {Machine review of arXiv:2506.22969}
}
read the original abstract
Sparse Tensor Cores offer exceptional performance gains for AI workloads by exploiting structured 2:4 sparsity. However, their potential remains untapped for core scientific workloads such as stencil computations, which exhibit irregular sparsity patterns.This paper presents SparStencil, the first system to retarget sparse TCUs for scientific stencil computations through structured sparsity transformation. SparStencil introduces three key techniques: (1) Adaptive Layout Morphing, which restructures stencil patterns into staircase-aligned sparse matrices via a flatten-and-crush pipeline; (2) Structured Sparsity Conversion, which formulates transformation as a graph matching problem to ensure compatibility with 2:4 sparsity constraints; (3) Automatic Kernel Generation, which compiles transformed stencils into optimized sparse MMA kernels via layout search and table-driven memory mapping. Evaluated on 79 stencil kernels spanning diverse scientific domains, SparStencil achieves up to 7.1x speedup (3.1x on average) over state-of-the-art framework while reducing code complexity and matching or exceeding expert-tuned performance in both compute throughput and memory efficiency.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Yulong Ao, Chao Yang, Xinliang Wang, Wei Xue, Haohuan Fu, Fangfang Liu, Lin Gan, Ping Xu, and Wenjing Ma. 2017. 26 PFLOPS Stencil Computations for Atmospheric Modeling on Sunway TaihuLight. In 2017 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 535–544. doi:10.1109/ IPDPS.2017.9
work page 2017
-
[2]
Krste Asanovic, Ras Bodik, Bryan Christopher Catanzaro, Joseph James Gebis, Parry Husbands, Kurt Keutzer, David A Patterson, William Lester Plishker, John Shalf, Samuel Webb Williams, et al. 2006. The landscape of parallel computing research: A view from berkeley. (2006)
work page 2006
-
[3]
Krste Asanovic, Rastislav Bodik, James Demmel, Tony Keaveny, Kurt Keutzer, John Kubiatowicz, Nelson Morgan, David Patterson, Koushik Sen, John Wawrzynek, David Wessel, and Katherine Yelick. 2009. A View of the Par- allel Computing Landscape. Commun. ACM 52, 10 (oct 2009), 56–67. doi:10. 1145/1562764.1562783
arXiv 2009
-
[4]
Vinayaka Bandishti, Irshad Pananilath, and Uday Bondhugula. 2012. Tiling Stencil Computations to Maximize Parallelism. In Proceedings of the 2012 International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’12). IEEE Computer Society, USA, 1–11. doi:10.1109/ SC.2012.107
work page 2012
- [6]
-
[7]
Uday Bondhugula, Albert Hartono, J. Ramanujam, and P. Sadayappan. 2008. A Practical Automatic Polyhedral Parallelizer and Locality Optimizer. In Proceedings of the 29th ACM SIGPLAN Conference on Programming Language Design and Implementation (Tucson, AZ, USA) (PLDI ’08). Association for Com- puting Machinery, New York, NY, USA, 101–113. doi:10.1145/137558...
-
[8]
Peng Chen, Mohamed Wahib, Shinichiro Takizawa, Ryousei Takano, and Satoshi Matsuoka. 2019. A Versatile Software Systolic Execution Model for GPU Memory-Bound Kernels. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Col- orado) (SC ’19). Association for Computing Machinery, New York, ...
arXiv 2019
-
[9]
Yuetao Chen, Kun Li, Yuhao Wang, Donglin Bai, Lei Wang, Lingxiao Ma, Liang Yuan, Yunquan Zhang, Ting Cao, and Mao Yang. 2024. ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores. In Proceedings of the 29th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (Edinburgh, United Kingdom) (PPoPP ’24)...
Show all 53 references
-
[10]
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014)
2014 arXiv
-
[11]
Oliveira, Juan Gómez-Luna, and Onur Mutlu
Alain Denzler, Rahul Bera, Nastaran Hajinazar, Gagandeep Singh, Geraldo F. Oliveira, Juan Gómez-Luna, and Onur Mutlu. 2023. Casper: Accelerating Stencil Computation using Near-cache Processing. arXiv:2112.14216 [cs.AR] https: //arxiv.org/abs/2112.14216
2023 arXiv
-
[12]
Jack Edmonds. 1965. Paths, Trees, and Flowers.Canadian Journal of Mathematics 17 (1965), 449–467. doi:10.4153/CJM-1965-045-4
1965 doi
-
[13]
Falch and Anne C
Thomas L. Falch and Anne C. Elster. 2014. Register Caching for Stencil Computa- tions on GPUs. In2014 16th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing. 479–486. doi:10.1109/SYNASC.2014.70
2014 doi
-
[14]
Tobias Grosser, Albert Cohen, Paul H. J. Kelly, J. Ramanujam, P. Sadayappan, and Sven Verdoolaege. 2013. Split Tiling for GPUs: Automatic Parallelization Using Trapezoidal Tiles. InProceedings of the 6th Workshopon General Purpose Processor Using Graphics Processing Units (Hou...
2013
-
[15]
Tobias Gysi, Christoph Müller, Oleksandr Zinenko, Stephan Herhut, Eddie Davis, Tobias Wicky, Oliver Fuhrer, Torsten Hoefler, and Tobias Grosser. 2021. Domain- Specific Multi-Level IR Rewriting for GPU: The Open Earth Compiler for GPU- Accelerated Climate Simulation. ACM Trans....
2021 doi
-
[16]
Haozhi Han, Kun Li, Wei Cui, Donglin Bai, Yiwei Zhang, Liang Yuan, Yifeng Chen, Yunquan Zhang, Ting Cao, and Mao Yang. 2025. FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Ten- sor Core Units. In Proceedings of the 30th ACM SIGPLA...
2025
-
[17]
Ramanujam, and P
Tom Henretty, Kevin Stock, Louis-Noël Pouchet, Franz Franchetti, J. Ramanujam, and P. Sadayappan. 2011. Data Layout Transformation for Stencil Computations on Short-Vector SIMD Architectures. In Compiler Construction, Jens Knoop (Ed.). Springer Berlin Heidelberg, Berlin, Heide...
2011
-
[18]
Ramanu- jam, and P
Tom Henretty, Richard Veras, Franz Franchetti, Louis-Noël Pouchet, J. Ramanu- jam, and P. Sadayappan. 2013. A Stencil Compiler for Short-Vector SIMD Architectures. In Proceedings of the 27th International ACM Conference on International Conference on Supercomputing (Eugene, Or...
2013
-
[19]
Sadayappan
Justin Holewinski, Louis-Noël Pouchet, and P. Sadayappan. 2012. High- Performance Code Generation for Stencil Computations on GPU Architectures. In Proceedings of the 26th ACM International Conference on Supercomputing (San Servolo Island, Venice, Italy) (ICS ’12). Association...
2012
-
[20]
Huynh, Z.J
H.T. Huynh, Z.J. Wang, and P.E. Vincent. 2014. High-order methods for com- putational fluid dynamics: A brief review of compact differential formulations on unstructured grids. Computers & Fluids 98 (2014), 209–220. doi:10.1016/j. compfluid.2013.12.007 12th USNCCM mini-symposi...
2014 doi
-
[21]
Mathias Jacquelin, Mauricio Araya–Polo, and Jie Meng. 2022. Scalable Distributed High-Order Stencil Computations. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis. 1–13. doi:10.1109/ SC41404.2022.00035
2022 arXiv
-
[22]
Mellor-Crummey, and R
Guohua Jin, J. Mellor-Crummey, and R. Fowler. 2001. Increasing Temporal Locality with Skewing and Recursive Blocking. InSC ’01: Proceedings of the 2001 ACM/IEEE Conference on Supercomputing. 57–57. doi:10.1109/SC.2001.10041
2001 arXiv
-
[23]
Kun Li, Liang Yuan, Yunquan Zhang, and Yue Yue. 2021. Reducing Redundancy in Data Organization and Arithmetic Calculation for Stencil Computations. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (St. Louis, Misso...
2021
-
[24]
Xiaoyan Liu, Yi Liu, Hailong Yang, Jianjin Liao, Mingzhen Li, Zhongzhi Luan, and Depei Qian. 2022. Toward accelerated stencil computation by adapting tensor core unit on GPU. In Proceedings of the 36th ACM International Conference on Supercomputing. 1–12
2022
-
[25]
Lusher, Satya P
David J. Lusher, Satya P. Jammy, and Neil D. Sandham. 2021. OpenSBLI: Automated code-generation for heterogeneous computing architectures ap- plied to compressible fluid dynamics on structured grids. Computer Physics Communications 267 (2021), 108063. doi:10.1016/j.cpc.2021.108063
2021
-
[26]
Naoya Maruyama and Takayuki Aoki. 2014. Optimizing stencil computations for NVIDIA Kepler GPUs. In Proceedings of the 1st international workshop on high-performance stencil computations, Vienna. Citeseer, 89–95
2014
-
[27]
Naoya Maruyama, Kento Sato, Tatsuo Nomura, and Satoshi Matsuoka. 2011. Physis: An implicitly parallel programming model for stencil computations on large-scale GPU-accelerated supercomputers. In SC ’11: Proceedings of 2011 International Conference for High Performance Computin...
2011
-
[28]
Kazuaki Matsumura, Hamid Reza Zohouri, Mohamed Wahib, Toshio Endo, and Satoshi Matsuoka. 2020. AN5D: automated stencil framework for high-degree temporal blocking on GPUs. InProceedings of the 18th ACM/IEEE International Symposium on Code Generation and Optimization. 199–211
2020
-
[29]
Jiayuan Meng and Kevin Skadron. 2009. Performance Modeling and Automatic Ghost Zone Optimization for Iterative Stencil Loops on GPUs. In Proceedings of the 23rd International Conference on Supercomputing (Yorktown Heights, NY, USA) (ICS ’09). Association for Computing Machiner...
2009
-
[30]
Nvidia. 2023. cuDNN. https://developer.nvidia.com/cudnn, Last accessed on 2023-7-24
2023
-
[31]
NVIDIA Corporation. 2020. NVIDIA A100 Tensor Core GPU Datasheet. https: //www.nvidia.com/en-us/data-center/a100/ Accessed: 2024-11-21
2020
-
[32]
NVIDIA Corporation. 2024. Parallel Thread Execution ISA Version 8.5. https: //docs.nvidia.com/cuda/parallel-thread-execution/index.html Accessed: 2024-11- 21
2024
-
[33]
Jeff Pool, Abhishek Sawarkar, and Jay Rodge. 2021. Accelerating Inference with Sparsity Using the NVIDIA Ampere Architecture and NVIDIA Ten- sorRT. https://developer.nvidia.com/blog/accelerating-inference-with-sparsity- using-ampere-and-tensorrt/ Accessed: 2024-11-22
2021
-
[34]
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013. Halide: A Language and Com- piler for Optimizing Parallelism, Locality, and Recomputation in Image Pro- cessing Pipelines. In Proceedings of the 34th ACM SIGPLAN Con...
2013
-
[35]
Ramanujam, Atanas Rountev, and P
Prashant Rawat, Martin Kong, Tom Henretty, Justin Holewinski, Kevin Stock, Louis-Noël Pouchet, J. Ramanujam, Atanas Rountev, and P. Sadayappan. 2015. SDSLc: A Multi-Target Domain-Specific Compiler for Stencil Computations. In Proceedings of the 5th International Workshop on Do...
2015
-
[36]
Sadayappan
Prashant Singh Rawat, Aravind Sukumaran-Rajam, Atanas Rountev, Fabrice Rastello, Louis-Noël Pouchet, and P. Sadayappan. 2018. Associative Instruction Reordering to Alleviate Register Pressure. In SC18: International Conference for High Performance Computing, Networking, Storag...
2018
-
[37]
Sadayap- pan
Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Mahesh Rav- ishankar, Vinod Grover, Atanas Rountev, Louis-Noël Pouchet, and P. Sadayap- pan. 2018. Domain-Specific Optimization and Generation of High-Performance GPU Code for Stencil Computations. Proc. IEEE 106, 1...
2018
-
[38]
Sadayappan
Prashant Singh Rawat, Miheer Vaidya, Aravind Sukumaran-Rajam, Atanas Roun- tev, Louis-Noël Pouchet, and P. Sadayappan. 2019. On Optimizing Complex Stencils on GPUs. In2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 641–652. doi:10.1109/IPDPS.2019.00073
2019
-
[39]
Rivera and Chau-Wen Tseng
G. Rivera and Chau-Wen Tseng. 2000. Tiling Optimizations for 3D Scientific Computations. In SC ’00: Proceedings of the 2000 ACM/IEEE Conference on Supercomputing. 32–32. doi:10.1109/SC.2000.10015
2000
-
[40]
Ramanujam, and P
Kevin Stock, Martin Kong, Tobias Grosser, Louis-Noël Pouchet, Fabrice Rastello, J. Ramanujam, and P. Sadayappan. 2014. A Framework for Enhancing Data Reuse via Associative Reordering. SIGPLAN Not. 49, 6 (jun 2014), 65–76. doi:10.1145/ 2666356.2594342
2014
-
[41]
Sven Verdoolaege, Juan Carlos Juega, Albert Cohen, José Ignacio Gómez, Chris- tian Tenllado, and Francky Catthoor. 2013. Polyhedral Parallel Code Generation for CUDA. ACM Trans. Archit. Code Optim. 9, 4, Article 54 (jan 2013), 23 pages. doi:10.1145/2400682.2400713
2013
-
[42]
Luhan Wang, Haipeng Jia, Lei Xu, Cunyang Wei, Kun Li, Xianmeng Jiang, and Yunquan Zhang. 2024. VNEC: A Vectorized Non-Empty Column Format for SpMV on CPUs. In 2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). 14–25. doi:10.1109/IPDPS57955.2024.00011
2024
-
[43]
David Wonnacott. 2002. Achieving Scalable Locality with Time Skewing. Int. J. Parallel Program. 30, 3 (jun 2002), 181–221. doi:10.1023/A:1015460304860
2002 doi
-
[44]
Xin You, Hailong Yang, Zhonghui Jiang, Zhongzhi Luan, and Depei Qian. 2021. DRStencil: Exploiting Data Reuse within Low-order Stencil on GPU. In 2021 IEEE 23rd Int Conf on High Performance Computing & Communications; 7th Int Conf on Data Science & Systems; 19th Int Conf on Sma...
2021
-
[45]
Liang Yuan, Shan Huang, Yunquan Zhang, and Hang Cao. 2019. Tessellating Star Stencils. In Proceedings of the 48th International Conference on Parallel Processing (Kyoto, Japan) (ICPP ’19). Association for Computing Machinery, New York, NY, USA, Article 43, 10 pages. doi:10.114...
2019
-
[46]
Liang Yuan, Yunquan Zhang, Peng Guo, and Shan Huang. 2017. Tessellating Stencils. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Colorado) (SC ’17). As- sociation for Computing Machinery, New York, NY, U...
2017
-
[47]
Lingqi Zhang, Mohamed Wahib, Peng Chen, Jintao Meng, Xiao Wang, Toshio Endo, and Satoshi Matsuoka. 2023. PERKS: a Locality-Optimized Execution Model for Iterative Memory-bound GPU Applications. In Proceedings of the 37th International Conference on Supercomputing. 167–179
2023
-
[48]
Lingqi Zhang, Mohamed Wahib, Peng Chen, Jintao Meng, Xiao Wang, Toshio Endo, and Satoshi Matsuoka. 2023. Revisiting Temporal Blocking Sten- cil Optimizations. In Proceedings of the 37th International Conference on Supercomputing. 251–263
2023
-
[49]
Yiwei Zhang, Kun Li, Liang Yuan, Jiawen Cheng, Yunquan Zhang, Ting Cao, and Mao Yang. 2024. LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor Cores . In2024 SC24: International Conference for High Performance Computing, Networking, Storage and Analysis SC. IEEE...
2024 arXiv
-
[50]
Tuowen Zhao, Protonu Basu, Samuel Williams, Mary Hall, and Hans Johansen
-
[51]
Tuowen Zhao, Mary Hall, Hans Johansen, and Samuel Williams. 2021. Improving Communication by Optimizing On-Node Data Movement with Data Layout. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (Virtual Event, Republic of Korea...
2021
-
[52]
Tuowen Zhao, Samuel Williams, Mary Hall, and Hans Johansen. 2018. Deliv- ering Performance-Portable Stencil Computations on CPUs and GPUs Using Bricks. In 2018 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (P3HPC). 59–70. doi:10.1109/P3HPC...
2018
-
[53]
Size Zheng, Renze Chen, Anjiang Wei, Yicheng Jin, Qin Han, Liqiang Lu, Bingyang Wu, Xiuhong Li, Shengen Yan, and Yun Liang. 2022. AMOS: enabling automatic mapping for tensor computations on spatial accelerators with hard- ware abstraction. In Proceedings of the 49th Annual Int...
2022
-
[2019]
In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Colorado) (SC ’19)
Exploiting Reuse and Vectorization in Blocked Stencil Computations on CPUs and GPUs. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Colorado) (SC ’19). Association for Computing Machinery, New York, NY, ...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.