REVIEW 3 major objections 5 minor 1 cited by
Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read One portable GPU routine for singular values, based on two-stage QR reduction, beats most vendor libraries beyond 1024×1024 and reaches 80–90% of cuSOLVER on large matrices.
desk verdict A genuine engineering contribution—portable GPU SVD in Julia with first Apple Metal and FP16 support—but the headline performance claim is scoped too broadly and the MAGMA/SLATE comparison needs a fairness pass. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine is the tile-wise two-stage QR reduction: for each diagonal tile, an RQ sweep applies a block Householder QR to the diagonal tile and annihilates below-diagonal tiles, then an LQ sweep (obtained by QR on a lazy transpose, avoiding data movement) annihilates tiles to the right. The kernels are specialized at compile time through Julia's multiple dispatch, with three tuned parameters: TILESIZE (algorithmic tile width), SPLITK (threads per column in panel factorization), and COLPERBLOCK (columns per thread block in the update). Fusing the TSQRT and TSMQR row loops turns a quadratic number of kernel launches into a linear number. This machinery is what converts the generic algorithm in
What would settle it
Re-benchmark the same matrix sizes with all libraries computing the same requested outputs, including singular vectors, and with MAGMA/SLATE in their native multi-GPU configurations. If the unified implementation no longer beats MAGMA, SLATE, rocSOLVER, and oneMKL above 1024×1024, or if cuSOLVER's lead exceeds 20–50% on large matrices, the central portability claim is undercut.
Extended reading notes
Core claim
The paper claims that a single, type- and hardware-generic GPU implementation of the singular value decomposition can be made fast enough to compete with vendor-tuned libraries, and backs that claim with a specific construction. The implementation uses the classical two-stage QR scheme: a compute-bound reduction of the dense matrix to band form (phase one), a GPU tile-based reduction from band to bidiagonal form (phase two), and a final bidiagonal-to-diagonal stage delegated to LAPACK's divide-and-conquer on the CPU. The novelty is in the kernels: a register-resident tile Householder QR with tunable split-K parallelism, a column-blocked trailing-submatrix update, and a fused kernel that proc
Load-bearing premise
The performance results stand on the benchmark configuration: vendor libraries were run in comparable single-GPU modes without singular vectors, so if those settings understate what the vendor solvers can do (with vectors requested, multi-GPU, or 64-bit addressing), the reported speedups could shrink or reverse.
Editorial extensions
If this is right
- For matrices above 1024×1024, the unified routine beats four of the five reference libraries; only cuSOLVER remains ahead, and by at most 10–20% on large matrices.
- AMD, Intel, and Apple GPU users gain a GPU SVD without vendor-specific rewrites, and Apple Metal gains SVD support for the first time.
- Half-precision users can keep larger matrices GPU-resident (up to $131{,}000 \times 131{,}000$ on H100) because FP16 storage halves memory, even though current compute paths upcast to FP32.
- Hyperparameter tuning (TILESIZE, SPLITK, COLPERBLOCK) recovers most of the lost performance, so new hardware can be supported by re-tuning rather than rewriting kernels.
- Accuracy tracks cuSOLVER: relative Frobenius errors around $10^{-16}$ in FP64, $10^{-8}$ in FP32, and $10^{-3}$ in FP16.
Reading between the lines
- Because the comparison computes singular values only, without U and V, a full-SVD variant would stress the memory-bound band-to-bidiagonal stage more heavily and could shift the ranking relative to cuSOLVER.
- The same 'one kernel, many backends, tuned hyperparameters' recipe likely transfers to other dense factorizations in the same abstraction layer, making the portability argument a template rather than a one-off result.
- The FP16 path currently upcasts to FP32; on GPUs with native scalar FP16 or a Tensor Core path, the same code could become substantially faster than FP32, which would matter for LoRA-style LLM adapters.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a Julia implementation of a two-stage QR-based SVD solver for GPUs, built on GPUArrays.jl and KernelAbstractions.jl. The solver is hardware- and precision-generic, supporting NVIDIA, AMD, Intel, and Apple GPUs in FP16/FP32/FP64, and is claimed to be the first GPU-accelerated SVD for Apple Metal and the first GPU SVD supporting half precision. The authors describe custom CUDA/HIP/Metal kernels for the panel factorization and trailing update, a fused TSQRT/TSMQR kernel, and a hyperparameter-tuning strategy (TILESIZE, COLPERBLOCK, SPLITK). Accuracy is evaluated against synthetic matrices with known singular values. Performance is benchmarked against MAGMA, SLATE, rocSOLVER, oneMKL, and cuSOLVER on several GPUs; the abstract claims the unified function outperforms most libraries for matrices larger than 1024x1024 and reaches 80-90% of cuSOLVER on large matrices.
Significance. If the performance claims hold, the paper demonstrates a meaningful milestone for portable GPU linear algebra: a single open-source implementation can approach vendor-tuned libraries across diverse architectures and precisions. The strong points are the breadth of hardware and type support, the concrete kernel-level description, the empirical accuracy validation across three singular-value distributions, the open-source release, and the honest disclosure of tuned hyperparameters. The central contribution is an engineering/implementation result rather than a new numerical algorithm; its significance depends on the fairness and reproducibility of the performance comparison against MAGMA, SLATE, and cuSOLVER.
major comments (3)
- [Section 3.4 / Section 4.1] The headline claim in the abstract 'outperforms most linear algebra libraries (MAGMA, SLATE, rocSOLVER, oneMKL) for matrix sizes larger than 1024x1024' is not established for MAGMA and SLATE as configured. Section 3.4 states MAGMA is run 'with 1 GPU and no singular vectors specified' and SLATE 'with as options the target and origin being the device' and no vectors. Both are hybrid CPU-GPU/distributed-memory libraries; the paper itself concedes in Section 4.1 that they 'are designed for large-scale problems that leverage multi-GPU and hybrid CPU-GPU systems, which may explain SLATE's lower performance on consumer-grade laptops.' The report omits CPU-thread counts for MAGMA and MPI/task configuration for SLATE, so the comparison may not reflect their intended operating regime. The authors should either benchmark these libraries in representative single-GPU hybrid configurations with the sa
- [Abstract / Section 4.1 / Section 5] The crossover size for the MAGMA comparison is inconsistent. The abstract and conclusion state the unified function outperforms MAGMA/SLATE for matrices 'larger than 1024', but Section 4.1 states 'the unified implementation exceeding equal runtime for all hardware on matrix sizes larger than 2048 x 2048', and then later concludes 'for matrix sizes larger than 256, the unified API matches or surpasses the performance of MAGMA.' These thresholds are not equivalent and affect the central claim. Please reconcile the abstract, Section 4.1, and the conclusion with the actual plotted data.
- [Section 3.4 / Figures 3-4] The performance measurements are presented without any measure of variance. The text says 'benchmarked over 20 runs with a single synchronization at the end... repeated until 2 seconds total benchmark time' and 'Each measurement is run twice after each other,' but the ratio plots in Figures 3 and 4 show single values with no error bars, and Table 4 gives only geometric means and ranges. Since several comparisons are close to 1.0 (e.g., A100/H100 vs. cuSOLVER at 0.8-0.9), the reader cannot tell whether observed differences are beyond run-to-run noise. Please report medians or means with standard error/min-max over independent repetitions, and clarify the exact timing protocol.
minor comments (5)
- [Table 4] The A100/SLATE row reports a geometric mean of 2.5 with range (3.2 - 5.7); this is internally inconsistent since the mean must lie within the range if the range is the min-max. Please correct the values or the range.
- [Section 4.3 / Table 2] Figure 5 and the text reference 'Apple Metal M3' and 'Apple M1' inconsistently; Table 2 lists 'Apple M1 Pro'. Please unify the hardware designation.
- [Table 3] The table formatting makes the tuned values hard to parse: the row labels 'TILESIZE 64 to 128' and 'COLPERBLOCK 32 to 16' are ambiguous, and the caption says 'Hyperparamter' (typo). Please restructure the table so the matrix sizes, varied parameter, and direction of change are explicit.
- [Section 4.1] The sentence 'we see the unified implementation exceeding equal runtime for all hardware' should be reworded to 'exceeding equal runtime' or 'attaining speedups greater than 1' for clarity.
- [Section 3.4] Please give the exact versions and build options for MAGMA, SLATE, rocSOLVER, oneMKL, and cuSOLVER, and specify whether the vendor routines were called with the same data layout/pointer mode as the unified solver.
Circularity Check
No circularity found: the paper reports measured performance against external libraries; tuned hyperparameters are disclosed, and no fitted quantity is back-labeled as a prediction.
full rationale
This is an implementation-and-benchmark paper rather than a mathematical derivation, so the circularity patterns that apply to derived predictions do not arise. The central claims are empirical: the unified implementation is benchmarked against MAGMA, SLATE, rocSOLVER, oneMKL, and cuSOLVER, and the reported runtimes are direct measurements, not quantities derived from fitted parameters. The tuning knobs TILESIZE, COLPERBLOCK, and SPLITK are explicitly disclosed in Sections 3.2–3.3 as hyperparameters searched per hardware and precision; the paper does not fit a parameter to a subset of benchmark outcomes and then claim those outcomes as predicted. Accuracy is checked against externally constructed matrices with known singular values and compared to cuSOLVER, not against the implementation itself. The algorithm is attributed to external work (Haidar et al., Ballard et al., LAPACK divide-and-conquer), and self-citations to NextLA.jl, KernelAbstractions.jl, and GPUArrays.jl provide ecosystem context rather than load-bearing evidence for the measured speedups. The manuscript also contains an explicit limitation statement in Section 4.1: 'MAGMA and SLATE libraries are designed for large-scale problems that leverage multi-GPU and hybrid CPU–GPU systems, which may explain SLATE’s lower performance on consumer-grade laptops.' That is a benchmarking-fairness caveat that could affect the strength of the comparative claim, but it is not circularity: the comparison is still an external, falsifiable measurement. No equation or result in the paper reduces to its own input by construction.
Assumptions & free parameters
free parameters (4)
- TILESIZE =
per-hardware tuned, values not reported for final benchmarks
- COLPERBLOCK =
per-hardware tuned, values not reported
- SPLITK =
per-hardware tuned, values not reported
- 10*epsilon threshold =
10 x machine epsilon for each precision
assumptions (4)
- standard math Householder QR is backward stable
- standard math LAPACK divide-and-conquer bidiagonal SVD is correct
- domain assumption cuSOLVER and rocSOLVER have the cited 64-bit addressing limitations
- domain assumption GPUArrays.jl and KernelAbstractions.jl generate correct device code
Cite this review
Pith. "Pith review of Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision." pith.science (2026). https://pith.science/paper/ZEESZI5S
@misc{pith2026250806339,
author = {Pith},
title = {Pith review of: Performant Unified GPU Kernels for Portable Singular Value Computation Across Hardware and Precision},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZEESZI5S}},
note = {Machine review of arXiv:2508.06339}
}
read the original abstract
This paper presents a portable, GPU-accelerated implementation of a QR-based singular value computation algorithm in Julia. The singular value ecomposition (SVD) is a fundamental numerical tool in scientific computing and machine learning, providing optimal low-rank matrix approximations. Its importance has increased even more in large-scale machine learning pipelines, including large language models (LLMs), where it enables low-rank adaptation (LoRA). The implemented algorithm is based on the classic two-stage QR reduction, consisting of successive matrix reduction to band form and bidiagonal form. Our implementation leverages Julia's multiple dispatch and metaprogramming capabilities, integrating with the GPUArrays and KernelAbstractions frameworks to provide a unified type and hardware-agnostic function. It supports diverse GPU architectures and data types, and is, to our knowledge, the first GPU-accelerated singular value implementation to support Apple Metal GPUs and half precision. Performance results on multiple GPU backends and data types demonstrate that portability does not require sacrificing performance: the unified function outperforms most linear algebra libraries (MAGMA, SLATE, rocSOLVER, oneMKL) for matrix sizes larger than 1024x1024, and achieves 80%-90% of the performance of cuSOLVER for large matrices.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Accelerating Bidiagonalization of Banded Matrices through Memory-Aware Bulge-Chasing on GPUs
A memory-aware GPU bulge-chasing algorithm reduces banded matrices to bidiagonal form, achieving >100x speedups over CPU libraries at 32k sizes.
Reference graph
Works this paper leans on
-
[1]
Ahmad Abdelfattah, Natalie Beams, Robert Carson, Pieter Ghysels, Tzanio Kolev, Thomas Stitt, Arturo Vargas, Stanimire Tomov, and Jack Dongarra. 2024. MAGMA: Enabling Exascale Performance with Accelerated BLAS and LAPACK for Diverse GPU Architectures. The International Journal of High Performance Computing Applications 38, 5 (2024), 468–490
work page 2024
- [2]
-
[3]
Rabab Alomairy, Mark Gates, Sebastien Cayrols, and Dalal Sukkari. 2022. Com- munication Avoiding{LU} with Tournament Pivoting in{SLATE},{SWAN} No. 18. https://icl.utk.edu/files/publications/2022/icl-utk-1533-2022.pdf
work page 2022
-
[4]
Rabab Alomairy, Hatem Ltaief, Mustafa Abduljabbar, and David Keyes. 2020. Abstraction layer for standardizing APIs of task-based engines. IEEE Transactions on Parallel and Distributed Systems 31, 11 (2020), 2482–2495
work page 2020
-
[5]
2025.NextLA.jl: Next-Gen Linear Algebra
Rabab Alomairy, Evelyne Ringoot, Sophie Xuan, Vicki Carrica, Maxwell Onyango, and Julian Samaroo. 2025.NextLA.jl: Next-Gen Linear Algebra. NextLinearAlgebra. https://doi.org/10.5281/zenodo.15049222
-
[6]
Rabab Alomairy, Felipe Tome, Julian Samaroo, and Alan Edelman. 2024. Dynamic Task Scheduling with Data Dependency Awareness Using Julia. In 2024 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 1–7. doi:10.1109/ HPEC62836.2024.10938467
arXiv 2024
-
[7]
Aksel Alpay, Bálint Soproni, Holger Wünsche, and Vincent Heuveline. 2022. Exploring the possibility of a hipSYCL-based implementation of oneAPI. In Pro- ceedings of the 10th International Workshop on OpenCL (Bristol, United Kingdom, United Kingdom) (IWOCL ’22). Association for Computing Machinery, New York, NY, USA, Article 10, 12 pages. doi:10.1145/35295...
-
[8]
Edward Anderson, Zhaojun Bai, Christian Bischof, L Susan Blackford, James Demmel, Jack Dongarra, Jeremy Du Croz, Anne Greenbaum, Sven Hammarling, Alan McKenney, et al. 1999. LAPACK Users’ Guide. SIAM
work page 1999
Show all 66 references
-
[9]
Apple. 2025. Metal Performance Shaders. Apple. https://developer.apple.com/ documentation/metalperformanceshaders
2025
-
[10]
Alex Arslan, Stefan Karpinski, Kristoffer Carlsson, and all. 2023. Embedding julia. Julia language. Retrieved July 28, 2025 from https://docs.julialang.org/en/v1/ manual/embedding/
2023
-
[11]
Cédric Augonnet, Samuel Thibault, Raymond Namyst, and Pierre-André Wacre- nier. 2009. StarPU: A Unified Platform for Task Scheduling on Heterogeneous Multicore Architectures. In Euro-Par 2009 Parallel Processing (Berlin, Heidelberg), Henk Sips, Dick Epema, and Hai-Xiang Lin (E...
2009
-
[12]
Grey Ballard, James Demmel, and Nicholas Knight. 2012. Communication avoid- ing successive band reduction. In Proceedings of the 17th ACM SIGPLAN Sympo- sium on Principles and Practice of Parallel Programming (New Orleans, Louisiana, USA) (PPoPP ’12). Association for Computing...
2012
-
[13]
Tim Besard. 2025. oneAPI.jl. JuliaGPU. doi:10.5281/zenodo.14615352
2025 doi
-
[14]
Tim Besard, Valentin Churavy, Alan Edelman, and Bjorn De Sutter. 2019. Rapid Software Prototyping for Heterogeneous and Distributed Platforms. Advances in Engineering Software 132 (2019), 29–46
2019
-
[15]
Tim Besard, Christophe Foket, and Bjorn De Sutter. 2019. Effective Extensible Programming: Unleashing Julia on GPUs. IEEE Transactions on Parallel and Distributed Systems 30, 4 (2019), 827–841. doi:10.1109/TPDS.2018.2872064
2019
-
[16]
Tim Besard and Max Hawkins. 2025. Metal.jl. JuliaGPU. doi:10.5281/zenodo. 14615291
2025 doi
-
[17]
Shah, Jan Vitek, and Lionel Zoubritzky
Jeff Bezanson, Jiahao Chen, Benjamin Chung, Stefan Karpinski, Viral B. Shah, Jan Vitek, and Lionel Zoubritzky. 2018. Julia: dynamism and performance reconciled by design. Proc. ACM Program. Lang. 2, OOPSLA, Article 120 (Oct. 2018), 23 pages. doi:10.1145/3276490
2018 doi
-
[18]
L Susan Blackford, Jaeyoung Choi, Andy Cleary, Eduardo D’Azevedo, James Demmel, Inderjit Dhillon, Jack Dongarra, Sven Hammarling, Greg Henry, Antoine Petitet, et al. 1997. ScaLAPACK Users’ Guide. SIAM
1997
-
[19]
Cory Bloor, Juan Zuniga-Anaya, Troy Alderso, and all. 2025. rocSOLVER. AMD. https://github.com/ROCm/rocSOLVER
2025
-
[20]
George Bosilca, Aurelien Bouteiller, Anthony Danalis, Mathieu Faverge, Azzam Haidar, Thomas Herault, Jakub Kurzak, Julien Langou, Pierre Lemarinier, Hatem Ltaief, et al. 2011. Flexible Development of Dense Linear Algebra Algorithms on Massively Parallel Architectures with DPLA...
2011
-
[21]
George Bosilca, Aurelien Bouteiller, Anthony Danalis, Mathieu Faverge, Azzam Haidar, Thomas Herault, Jakub Kurzak, Julien Langou, Pierre Lemarinier, Hatem Ltaief, Piotr Luszczek, Asim YarKhan, and Jack Dongarra. 2011. Flexible Devel- opment of Dense Linear Algebra Algorithms o...
2011 doi
-
[22]
William Brandon. 2024. Matmul3. https://github.com/accelerated-computing- class/lab6/blob/main/matmul_3.cu
2024
-
[23]
Alfredo Buttari, Julien Langou, Jakub Kurzak, and Jack Dongarra. 2009. A class of parallel tiled linear algebra algorithms for multicore architectures. Parallel Comput. 35, 1 (2009), 38–53. doi:10.1016/j.parco.2008.10.002 2025-08-11 01:00. Page 10 of 1–12. Performant Unified G...
2009 doi
-
[24]
Qinglei Cao, Rabab Alomairy, Yu Pei, George Bosilca, Hatem Ltaief, David Keyes, and Jack Dongarra. 2022. A Framework to Exploit Data Sparsity in Tile Low- Rank Cholesky Factorization. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 414–424
2022
-
[25]
Vicki Carrica, Maxwell Onyango, Rabab Alomairy, Evelyne Ringoot, James Schloss, and Alan Edelman. 2025. Toward Portable GPU Performance: Julia Recursive Implementation of TRMM and TRSM. arXiv preprint arXiv:2504.13821 (2025), 10
2025 arXiv
-
[26]
Jiahao Chen. 2024. Randommatrices.jl. https://github.com/JuliaMath/ RandomMatrices.jl
2024
-
[27]
Zihan Chen. 2018. Singular Value Decomposition and Its Applications in Image Processing. In Proceedings of the 2018 1st International Conference on Mathematics and Statistics (Porto, Portugal) (ICoMS ’18). Association for Computing Machinery, New York, NY, USA, 16–22. doi:10.1...
2018
-
[28]
Valentin Churavy. 2019. Transparent distributed programming in Julia . Master’s thesis. Massachusetts Institute of Technology
2019
-
[29]
Valentin Churavy. 2023. KernelAbstractions.jl. https://github.com/JuliaGPU/ KernelAbstractions.jl
2023
-
[30]
Connolly and Nicholas J
Michael P. Connolly and Nicholas J. Higham. 2023. Probabilistic Rounding Error Analysis of Householder QR Factorization. SIAM J. Matrix Anal. Appl. 44, 3 (2023), 1146–1163. doi:10.1137/22M1514817 arXiv:https://doi.org/10.1137/22M1514817
2023 doi
-
[31]
Joshua H Davis, Pranav Sivaraman, Isaac Minn, Konstantinos Parasyris, Harshitha Menon, Giorgis Georgakoudis, and Abhinav Bhatele. 2024. An Evaluative Com- parison of Performance Portability across GPU Programming Models. arXiv preprint arXiv:2402.08950 (2024), 12
2024 arXiv
-
[32]
Tom Deakin, Simon McIntosh-Smith, James Price, Andrei Poenaru, Patrick Atkin- son, Codrin Popa, and Justin Salmon. 2019. Performance Portability across Diverse Computer Architectures. In 2019 IEEE/ACM International Workshop on Performance, Portability and Productivity in HPC (...
2019
-
[33]
J. Demmel. 1989. LAPACK: a portable linear algebra library for supercomputers. In IEEE Control Systems Society Workshop on Computer-Aided Control System Design. 1–7. doi:10.1109/CACSD.1989.69824
1989
-
[34]
Jack Dongarra, Mark Gates, Azzam Haidar, Jakub Kurzak, Piotr Luszczek, Stan- imire Tomov, and Ichitaro Yamazaki. 2014. Accelerating Numerical Dense Linear Algebra Calculations with GPUs . Springer International Publishing, Cham, 3–28. doi:10.1007/978-3-319-06548-9_1
2014 doi
-
[35]
Jack Dongarra, Mark Gates, Azzam Haidar, Jakub Kurzak, Piotr Luszczek, Stan- imire Tomov, and Ichitaro Yamazaki. 2018. The Singular Value Decomposition: Anatomy of Optimizing an Algorithm for Extreme Scale. SIAM Rev. 60, 4 (2018), 808–865. doi:10.1137/17M1117732 arXiv:https://...
2018 doi
-
[36]
Jack Dongarra, Mark Gates, Azzam Haidar, Jakub Kurzak, Piotr Luszczek, Panruo Wu, Ichitaro Yamazaki, Asim Yarkhan, Maksims Abalenkovs, Negin Bagherpour, Sven Hammarling, Jakub Šístek, David Stevens, Mawussi Zounon, and Samuel D. Relton. 2019. PLASMA: Parallel Linear Algebra So...
2019
-
[37]
Dufek, Rahulkumar Gayatri, Neil Mehta, Douglas Doerfler, Brandon Cook, Yasaman Ghadar, and Carleton DeTar
Amanda S. Dufek, Rahulkumar Gayatri, Neil Mehta, Douglas Doerfler, Brandon Cook, Yasaman Ghadar, and Carleton DeTar. 2021. Case Study of Using Kokkos and SYCL as Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs. In 2021 International Work...
2021
-
[38]
Mathieu Faverge, Nathalie Furmento, Abdou Guermouche, Gwenolé Lucas, Ray- mond Namyst, Samuel Thibault, and Pierre-andré Wacrenier. 2023. Programming Heterogeneous Architectures Uing Hierarchical Tasks. Concurrency and Compu- tation: Practice and Experience 35, 25 (2023), e7811
2023
-
[39]
Mark Gates, Ahmad Abdelfattah, Kadir Akbudak, Mohammed Al Farhan, Rabab Alomairy, Daniel Bielich, Treece Burgess, Sébastien Cayrols, Neil Lindquist, Dalal Sukkari, et al. 2025. Evolution of the SLATE Linear Algebra Library. The International Journal of High Performance Computi...
2025
-
[40]
Mark Gates, Ali Charara, Jakub Kurzak, Asim YarKhan, Mohammed Al Farhan, Dalal Sukkari, and Jack Dongarra. 2020. SLATE Users’ Guide, SW AN No. 10 . Technical Report ICL-UT-19-01. Innovative Computing Laboratory, University of Tennessee. revision 07-2020
2020
-
[41]
Mark Gates, Stanimire Tomov, and Jack Dongarra. 2018. Accelerating the SVD two stage bidiagonal reduction and divide and conquer using GPUs. Parallel Comput. 74 (2018), 3–18. doi:10.1016/j.parco.2017.10.004 Parallel Matrix Algorithms and Applications (PMAA’16)
2018 doi
-
[42]
Godoy, Pedro Valero-Lara, T
William F. Godoy, Pedro Valero-Lara, T. Elise Dettling, Christian Trefftz, Ian Jorquera, Thomas Sheehy, Ross G. Miller, Marc Gonzalez-Tallada, Jeffrey S. Vetter, and Valentin Churavy. 2023. Evaluating performance and portability of high-level programming models: Julia, Python/...
2023
-
[43]
Azzam Haidar, Hatem Ltaief, and Jack Dongarra. 2011. Parallel reduction to con- densed forms for symmetric eigenvalue problems using aggregated fine-grained and memory-aware kernels. In Proceedings of 2011 International Conference for High Performance Computing, Networking, St...
2011
-
[44]
Azzam Haidar, Hatem Ltaief, Piotr Luszczek, and Jack Dongarra. 2012. A Com- prehensive Study of Task Coalescing for Selecting Parallelism Granularity in a Two-Stage Bidiagonal Reduction. In 2012 IEEE 26th International Parallel and Distributed Processing Symposium. 25–35. doi:...
2012 doi
-
[45]
Behnam Hashemi and Yuji Nakatsukasa. 2022. Least-Squares Spec- tral Methods for ODE Eigenvalue Problems. SIAM Journal on Sci- entific Computing 44, 5 (2022), A3244–A3264. doi:10.1137/21M1445934 arXiv:https://doi.org/10.1137/21M1445934
2022 doi
-
[46]
Igual, Ernie Chan, Enrique S
Francisco D. Igual, Ernie Chan, Enrique S. Quintana-Ortí, Gregorio Quintana- Ortí, Robert A. van de Geijn, and Field G. Van Zee. 2012. The FLAME Approach: From Dense Linear Algebra Algorithms to High-Performance Multi-Accelerator Implementations. J. Parallel and Distrib. Compu...
2012
-
[47]
Ronan Keryell, Ruyman Reyes, and Lee Howes. 2015. Khronos SYCL for OpenCL: a tutorial. In Proceedings of the 3rd International Workshop on OpenCL (Palo Alto, California) (IWOCL ’15). Association for Computing Machinery, New York, NY, USA, Article 24, 1 pages. doi:10.1145/27913...
2015
-
[48]
Christoph Klein. 2025. NVIDIA CUDALibrarySamples Issues: eigenvalue solver xsyevd has limit on matrix siz. Retrieved April 30, 2025 from https://github.com/ NVIDIA/CUDALibrarySamples/issues/246
2025
-
[49]
Zhiteng Li, Mingyuan Xia, Jingyuan Zhang, Zheng Hui, Linghe Kong, Yu- lun Zhang, and Xiaokang Yang. 2025. AdaSVD: Adaptive Singular Value De- composition for Large Language Models. arXiv:2502.01403 [cs.CV] https: //arxiv.org/abs/2502.01403
2025
-
[50]
Milo Lurati, Stijn Heldens, Alessio Sclocco, and Ben van Werkhoven. 2024. Bring- ing Auto-Tuning to HIP: Analysis of Tuning Impact and Difficulty on AMD and Nvidia GPUs. In Euro-Par 2024: Parallel Processing , Jesus Carretero, Sameer Shende, Javier Garcia-Blas, Ivona Brandic, ...
2024
-
[51]
Giulio Malenza, Valentina Cesare, Marco Edoardo Santimaria, Robert Birke, Al- berto Vecchiato, Ugo Becciani, and Marco Aldinucci. 2024. Performance portabil- ity via C++ PSTL, SYCL, OpenMP, and HIP: the Gaia AVU-GSR case study. InSC24- W: Workshops of the International Confere...
2024
-
[52]
Matthew Martineau, Simon McIntosh-Smith, and Wayne Gaudin
-
[53]
JAROSŁAW ADAM MISZCZAK. 2011. SINGULAR VALUE DECOMPOSITION AND MATRIX REORDERINGS IN QUANTUM INFORMATION THEORY. In- ternational Journal of Modern Physics C 22, 09 (2011), 897–918. doi:10.1142/ S0129183111016683
2011
-
[54]
NVIDIA. 2025. cuSOLVER. NVIDIA. https://developer.nvidia.com/cusolver
2025
-
[55]
Julian Samaroo, Anton Smirnov, Valentin Churavy, Ludovic Räss, Torrance Hodg- son, Alexis Montoison, Wiktor Phillips, Ali Ramadhan, Jason Barmparesos, Tim Besard, Julia TagBot, Michel Schanen, Carsten Bauer, Mosè Giordano, Takafumi Arakaki, Stephan Antholzer, Alessandro, Chris...
2023 doi
-
[56]
Tim Besard Simeon Danisch. 2023. GPUArrays.jl. https://github.com/JuliaGPU/ GPUArrays.jl
2023
-
[57]
John E Stone, David Gohara, and Guochun Shi. 2010. OpenCL: A parallel pro- gramming standard for heterogeneous computing systems. Computing in science & engineering 12, 3 (2010), 66
2010
-
[58]
Qi Sun, Edoardo Cetin, and Yujin Tang. 2025. Transformer-Squared: Self-adaptive LLMs. 19. arXiv:2501.06252 [cs.LG] https://arxiv.org/abs/2501.06252
2025 arXiv
-
[59]
Trott, Damien Lebrun-Grandié, Daniel Arndt, Jan Ciesko, Vinh Dang, Nathan Ellingwood, Rahulkumar Gayatri, Evan Harvey, Daisy S
Christian R. Trott, Damien Lebrun-Grandié, Daniel Arndt, Jan Ciesko, Vinh Dang, Nathan Ellingwood, Rahulkumar Gayatri, Evan Harvey, Daisy S. Hollman, Dan Ibanez, Nevin Liber, Jonathan Madsen, Jeff Miles, David Poliakoff, Amy Powell, Sivasankaran Rajamanickam, Mikael Simberg, D...
2022
-
[60]
Gerlach, Alan Edelman, George Barbastathis, Richard D
Utkarsh Utkarsh, Valentin Churavy, Yingbo Ma, Tim Besard, Prakitr Srisuma, Tim Gymnich, Adam R. Gerlach, Alan Edelman, George Barbastathis, Richard D. Braatz, and Christopher Rackauckas. 2024. Automated translation and accelerated solving of differential equations on multiple ...
2024 doi
-
[61]
Van Zee and Robert A
Field G. Van Zee and Robert A. van de Geijn. 2015. BLIS: A Framework for Rapidly Instantiating BLAS Functionality. ACM Trans. Math. Softw. 41, 3 (06 2025-08-11 01:00. Page 11 of 1–12. ICPP ’25, September 08–11, 2025, San Diego, CA, USA Ringoot, Alomairy, Churavy and Edelman 20...
2015 doi
-
[62]
Qinsi Wang, Jinghan Ke, Masayoshi Tomizuka, Yiran Chen, Kurt Keutzer, and Chenfeng Xu. 2025. Dobi-SVD: Differentiable SVD for LLM Compression and Some New Perspectives. arXiv:2502.02723 [cs.LG] https://arxiv.org/abs/2502. 02723
2025 arXiv
-
[63]
Junmin Xiao, Yunfei Pang, Qing Xue, Chaoyang Shui, Ke Meng, Hui Ma, Mingyi Li, Xiaoyang Zhang, and Guangming Tan. 2022. W-Cycle SVD: A Multilevel Algorithm for Batched SVD on GPUs. In SC22: International Conference for High Performance Computing, Networking, Storage and Analys...
2022 arXiv
-
[64]
Sophie Xuan, Evelyne Ringoot, Rabab Alomairy, Felipe Tome, Julian Samaroo, and Alan Edelman. 2024. Synthesizing Numerical Linear Algebra using Julia. In 2024 IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 2
2024
-
[65]
Jisheng Zhao, Colleen Bertoni, Jeffrey Young, Kevin Harms, Vivek Sarkar, and Brice Videau. 2023. HIPLZ: Enabling performance portability for exascale systems. Concurrency and Computation: Practice and Experience 35, 25 (2023), e7866. doi:10. 1002/cpe.7866 arXiv:https://onlinel...
2023 doi
-
[2017]
Concurrency and Computation: Practice and Experience 29, 15 (2017), e4117
Assessing the performance portability of modern parallel pro- gramming models using TeaLeaf. Concurrency and Computation: Practice and Experience 29, 15 (2017), e4117. doi:10.1002/cpe.4117 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1002/cpe.4117 e4117 cpe.4117
2017 doi
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.