A scale-invariant revision to the fast-mode scaling formula in Ozaki scheme II ensures the CRT uniqueness condition holds for all input scalings while preserving the speed of the original fast mode.
and Mary, Theo , year=
7 Pith papers cite this work, alongside 123 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
verdicts
UNVERDICTED 7roles
background 2polarities
background 2representative citing papers
A fused gather-GEMM-scatter CUDA kernel achieves 4.6-7.3x end-to-end speedup and 3.2-4.9x lower energy for matrix-free 3D SIMP topology optimization on RTX 4090 compared to three-stage baselines.
The paper introduces matrix-multiplication-based iterative refinement for diagonalizable non-Hermitian eigendecompositions that achieves quadratic residual reduction for simple eigenvalues and includes cluster stabilization.
Develops mixed-precision iterative refinement for low-rank Lyapunov equations with rounding error analysis enabling reduced precision for moderately conditioned problems.
Analog-aware block Jacobi schemes in flexible GMRES maintain convergence under simulated device non-idealities when block size, damping, and approximation accuracy are chosen to account for analog scaling, noise, quantization, and clipping.
Error analysis and cost estimator for recasting floating-point matrix multiplication as accumulated integer products on mixed-precision hardware.
PROMISE tool automates mixed-precision tuning with user-defined floating-point formats, validated on linear solvers and Rodinia benchmarks showing many variables can use lower precision safely.
citing papers explorer
-
Improved Scaling for Fast Mode of Ozaki Scheme II
A scale-invariant revision to the fast-mode scaling formula in Ozaki scheme II ensures the CRT uniqueness condition holds for all input scalings while preserving the speed of the original fast mode.
-
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
A fused gather-GEMM-scatter CUDA kernel achieves 4.6-7.3x end-to-end speedup and 3.2-4.9x lower energy for matrix-free 3D SIMP topology optimization on RTX 4090 compared to three-stage baselines.
-
Iterative Refinement for Diagonalizable Non-Hermitian Eigendecompositions
The paper introduces matrix-multiplication-based iterative refinement for diagonalizable non-Hermitian eigendecompositions that achieves quadratic residual reduction for simple eigenvalues and includes cluster stabilization.
-
Mixed-precision iterative refinement for low-rank Lyapunov equations
Develops mixed-precision iterative refinement for low-rank Lyapunov equations with rounding error analysis enabling reduced precision for moderately conditioned problems.
-
Hybrid Digital-Analog Approximate Inverse Preconditioning for Krylov Methods
Analog-aware block Jacobi schemes in flexible GMRES maintain convergence under simulated device non-idealities when block size, damping, and approximation accuracy are chosen to account for analog scaling, noise, quantization, and clipping.
-
Analysis of Floating-Point Matrix Multiplication Computed via Integer Arithmetic
Error analysis and cost estimator for recasting floating-point matrix multiplication as accumulated integer products on mixed-precision hardware.
-
Floating-point autotuning with customized precisions
PROMISE tool automates mixed-precision tuning with user-defined floating-point formats, validated on linear solvers and Rodinia benchmarks showing many variables can use lower precision safely.