Mixed-precision CA-SGD for GLMs on A100 GPUs matches FP32 loss within 0.5% while delivering 5.1-6.8x speedup via a nine-choice finite-precision error recipe.
Title resolution pending
9 Pith papers cite this work, alongside 185 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
A new spectral truncation method for filtering 3D EFIE integral operators is introduced using the spherical Hankel transform representation of the Green's function, supported by semi-analytical and numerical evidence on operator spectra.
A fused gather-GEMM-scatter CUDA kernel achieves 4.6-7.3x end-to-end speedup and 3.2-4.9x lower energy for matrix-free 3D SIMP topology optimization on RTX 4090 compared to three-stage baselines.
The paper introduces matrix-multiplication-based iterative refinement for diagonalizable non-Hermitian eigendecompositions that achieves quadratic residual reduction for simple eigenvalues and includes cluster stabilization.
Develops mixed-precision iterative refinement for low-rank Lyapunov equations with rounding error analysis enabling reduced precision for moderately conditioned problems.
FP16 SparseStack on GPUs shows embedding quality insensitive to rounding method, with sketch distribution as the primary accuracy driver across tested inputs.
Error analysis and cost estimator for recasting floating-point matrix multiplication as accumulated integer products on mixed-precision hardware.
A preconditioning technique for the shifted Helmholtz operator stabilizes EFIE iterative solvers across multiple frequency and discretization regimes, enabling quasi-linear complexity.
citing papers explorer
-
Mixed-Precision Communication-Avoiding SGD for Generalized Linear Models on GPUs
Mixed-precision CA-SGD for GLMs on A100 GPUs matches FP32 loss within 0.5% while delivering 5.1-6.8x speedup via a nine-choice finite-precision error recipe.
-
Spectral Filtering of 3D Integral Operators Using Modified Green's Functions
A new spectral truncation method for filtering 3D EFIE integral operators is introduced using the spherical Hankel transform representation of the Green's function, supported by semi-analytical and numerical evidence on operator spectra.
-
Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels
A fused gather-GEMM-scatter CUDA kernel achieves 4.6-7.3x end-to-end speedup and 3.2-4.9x lower energy for matrix-free 3D SIMP topology optimization on RTX 4090 compared to three-stage baselines.
-
Iterative Refinement for Diagonalizable Non-Hermitian Eigendecompositions
The paper introduces matrix-multiplication-based iterative refinement for diagonalizable non-Hermitian eigendecompositions that achieves quadratic residual reduction for simple eigenvalues and includes cluster stabilization.
-
Mixed-precision iterative refinement for low-rank Lyapunov equations
Develops mixed-precision iterative refinement for low-rank Lyapunov equations with rounding error analysis enabling reduced precision for moderately conditioned problems.
-
Randomized Sketching is Robust to Low-Precision Rounding on GPUs
FP16 SparseStack on GPUs shows embedding quality insensitive to rounding method, with sketch distribution as the primary accuracy driver across tested inputs.
-
Analysis of Floating-Point Matrix Multiplication Computed via Integer Arithmetic
Error analysis and cost estimator for recasting floating-point matrix multiplication as accumulated integer products on mixed-precision hardware.
-
High-Frequency Preconditioners for Electromagnetic Integral Equations Based on Helmholtz Regularizations
A preconditioning technique for the shifted Helmholtz operator stabilizes EFIE iterative solvers across multiple frequency and discretization regimes, enabling quasi-linear complexity.
- Computing k-means in mixed precision