Mixed-precision CA-SGD for GLMs on A100 GPUs matches FP32 loss within 0.5% while delivering 5.1-6.8x speedup via a nine-choice finite-precision error recipe.
and Stich, S
9 Pith papers cite this work, alongside 16 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Derives convergence bounds for DQL with partial participation and non-convex losses; shows adaptive QNN-powered post-quantum security reduces execution time by 49% on hardware while detecting threats at over 91% accuracy.
UB-SMoE balances expert utilization in heterogeneous federated SMoE fine-tuning via Dynamic Modulated Routing and Universal Pseudo-Gradient, delivering up to 45% compute reduction and 8.7x performance gains for low-resource clients over prior LoRA-rank methods.
Recasts sampling-based nonconvex optimization as smoothed gradient descent to obtain non-asymptotic convergence guarantees and introduces the DIDA annealed algorithm that converges to the global optimum.
Rescaled ASGD recovers convergence to the true global objective by rescaling worker stepsizes proportional to computation times, matching the known time lower bound in the leading term under non-convex smoothness and bounded heterogeneity.
Pre-training provides a geometric warm start in a single-index model that enables weak-to-strong generalization up to a supervisor-limited bound, with empirical phase-transition evidence in LLMs.
StoSignSGD resolves SignSGD divergence on non-smooth objectives via structural stochasticity, matching optimal convex rates and improving non-convex bounds while delivering 1.44-2.14x speedups in FP8 LLM pretraining.
Verification of machine unlearning is fragile because model providers can use adversarial unlearning to pass checks while keeping data influence.
A survey equating offline Monte-Carlo/SAA and online stochastic-approximation sample complexities for convex stochastic optimization arising in statistics and ML.
citing papers explorer
-
Mixed-Precision Communication-Avoiding SGD for Generalized Linear Models on GPUs
Mixed-precision CA-SGD for GLMs on A100 GPUs matches FP32 loss within 0.5% while delivering 5.1-6.8x speedup via a nine-choice finite-precision error recipe.
-
Distributed Quantum Learning over Near-term Devices: Convergence Analysis and Security Design
Derives convergence bounds for DQL with partial participation and non-convex losses; shows adaptive QNN-powered post-quantum security reduces execution time by 49% on hardware while detecting threats at over 91% accuracy.
-
UB-SMoE: Universally Balanced Sparse Mixture-of-Experts for Resource-adaptive Federated Fine-tuning of Foundation Models
UB-SMoE balances expert utilization in heterogeneous federated SMoE fine-tuning via Dynamic Modulated Routing and Universal Pseudo-Gradient, delivering up to 45% compute reduction and 8.7x performance gains for low-resource clients over prior LoRA-rank methods.
-
Global Convergence of Sampling-Based Nonconvex Optimization through Diffusion-Style Smoothing
Recasts sampling-based nonconvex optimization as smoothed gradient descent to obtain non-asymptotic convergence guarantees and introduces the DIDA annealed algorithm that converges to the global optimum.
-
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
Rescaled ASGD recovers convergence to the true global objective by rescaling worker stepsizes proportional to computation times, matching the known time lower bound in the leading term under non-convex smoothness and bounded heterogeneity.
-
On the Blessing of Pre-training in Weak-to-Strong Generalization
Pre-training provides a geometric warm start in a single-index model that enables weak-to-strong generalization up to a supervisor-limited bound, with empirical phase-transition evidence in LLMs.
-
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
StoSignSGD resolves SignSGD divergence on non-smooth objectives via structural stochasticity, matching optimal convex rates and improving non-convex bounds while delivering 1.44-2.14x speedups in FP8 LLM pretraining.
-
Verification of Machine Unlearning is Fragile
Verification of machine unlearning is fragile because model providers can use adversarial unlearning to pass checks while keeping data influence.
-
Stochastic Optimization and Data Science
A survey equating offline Monte-Carlo/SAA and online stochastic-approximation sample complexities for convex stochastic optimization arising in statistics and ML.