GP-FL estimates the global Hessian at the server with Gaussian process regression on recent noisy gradient differences, claiming linear-quadratic convergence in over-the-air federated learning, but the required approximation property is assumed, not derived.
Effective Federated Adaptive Gradient Methods with Non-IID Decentralized Data
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Federated learning allows loads of edge computing devices to collaboratively learn a global model without data sharing. The analysis with partial device participation under non-IID and unbalanced data reflects more reality. In this work, we propose federated learning versions of adaptive gradient methods - Federated AGMs - which employ both the first-order and second-order momenta, to alleviate generalization performance deterioration caused by dissimilarity of data population among devices. To further improve the test performance, we compare several schemes of calibration for the adaptive learning rate, including the standard Adam calibrated by $\epsilon$, $p$-Adam, and one calibrated by an activation function. Our analysis provides the first set of theoretical results that the proposed (calibrated) Federated AGMs converge to a first-order stationary point under non-IID and unbalanced data settings for nonconvex optimization. We perform extensive experiments to compare these federated learning methods with the state-of-the-art FedAvg, FedMomentum and SCAFFOLD and to assess the different calibration schemes and the advantages of AGMs over the current federated learning methods.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2024 1verdicts
REJECT 1roles
background 1polarities
support 1representative citing papers
citing papers explorer
-
GP-FL: Model-Based Hessian Estimation for Second-Order Over-the-Air Federated Learning
GP-FL estimates the global Hessian at the server with Gaussian process regression on recent noisy gradient differences, claiming linear-quadratic convergence in over-the-air federated learning, but the required approximation property is assumed, not derived.