Pith. sign in

REVIEW 3 cited by

Machine Learning at the Wireless Edge: Distributed Stochastic Gradient Descent Over-the-Air

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.00844 v3 pith:MMIBGYFP submitted 2019-01-03 cs.DC cs.ITcs.LGmath.IT

classification cs.DCcs.ITcs.LGmath.IT
keywords gradientdevicesa-dsgdwirelessd-dsgdestimatesbandwidthchannel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We study federated machine learning (ML) at the wireless edge, where power- and bandwidth-limited wireless devices with local datasets carry out distributed stochastic gradient descent (DSGD) with the help of a remote parameter server (PS). Standard approaches assume separate computation and communication, where local gradient estimates are compressed and transmitted to the PS over orthogonal links. Following this digital approach, we introduce D-DSGD, in which the wireless devices employ gradient quantization and error accumulation, and transmit their gradient estimates to the PS over a multiple access channel (MAC). We then introduce a novel analog scheme, called A-DSGD, which exploits the additive nature of the wireless MAC for over-the-air gradient computation, and provide convergence analysis for this approach. In A-DSGD, the devices first sparsify their gradient estimates, and then project them to a lower dimensional space imposed by the available channel bandwidth. These projections are sent directly over the MAC without employing any digital code. Numerical results show that A-DSGD converges faster than D-DSGD thanks to its more efficient use of the limited bandwidth and the natural alignment of the gradient estimates over the channel. The improvement is particularly compelling at low power and low bandwidth regimes. We also illustrate for a classification problem that, A-DSGD is more robust to bias in data distribution across devices, while D-DSGD significantly outperforms other digital schemes in the literature. We also observe that both D-DSGD and A-DSGD perform better by increasing the number of devices (while keeping the total dataset size constant), showing their ability in harnessing the computation power of edge devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hierarchical Federated Learning Across Heterogeneous Cellular Networks

    cs.LG 2019-09 conditional novelty 6.0 of 10

    A hierarchical federated learning framework with gradient sparsification reduces modeled communication latency in heterogeneous cellular networks while keeping CIFAR-10 accuracy close to a flat baseline.

  2. On Analog Gradient Descent Learning over Multiple Access Fading Channels

    cs.LG 2019-08 conditional novelty 6.0 of 10

    The GBMA algorithm lets distributed nodes transmit analog gradients over a fading multiple access channel without power control, and provably approaches centralized gradient descent convergence as the number of nodes grows.

  3. Over-the-Air Computation Systems: Optimization, Analysis and Scaling Laws

    cs.IT 2019-09 conditional novelty 5.0 of 10

    For a single-antenna over-the-air computation system, the paper gives a closed-form optimal transmit-receive policy and proves the average computation error decays as O(1/√K).

Pith tools