Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

Federated Learning: Challenges, Methods, and Future Directions

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Federated learning is a distinct learning regime whose four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—require moving beyond classical distributed optimization and…

desk verdict A solid, useful organizing survey of early federated learning; its claim that FL fundamentally departs from classical distributed learning is asserted rather than shown, but the taxonomy still earns its place. read the letter →

arxiv 1908.07873 v1 pith:X7L7I6NU submitted 2019-08-21 cs.LG cs.DCstat.ML

classification cs.LGcs.DCstat.ML
keywords federatedlearningdistributedoptimizationnon-IIDdatasystemsheterogeneitycommunicationefficiencydifferentialprivacysecureaggregationpersonalization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Federated learning trains one statistical model from data scattered across phones, hospitals, or sensors, with raw data never leaving the device. The paper's central claim is that this setting is not a routine variant of distributed optimization or privacy-preserving machine learning: the combination of expensive communication, heterogeneous hardware, non-identically distributed data, and privacy constraints changes the problem in kind, not just in degree. The article builds a four-part taxonomy that explains why classical mini-batch SGD, bounded-delay asynchronous methods, and standard differential privacy cannot be transplanted unchanged. If the taxonomy holds, it gives the field a common research agenda: every method must be judged under low participation, unreliable devices, non-IID data, and privacy leakage through shared updates.

What carries the argument

The central object is the weighted federated objective $F(w) := \sum_{k=1}^{m} p_k F_k(w)$ together with the star-network training loop it implies: selected devices perform local training, send updates to a central server, and receive the new global model in return. The paper uses this formulation as the reference point that makes the four challenges concrete: communication cost is the number and size of messages, systems heterogeneity is device capacity and dropout, statistical heterogeneity is the failure of the IID assumption behind the local objectives, and privacy is the residual information contained in the updates themselves. The article also admits alternative objectives, such as personalized multi-task or meta-learning formulations, but always measures them against this canonical problem.

What would settle it

Run a head-to-head comparison on a realistic non-IID mobile dataset: Federated Averaging versus an ordinary data-center distributed SGD with the same communication budget, the same per-round device availability, and the same differential-privacy guarantee. If the classical baseline matches or beats the federated method on accuracy and privacy across the board, the claimed fundamental departure—that federated settings demand new algorithms—would be hard to defend.

Watch

Extended reading notes

Core claim

The paper proposes that the canonical federated learning problem is to minimize a weighted sum of device-local empirical risks, $F(w) := \sum_{k=1}^{m} p_k F_k(w)$, under the constraints that local data stay on each device and only intermediate updates are communicated. Around this objective it identifies four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—and asserts that together they make federated learning fundamentally different from data-center distributed learning and classical privacy-preserving analysis. The survey then maps existing work onto each challenge, showing where classical tools carry over and where they break: local-updating methods and compression reduce communication; active sampling and fault tolerance address systems variability; meta-learning, multi-task learning, and proximal terms address non-IID data; secure aggregation and differential privacy address leakage, each at some cost to efficiency or accuracy. The contribution is therefore not a new algorithm or theorem but a problem definition and an organizing map of the solution space, plus a list of open directions for the field.

Load-bearing premise

The whole taxonomy assumes the canonical problem is a star network with one global model and local data that never leaves devices, as in Eq. (1); if real federated use cases turn out to be decentralized, personalized, or one-shot, the challenge ordering would have to change.

Editorial extensions

If this is right

  • A federated algorithm should be validated under all four constraints together, not under IID data-center assumptions, since each classical assumption is violated in the federated setting.
  • Communication-efficient techniques such as local updating and compression must be compared on a communication-accuracy Pareto frontier, because their benefits can compose or cancel in ways that single-method analyses miss.
  • Convergence guarantees for federated methods need to account for low device participation and dropped devices; FedAvg lacks guarantees and can diverge on heterogeneous data, and the paper points to proximal, multi-task, or meta-learning variants as correctives.
  • Privacy in federated learning comes with a real trade-off: secure aggregation protects updates but adds communication, while differential privacy reduces accuracy, so practical systems must explicitly balance privacy, accuracy, and efficiency.
  • Standardized benchmarks and pre-training heterogeneity diagnostics are needed before empirical results in federated learning can be reliably compared across methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the taxonomy is right, the field's evaluation culture should shift from accuracy alone to a three-way balance of accuracy, communication, and privacy, with heterogeneity as an explicit experimental axis.
  • A testable extension: diagnostic proxies for non-IID-ness that can be computed before training would let practitioners decide up front between a single global model and personalized multi-task methods.
  • The star-network, single-global-model framing may be the least durable part of the taxonomy; decentralized and personalized formulations already appear as alternatives, and if they become canonical, the relative weight of the four challenges would shift.
  • The interaction of secure aggregation, compression, and differential privacy is an open design space the survey implicitly maps but does not explore; studying these mechanisms jointly could yield new accuracy-privacy-communication trade-offs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript is a survey of federated learning. It formulates the canonical problem as minimizing a weighted sum of local objective functions over a star communication topology (Eq. (1)), identifies four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—and reviews classical and recent methods for each in Section 2 before outlining future research directions in Section 3. The paper's central thesis is that these challenges make federated learning fundamentally different from traditional distributed optimization and privacy-preserving data analysis.

Significance. The survey is well organized and extensively referenced, and it provides a genuinely useful taxonomy for a nascent field. Its strengths include a broad and balanced coverage of algorithmic, systems, and privacy topics, explicit attention to limitations of existing methods, and pointers to practical tools such as LEAF and TensorFlow Federated. Descriptive claims are consistently attributed to the literature, and the paper does not overstate what is known. Because the contribution is organizational and bibliographic rather than technical, the main risk lies in the framing of the central claim: the assertion that federated learning requires a 'fundamental departure' from standard approaches is not operationalized, and the survey's own Section 2 shows that each individual challenge has classical precedents. This is a fixable support gap rather than a technical error, but it should be addressed before publication.

major comments (2)
  1. [Abstract, §1.2, §2 (opening)] The central assertion that federated learning 'differs significantly' from traditional distributed environments and 'requires a fundamental departure' from standard approaches is never given a precise comparison class or criterion. The survey itself shows that each individual challenge has classical precedents—local-updating SGD, compression schemes, asynchronous parameter servers, fault tolerance, and differential privacy are all discussed as prior work—so the novelty must reside in the scale or in the combination of these challenges. As written, the claim is unfalsifiable: a reader cannot determine what evidence would count against it. Please either provide a systematic comparison against a well-defined baseline (for example, a table of assumptions that federated learning violates relative to data-center distributed learning) or soften the claim to refer to the combination of challenges at scale. The material for such a comparison is already present in Section 2.
  2. [§1.1 and §2.1.3] The paper defines the canonical federated learning problem via Eq. (1) and the star-network topology, and the entire four-challenge taxonomy is built on this choice. However, the choice is not defended beyond the statement that the star network is 'predominant.' If decentralized topologies or personalized multi-task objectives become the primary use cases, the taxonomy would need substantial reordering. Please state this scope restriction more prominently—ideally in the abstract and introduction—and either justify the predominance claim with citations or explicitly frame the survey as covering the star-network, single-global-model variant of federated learning. The passing acknowledgments in Section 1.1 and Section 2.1.3 are not sufficient given the generality of the paper's language elsewhere.
minor comments (5)
  1. [§2.2.2] The introductory sentence of Section 2.2 lists the three directions as '(i) asynchronous communication, (ii) active device sampling, and (ii) fault tolerance'; the second '(ii)' should be '(iii)'.
  2. [§4 (Conclusion)] The phrase 'we have outlined out a handful of open problems' contains a typo; 'out' should be deleted.
  3. [§2.4.1] Describing differential privacy as having 'strong information theoretic guarantees' is potentially misleading; the guarantees are probabilistic and compositional, not information-theoretic in the Shannon sense. Consider rewording to 'rigorous probabilistic guarantees.'
  4. [References] Several reference entries contain an erroneous space in author initials (for example, 'V . Chen', 'V . Ivanov', 'P . Richtárik'). These should be cleaned up for consistency.
  5. [§1 (Introduction)] The sentence about 'mobile user modeling and personalization [60, 90]' cites XNOR-Net [90], which is a compressed-inference paper and does not directly support the claim about mobile user modeling; please replace or add a more directly relevant reference.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper is a survey with no derivation chain whose claims reduce to their inputs.

full rationale

This paper is an expository survey, not a derivation or prediction pipeline. Its central claim that federated learning 'differs significantly' from traditional distributed environments is a framing assertion supported by a broad literature review, not by a chain of equations or fitted parameters. The survey does cite the authors' own prior work (MOCHA, FedProx, q-FFL, LEAF) as examples of current approaches, but these citations function as pointers to independent published methods, not as load-bearing evidence in an argument that reduces to self-citation. There is no objective function fitted to a subset of data and then renamed as a prediction, no uniqueness theorem imported from the authors' own prior work to force a modeling choice, and no ansatz smuggled in via citation. Each of the four challenges is explicitly tied to classical predecessors in Section 2, which further undercuts any claim that the paper's novelty depends on a circular definition. Because the survey makes no quantitative derivations and its organizational claims are not verified or refuted through the paper's own equations, none of the enumerated circularity patterns applies.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

No free parameters or invented entities appear because the paper is a survey. Its only structural assumption is the canonical star-network, single-global-model framing of federated learning.

assumptions (1)
  • domain assumption The canonical federated learning problem is the weighted-sum objective in Eq. (1) under a star topology with periodic central-server communication.
    The survey organizes all methods around minimizing F(w) = sum p_k F_k(w) with a central server and local data, treating decentralized or one-shot alternatives as secondary (Sections 1.1 and 2.1.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Federated Learning: Challenges, Methods, and Future Directions." pith.science (2026). https://pith.science/paper/X7L7I6NU

@misc{pith2026190807873,
  author       = {Pith},
  title        = {Pith review of: Federated Learning: Challenges, Methods, and Future Directions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X7L7I6NU}},
  note         = {Machine review of arXiv:1908.07873}
}
read the original abstract

Federated learning involves training statistical models over remote devices or siloed data centers, such as mobile phones or hospitals, while keeping data localized. Training in heterogeneous and potentially massive networks introduces novel challenges that require a fundamental departure from standard approaches for large-scale machine learning, distributed optimization, and privacy-preserving data analysis. In this article, we discuss the unique characteristics and challenges of federated learning, provide a broad overview of current approaches, and outline several directions of future work that are relevant to a wide range of research communities.

Figures

Figures reproduced from arXiv: 1908.07873 by the authors.

Figure 1
Figure 1. An example application of federated learning for the task of next-word prediction on mobile [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Left: Distributed (mini-batch) SGD. Each device, k, locally computes gradients from a mini-batch of data points to approximate ∇Fk (w), and the aggregated mini-batch updates are applied on the server. Right: Local updating schemes. Each device immediately applies local updates, e.g., gradients, after they are computed and a server performs a global aggregation after a variable number of local updates. Local-updating… view at source ↗
Figure 3
Figure 3. Centralized vs. decentralized topologies. In the typical federated learning setting and as a focus of [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Systems heterogeneity in federated learning. Devices may vary in terms of network connection, [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Different modeling approaches in federated networks. Depending on properties of the data, [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: An illustration of different privacy-enhancing mechanisms in one round of federated learning. [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Gradient Compression and Correlation Driven Federated Learning for Wireless Traffic Prediction

    cs.DC 2025-01 conditional novelty 4.0 of 10

    A federated learning method that compresses gradients and weights model aggregation by gradient correlation attains state-of-the-art traffic prediction at roughly one-fortieth of the communication cost.

Reference graph

Works this paper leans on

141 extracted references · 52 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL https://www.tensorflow.org/federated

    Tensorflow federated: Machine learning on decentralized data. URL https://www.tensorflow.org/federated

  2. [2]

    Abadi, A

    M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In Conference on Computer and Communications Security, 2016

  3. [3]

    Agarwal, A

    N. Agarwal, A. T. Suresh, F. X. X. Yu, S. Kumar, and B. McMahan. cpSGD: Communication-efficient and differentially-private distributed sgd. In Advances in Neural Information Processing Systems, 2018

  4. [4]

    Agrawal and R

    R. Agrawal and R. Srikant. Privacy-preserving data mining. In International Conference on Management of Data, 2000

  5. [5]

    Ammad-ud din, E

    M. Ammad-ud din, E. Ivannikova, S. A. Khan, W. Oyomno, Q. Fu, K. E. Tan, and A. Flanagan. Federated collab- orative filtering for privacy-preserving personalized recommendation system. arXiv preprint arXiv:1901.09888, 2019

  6. [6]

    Anguita, A

    D. Anguita, A. Ghio, L. Oneto, X. Parra, and J. L. Reyes-Ortiz. A public domain dataset for human activity recognition using smartphones. In European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, 2013

  7. [7]

    Bassily, A

    R. Bassily, A. Smith, and A. Thakurta. Private empirical risk minimization: Efficient algorithms and tight error bounds. In Foundations of Computer Science, 2014

  8. [8]

    Bhowmick, J

    A. Bhowmick, J. Duchi, J. Freudiger, G. Kapoor, and R. Rogers. Protection against reconstruction and its applications in private federated learning. arXiv preprint arXiv:1812.00984, 2018

Show all 141 references
  1. [9]

    Bolukbasi, J

    T. Bolukbasi, J. Wang, O. Dekel, and V . Saligrama. Adaptive neural networks for efficient inference. InInternational Conference on Machine Learning, 2017

  2. [10]

    Bonawitz, V

    K. Bonawitz, V . Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth. Prac- tical secure aggregation for privacy-preserving machine learning. In Conference on Computer and Communications Security, 2017

  3. [11]

    Bonawitz, H

    K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V . Ivanov, C. Kiddon, J. Konecny, S. Mazzocchi, H. B. McMahan, T. V . Overveldt, D. Petrou, D. Ramage, and J. Roselander. Towards federated learning at scale: system design. In Conference on Systems and Machine Lear...

  4. [12]

    Bonomi, R

    F. Bonomi, R. Milito, J. Zhu, and S. Addepalli. Fog computing and its role in the internet of things. In SIGCOMM Workshop on Mobile Cloud Computing, 2012. 14

  5. [13]

    R. Bost, R. A. Popa, S. Tu, and S. Goldwasser. Machine learning classification over encrypted data. In Network and Distributed System Security Symposium, 2015

  6. [14]

    S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein. Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends R© in Machine Learning, 3:1–122, 2011

  7. [15]

    Caldas, J

    S. Caldas, J. Koneˇ cny, H. B. McMahan, and A. Talwalkar. Expanding the reach of federated learning by reducing client resource requirements. arXiv preprint arXiv:1812.07210, 2018

  8. [16]

    Caldas, P

    S. Caldas, P . Wu, T. Li, J. Koneˇ cn`y, H. B. McMahan, V . Smith, and A. Talwalkar. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018

  9. [17]

    Carlini, C

    N. Carlini, C. Liu, J. Kos, Ú. Erlingsson, and D. Song. The secret sharer: Measuring unintended neural network memorization & extracting secrets. arXiv preprint arXiv:1802.08232, 2018

  10. [18]

    R. Caruana. Multitask learning. Machine Learning, 28:41–75, 1997

  11. [19]

    Castro, B

    M. Castro, B. Liskov, et al. Practical byzantine fault tolerance. In Operating Systems Design and Implementation, 1999

  12. [20]

    Charles and D

    Z. Charles and D. Papailiopoulos. Gradient coding using the stochastic block model. In International Symposium on Information Theory, 2018

  13. [21]

    Z. B. Charles, D. S. Papailiopoulos, and J. Ellenberg. Approximate gradient coding via sparse random graphs. arXiv preprint arXiv:1711.0677, 2017

  14. [22]

    Chaudhuri, C

    K. Chaudhuri, C. Monteleoni, and A. D. Sarwate. Differentially private empirical risk minimization. Journal of Machine Learning Research, 12:1069–1109, 2011

  15. [23]

    D. Chaum. The dining cryptographers problem: Unconditional sender and recipient untraceability. Journal of Cryptology, 1:65–75, 1988

  16. [24]

    F. Chen, Z. Dong, Z. Li, and X. He. Federated meta-learning for recommendation. arXiv preprint arXiv:1802.07876, 2018

  17. [25]

    V . Chen, V . Pastro, and M. Raykova. Secure computation for machine learning with spdz. arXiv preprint arXiv:1901.00329, 2019

  18. [26]

    Corinzia and J

    L. Corinzia and J. M. Buhmann. Variational federated multi-task learning. arXiv preprint arXiv:1906.06268, 2019

  19. [27]

    W. Dai, A. Kumar, J. Wei, Q. Ho, G. Gibson, and E. P . Xing. High-performance distributed ML at scale through parameter server consistency models. In AAAI Conference on Artificial Intelligence, 2015

  20. [28]

    Dekel, R

    O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao. Optimal distributed online prediction using mini-batches. Journal of Machine Learning Research, 13:165–202, 2012

  21. [29]

    Deshpande, C

    A. Deshpande, C. Guestrin, S. R. Madden, J. M. Hellerstein, and W. Hong. Model-based approximate querying in sensor networks. The VLDB Journal, 14:417–443, 2005

  22. [30]

    Duchi, M

    J. Duchi, M. I. Jordan, and B. McMahan. Estimation, optimization, and parallelism when data is sparse. In Advances in Neural Information Processing Systems, 2013

  23. [31]

    J. C. Duchi, M. I. Jordan, and M. J. Wainwright. Privacy aware learning. In Advances in Neural Information Processing Systems, 2012

  24. [32]

    C. Dwork. A firm foundation for private data analysis. Communications of the ACM, 54:86–95, 2011

  25. [33]

    Dwork and A

    C. Dwork and A. Roth. The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9:211–407, 2014. 15

  26. [34]

    Dwork, F

    C. Dwork, F. McSherry, K. Nissim, and A. Smith. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography Conference, 2006

  27. [35]

    Eichner, T

    H. Eichner, T. Koren, H. B. McMahan, N. Srebro, and K. Talwar. Semi-cyclic stochastic gradient descent. In International Conference on Machine Learning, 2019

  28. [36]

    El Emam and F

    K. El Emam and F. K. Dankar. Protecting privacy using k-anonymity. Journal of the American Medical Informatics Association, 15:627–637, 2008

  29. [37]

    Evgeniou and M

    T. Evgeniou and M. Pontil. Regularized multi–task learning. In Conference on Knowledge Discovery and Data Mining, 2004

  30. [38]

    Feldman, I

    V . Feldman, I. Mironov, K. Talwar, and A. Thakurta. Privacy amplification by iteration. In Foundations of Computer Science, 2018

  31. [39]

    Fredrikson, S

    M. Fredrikson, S. Jha, and T. Ristenpart. Model inversion attacks that exploit confidence information and basic countermeasures. In Conference on Computer and Communications Security, 2015

  32. [40]

    Garcia Lopez, A

    P . Garcia Lopez, A. Montresor, D. Epema, A. Datta, T. Higashino, A. Iamnitchi, M. Barcellos, P . Felber, and E. Riviere. Edge-centric computing: Vision and challenges. SIGCOMM Computer Communication Review, 45:37–42, 2015

  33. [41]

    R. C. Geyer, T. Klein, and M. Nabi. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557, 2017

  34. [42]

    Ghazi, R

    B. Ghazi, R. Pagh, and A. Velingker. Scalable and differentially private distributed aggregation in the shuffled model. arXiv preprint arXiv:1906.08320, 2019

  35. [43]

    Goryczka and L

    S. Goryczka and L. Xiong. A comprehensive comparison of multiparty secure additions with differential privacy. IEEE Transactions on Dependable and Secure Computing, 14:463–477, 2015

  36. [44]

    Guha and V

    N. Guha and V . Smith. Model aggregation via good-enough model spaces. arXiv preprint arXiv:1805.07782, 2018

  37. [45]

    N. Guha, A. Talwalkar, and V . Smith. One-shot federated learning. arXiv preprint arXiv:1902.11175, 2019

  38. [46]

    A. Hard, K. Rao, R. Mathews, F. Beaufays, S. Augenstein, H. Eichner, C. Kiddon, and D. Ramage. Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604, 2018

  39. [47]

    L. He, A. Bian, and M. Jaggi. Cola: Decentralized linear learning. In Advances in Neural Information Processing Systems, 2018

  40. [48]

    Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P . B. Gibbons, G. A. Gibson, G. Ganger, and E. P . Xing. More effective distributed ML via a stale synchronous parallel parameter server. In Advances in Neural Information Processing Systems, 2013

  41. [49]

    K. Hong, D. Lillethun, U. Ramachandran, B. Ottenwälder, and B. Koldehofe. Mobile fog: A programming model for large-scale applications on the internet of things. In SIGCOMM Workshop on Mobile Cloud Computing, 2013

  42. [50]

    Huang, F

    J. Huang, F. Qian, Y. Guo, Y. Zhou, Q. Xu, Z. M. Mao, S. Sen, and O. Spatscheck. An in-depth study of lte: effect of network protocol and application behavior on performance. SIGCOMM Computer Communication Review, 43: 363–374, 2013

  43. [51]

    Huang and D

    L. Huang and D. Liu. Patient clustering improves efficiency of federated machine learning to predict mortality and hospital stay time using distributed electronic medical records. arXiv preprint arXiv:1903.09296, 2019

  44. [52]

    Huang, Y

    L. Huang, Y. Yin, Z. Fu, S. Zhang, H. Deng, and D. Liu. Loadaboost: Loss-based adaboost federated machine learning on medical data. arXiv preprint arXiv:1811.12629, 2018

  45. [53]

    Iyengar, J

    R. Iyengar, J. P . Near, D. Song, O. Thakkar, A. Thakurta, and L. Wang. Towards practical differentially private convex optimization. In Conference on Computer and Communications Security, 2019. 16

  46. [54]

    Jaggi, V

    M. Jaggi, V . Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan. Communication-efficient distributed dual coordinate ascent. In Advances in Neural Information Processing Systems, 2014

  47. [55]

    Jeong, S

    E. Jeong, S. Oh, H. Kim, J. Park, M. Bennis, and S.-L. Kim. Communication-efficient on-device machine learning: Federated distillation and augmentation under non-iid private data. arXiv preprint arXiv:1811.11479, 2018

  48. [56]

    Jiang and G

    P . Jiang and G. Agrawal. A linear speedup analysis of distributed deep learning with sparse and quantized communication. In Advances in Neural Information Processing Systems, 2018

  49. [57]

    J. Kang, Z. Xiong, D. Niyato, H. Yu, Y.-C. Liang, and D. I. Kim. Incentive design for efficient federated learning in mobile networks: A contract theory approach. arXiv preprint arXiv:1905.07479, 2019

  50. [58]

    Khodak, M.-F

    M. Khodak, M.-F. Balcan, and A. Talwalkar. Adaptive gradient-based meta-learning methods. arXiv preprint arXiv:1906.02717, 2019

  51. [59]

    Koneˇ cn`y, H

    J. Koneˇ cn`y, H. B. McMahan, F. X. Yu, P . Richtárik, A. T. Suresh, and D. Bacon. Federated learning: strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492, 2016

  52. [60]

    Kuflik, J

    T. Kuflik, J. Kay, and B. Kummerfeld. Challenges and solutions of ubiquitous user modeling. In Ubiquitous Display Environments. 2012

  53. [61]

    Lalitha, X

    A. Lalitha, X. Wang, O. Kilinc, Y. Lu, T. Javidi, and F. Koushanfar. Decentralized bayesian learning over graphs. arXiv preprint arXiv:1905.10466, 2019

  54. [62]

    Lee and D

    C.-P . Lee and D. Roth. Distributed box-constrained quadratic optimization for dual linear svm. In International Conference on Machine Learning, 2015

  55. [63]

    K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran. Speeding up distributed machine learning using codes. IEEE Transactions on Information Theory, 64:1514–1529, 2017

  56. [64]

    J. Li, M. Khodak, S. Caldas, and A. Talwalkar. Differentially-private gradient-based meta-learning. Technical Report, 2019

  57. [65]

    T. Li, A. K. Sahu, M. Sanjabi, M. Zaheer, A. Talwalkar, and V . Smith. Federated optimization for heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018

  58. [66]

    T. Li, M. Sanjabi, and V . Smith. Fair resource allocation in federated learning. arXiv preprint arXiv:1905.10497, 2019

  59. [67]

    X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu. Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent. In Advances in Neural Information Processing Systems, 2017

  60. [68]

    T. Lin, S. U. Stich, and M. Jaggi. Don’t use large mini-batches, use local sgd. arXiv preprint arXiv:1808.07217, 2018

  61. [69]

    Lindell and B

    Y. Lindell and B. Pinkas. Privacy preserving data mining. In Advances in Cryptology, 2000

  62. [70]

    L. Liu, J. Zhang, S. Song, and K. B. Letaief. Edge-assisted hierarchical federated learning with non-iid data. arXiv preprint arXiv:1905.06641, 2019

  63. [71]

    Y. Liu, J. K. Muppala, M. Veeraraghavan, D. Lin, and M. Hamdi. Data center networks: Topologies, architectures and fault-tolerance characteristics. Springer Science & Business Media, 2013

  64. [72]

    C. Ma, V . Smith, M. Jaggi, M. I. Jordan, P . Richtárik, and M. Takᡠc. Adding vs. averaging in distributed primal-dual optimization. In International Conference on Machine Learning, 2015

  65. [73]

    L. W. Mackey, M. I. Jordan, and A. Talwalkar. Divide-and-conquer matrix factorization. In Advances in Neural Information Processing Systems, 2011

  66. [74]

    S. R. Madden, M. J. Franklin, J. M. Hellerstein, and W. Hong. Tinydb: an acquisitional query processing system for sensor networks. Transactions on Database Systems, 30:122–173, 2005. 17

  67. [75]

    H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas. Communication-efficient learning of deep networks from decentralized data. In Conference on Artificial Intelligence and Statistics , 2017

  68. [76]

    H. B. McMahan, D. Ramage, K. Talwar, and L. Zhang. Learning differentially private recurrent language models. In International Conference on Learning Representations, 2018

  69. [77]

    Melis, G

    L. Melis, G. Danezis, and E. D. Cristofaro. Efficient private statistics with succinct sketches. In Network and Distributed System Security Symposium, 2016

  70. [78]

    Melis, C

    L. Melis, C. Song, E. De Cristofaro, and V . Shmatikov. Exploiting unintended feature leakage in collaborative learning. In IEEE Symposium on Security & Privacy , 2019

  71. [79]

    Mohassel and P

    P . Mohassel and P . Rindal. Aby 3: a mixed protocol framework for machine learning. In Conference on Computer and Communications Security, 2018

  72. [80]

    Mohri, G

    M. Mohri, G. Sivek, and A. T. Suresh. Agnostic federated learning. In International Conference on Machine Learning, 2019

  73. [81]

    M. E. Nergiz and C. Clifton. δ-presence without complete world knowledge. IEEE Transactions on Knowledge and Data Engineering, 22:868–883, 2010

  74. [82]

    Nikolaenko, U

    V . Nikolaenko, U. Weinsberg, S. Ioannidis, M. Joye, D. Boneh, and N. Taft. Privacy-preserving ridge regression on hundreds of millions of records. In Symposium on Security and Privacy, 2013

  75. [83]

    Nishio and R

    T. Nishio and R. Yonetani. Client selection for federated learning with heterogeneous resources in mobile edge. In International Conference on Communications, 2019

  76. [84]

    Pantelopoulos and N

    A. Pantelopoulos and N. G. Bourbakis. A survey on wearable sensor-based systems for health monitoring and prognosis. IEEE Transactions on Systems, Man, and Cybernetics, 40:1–12, 2010

  77. [85]

    Papernot, M

    N. Papernot, M. Abadi, U. Erlingsson, I. Goodfellow, and K. Talwar. Semi-supervised knowledge transfer for deep learning from private training data. In International Conference on Learning Representations, 2017

  78. [86]

    Papernot, S

    N. Papernot, S. Song, I. Mironov, A. Raghunathan, K. Talwar, and Ú. Erlingsson. Scalable private learning with pate. In International Conference on Learning Representations, 2018

  79. [87]

    A. Qiao, B. Aragam, B. Zhang, and E. Xing. Fault tolerance in iterative-convergent machine learning. In International Conference on Machine Learning, 2019

  80. [88]

    Z. Qu, P . Richtárik, and T. Zhang. Quartz: Randomized dual coordinate ascent with arbitrary sampling. In Advances in Neural Information Processing Systems, 2015

  81. [89]

    Ramaswamy, R

    S. Ramaswamy, R. Mathews, K. Rao, and F. Beaufays. Federated learning for emoji prediction in a mobile keyboard. arXiv preprint arXiv:1906.04329, 2019

  82. [90]

    Rastegari, V

    M. Rastegari, V . Ordonez, J. Redmon, and A. Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In European Conference on Computer Vision, 2016

  83. [91]

    Ratner et al

    A. Ratner et al. SysML: The new frontier of machine learning systems. arXiv preprint arXiv:1904.03257, 2019

  84. [92]

    Recht, C

    B. Recht, C. Re, S. Wright, and F. Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems, 2011

  85. [93]

    S. J. Reddi, J. Koneˇ cn`y, P . Richtárik, B. Póczós, and A. Smola. Aide: Fast and communication efficient distributed optimization. arXiv preprint arXiv:1608.06879, 2016

  86. [94]

    Reisizadeh, S

    A. Reisizadeh, S. Prakash, R. Pedarsani, and A. S. Avestimehr. Coded computation over heterogeneous clusters. IEEE Transactions on Information Theory, 65:4227–4242, 2019. 18

  87. [95]

    M. S. Riazi, C. Weinert, O. Tkachenko, E. M. Songhori, T. Schneider, and F. Koushanfar. Chameleon: A hybrid secure computation framework for machine learning applications. In Asia Conference on Computer and Communications Security, 2018

  88. [96]

    Richtárik and M

    P . Richtárik and M. Takᡠc. Distributed coordinate descent method for learning with big data.Journal of Machine Learning Research, 17:2657–2681, 2016

  89. [97]

    B. D. Rouhani, M. S. Riazi, and F. Koushanfar. Deepsecure: Scalable provably-secure deep learning. In Design Automation Conference, 2018

  90. [98]

    Samarakoon, M

    S. Samarakoon, M. Bennis, W. Saad, and M. Debbah. Federated learning for ultra-reliable low-latency v2v communications. In Global Communications Conference, 2018

  91. [99]

    Sattler, S

    F. Sattler, S. Wiedemann, K.-R. Müller, and W. Samek. Robust and communication-efficient federated learning from non-iid data. arXiv preprint arXiv:1903.02891, 2019

  92. [100]

    Schmidt and N

    M. Schmidt and N. L. Roux. Fast convergence of stochastic gradient descent under a strong growth condition. arXiv preprint arXiv:1308.6370, 2013

  93. [101]

    Seide, H

    F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu. 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In International Speech Communication Association, 2014

  94. [102]

    Shalev-Shwartz and T

    S. Shalev-Shwartz and T. Zhang. Accelerated mini-batch stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, 2013

  95. [103]

    Shamir and N

    O. Shamir and N. Srebro. Distributed stochastic optimization and learning. InAllerton Conference on Communication, Control, and Computing, 2014

  96. [104]

    Shamir, N

    O. Shamir, N. Srebro, and T. Zhang. Communication-efficient distributed optimization using an approximate newton-type method. In International Conference on Machine Learning, 2014

  97. [105]

    Silva, B

    S. Silva, B. Gutman, E. Romero, P . M. Thompson, A. Altmann, and M. Lorenzi. Federated learning in distributed medical databases: Meta-analysis of large-scale subcortical brain data. arXiv preprint arXiv:1810.08553, 2018

  98. [106]

    Smith, C.-K

    V . Smith, C.-K. Chiang, M. Sanjabi, and A. Talwalkar. Federated multi-task learning. In Advances in Neural Information Processing Systems, 2017

  99. [107]

    Smith, S

    V . Smith, S. Forte, C. Ma, M. Takac, M. I. Jordan, and M. Jaggi. Cocoa: a general framework for communication- efficient distributed optimization. Journal of Machine Learning Research, 18:1–47, 2018

  100. [108]

    S. U. Stich. Local sgd converges fast and communicates little. In International Conference on Learning Representations, 2019

  101. [109]

    Tandon, Q

    R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis. Gradient coding: Avoiding stragglers in distributed learning. In International Conference on Machine Learning, 2017

  102. [110]

    A. S. Tanenbaum and M. Van Steen. Distributed systems: principles and paradigms . Prentice-Hall, 2007

  103. [111]

    H. Tang, S. Gan, C. Zhang, T. Zhang, and J. Liu. Communication compression for decentralized training. In Advances in Neural Information Processing Systems, 2018

  104. [112]

    H. Tang, C. Yu, C. Renggli, S. Kassing, A. Singla, D. Alistarh, J. Liu, and C. Zhang. Distributed learning over unreliable networks. In International Conference on Machine Learning, 2019

  105. [113]

    Thakkar, G

    O. Thakkar, G. Andrew, and H. B. McMahan. Differentially private learning with adaptive clipping. arXiv preprint arXiv:1905.03871, 2019

  106. [114]

    Thrun and L

    S. Thrun and L. Pratt. Learning to learn. Springer Science & Business Media, 2012

  107. [115]

    Van Berkel

    C. Van Berkel. Multi-core for mobile phones. In Conference on Design, Automation and Test in Europe, 2009. 19

  108. [116]

    Vaswani, F

    S. Vaswani, F. Bach, and M. Schmidt. Fast and faster convergence of sgd for over-parameterized models (and an accelerated perceptron). In Conference on Artificial Intelligence and Statistics , 2019

  109. [117]

    Vepakomma, O

    P . Vepakomma, O. Gupta, A. Dubey, and R. Raskar. Reducing leakage in distributed deep learning for sensitive health data. arXiv preprint arXiv:1812.00564, 2019

  110. [118]

    Wagner and D

    I. Wagner and D. Eckhoff. Technical privacy metrics: a systematic survey. ACM Computing Surveys, 51:57, 2018

  111. [119]

    H. Wang, S. Sievert, S. Liu, Z. Charles, D. Papailiopoulos, and S. Wright. Atomo: Communication-efficient learning via atomic sparsification. In Advances in Neural Information Processing Systems, 2018

  112. [120]

    Wang and G

    J. Wang and G. Joshi. Cooperative sgd: A unified framework for the design and analysis of communication- efficient sgd algorithms. arXiv preprint arXiv:1808.07576, 2018

  113. [121]

    Wang and G

    J. Wang and G. Joshi. Adaptive communication strategies to achieve the best error-runtime trade-off in local- update sgd. In Conference on Systems and Machine Learning, 2019

  114. [122]

    S. Wang, F. Roosta-Khorasani, P . Xu, and M. W. Mahoney. Giant: Globally improved approximate newton method for distributed optimization. In Advances in Neural Information Processing Systems, 2018

  115. [123]

    S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and K. Chan. Adaptive federated learning in resource constrained edge computing systems. Journal on Selected Areas in Communications, 37:1205–1221, 2019

  116. [124]

    Federated learning white paper v1.0

    WeBank AI Group. Federated learning white paper v1.0. 2018

  117. [125]

    Woodworth, J

    B. Woodworth, J. Wang, A. Smith, B. McMahan, and N. Srebro. Graph oracle models, lower bounds, and gaps for parallel stochastic optimization. In Advances in Neural Information Processing Systems, 2018

  118. [126]

    X. Wu, F. Li, A. Kumar, K. Chaudhuri, S. Jha, and J. Naughton. Bolt-on differential privacy for scalable stochastic gradient descent-based analytics. In International Conference on Management of Data, 2017

  119. [127]

    Q. Yang, Y. Liu, T. Chen, and Y. Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology, 10:12, 2019

  120. [128]

    T. Yang. Trading computation for communication: Distributed stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, 2013

  121. [129]

    Y. Yao, L. Rosasco, and A. Caponnetto. On early stopping in gradient descent learning. Constructive Approximation, 26:289–315, 2007

  122. [130]

    D. Yin, A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P . Bartlett. Gradient diversity: a key ingredient for scalable distributed learning. In Conference on Artificial Intelligence and Statistics , pages 1998–2007, 2018

  123. [131]

    H. Yu, S. Yang, and S. Zhu. Parallel restarted sgd for non-convex optimization with faster convergence and less communication. In AAAI Conference on Artificial Intelligence, 2018

  124. [132]

    H. Yu, R. Jin, and S. Yang. On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization. In International Conference on Machine Learning, 2019

  125. [133]

    Yuan and S

    J. Yuan and S. Yu. Privacy preserving back-propagation neural network learning made practical with cloud computing. IEEE Transactions on Parallel and Distributed Systems, 25:212–221, 2013

  126. [134]

    Yurochkin, M

    M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, T. N. Hoang, and Y. Khazaeni. Bayesian nonparametric federated learning of neural networks. In International Conference on Machine Learning, 2019

  127. [135]

    Zhang, J

    H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang. ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning. In International Conference on Machine Learning, 2017

  128. [136]

    Zhang, A

    S. Zhang, A. E. Choromanska, and Y. LeCun. Deep learning with elastic averaging sgd. In Advances in Neural Information Processing Systems, 2015. 20

  129. [137]

    Zhang, J

    Y. Zhang, J. Duchi, and M. Wainwright. Divide and conquer kernel ridge regression: A distributed algorithm with minimax optimal rates. Journal of Machine Learning Research, 16:3299–3340, 2015

  130. [138]

    Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V . Chandra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018

  131. [139]

    Y. Zhao, J. Zhao, L. Jiang, R. Tan, and D. Niyato. Mobile edge computing, blockchain and reputation-based crowdsourcing iot federated learning: A secure, decentralized and privacy-preserving system. arXiv preprint arXiv:1906.10893, 2019

  132. [140]

    Zhou and G

    F. Zhou and G. Cong. On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization. In International Joint Conference on Artificial Intelligence , 2018

  133. [141]

    Zinkevich, M

    M. Zinkevich, M. Weimer, L. Li, and A. J. Smola. Parallelized stochastic gradient descent. In Advances in Neural Information Processing Systems, 2010. 21

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.