REVIEW 2 major objections 5 minor 1 cited by
Federated Learning: Challenges, Methods, and Future Directions
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Federated learning is a distinct learning regime whose four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—require moving beyond classical distributed optimization and…
desk verdict A solid, useful organizing survey of early federated learning; its claim that FL fundamentally departs from classical distributed learning is asserted rather than shown, but the taxonomy still earns its place. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the weighted federated objective $F(w) := \sum_{k=1}^{m} p_k F_k(w)$ together with the star-network training loop it implies: selected devices perform local training, send updates to a central server, and receive the new global model in return. The paper uses this formulation as the reference point that makes the four challenges concrete: communication cost is the number and size of messages, systems heterogeneity is device capacity and dropout, statistical heterogeneity is the failure of the IID assumption behind the local objectives, and privacy is the residual information contained in the updates themselves. The article also admits alternative objectives, such as personalized multi-task or meta-learning formulations, but always measures them against this canonical problem.
What would settle it
Run a head-to-head comparison on a realistic non-IID mobile dataset: Federated Averaging versus an ordinary data-center distributed SGD with the same communication budget, the same per-round device availability, and the same differential-privacy guarantee. If the classical baseline matches or beats the federated method on accuracy and privacy across the board, the claimed fundamental departure—that federated settings demand new algorithms—would be hard to defend.
Extended reading notes
Core claim
The paper proposes that the canonical federated learning problem is to minimize a weighted sum of device-local empirical risks, $F(w) := \sum_{k=1}^{m} p_k F_k(w)$, under the constraints that local data stay on each device and only intermediate updates are communicated. Around this objective it identifies four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—and asserts that together they make federated learning fundamentally different from data-center distributed learning and classical privacy-preserving analysis. The survey then maps existing work onto each challenge, showing where classical tools carry over and where they break: local-updating methods and compression reduce communication; active sampling and fault tolerance address systems variability; meta-learning, multi-task learning, and proximal terms address non-IID data; secure aggregation and differential privacy address leakage, each at some cost to efficiency or accuracy. The contribution is therefore not a new algorithm or theorem but a problem definition and an organizing map of the solution space, plus a list of open directions for the field.
Load-bearing premise
The whole taxonomy assumes the canonical problem is a star network with one global model and local data that never leaves devices, as in Eq. (1); if real federated use cases turn out to be decentralized, personalized, or one-shot, the challenge ordering would have to change.
Editorial extensions
If this is right
- A federated algorithm should be validated under all four constraints together, not under IID data-center assumptions, since each classical assumption is violated in the federated setting.
- Communication-efficient techniques such as local updating and compression must be compared on a communication-accuracy Pareto frontier, because their benefits can compose or cancel in ways that single-method analyses miss.
- Convergence guarantees for federated methods need to account for low device participation and dropped devices; FedAvg lacks guarantees and can diverge on heterogeneous data, and the paper points to proximal, multi-task, or meta-learning variants as correctives.
- Privacy in federated learning comes with a real trade-off: secure aggregation protects updates but adds communication, while differential privacy reduces accuracy, so practical systems must explicitly balance privacy, accuracy, and efficiency.
- Standardized benchmarks and pre-training heterogeneity diagnostics are needed before empirical results in federated learning can be reliably compared across methods.
Reading between the lines
- If the taxonomy is right, the field's evaluation culture should shift from accuracy alone to a three-way balance of accuracy, communication, and privacy, with heterogeneity as an explicit experimental axis.
- A testable extension: diagnostic proxies for non-IID-ness that can be computed before training would let practitioners decide up front between a single global model and personalized multi-task methods.
- The star-network, single-global-model framing may be the least durable part of the taxonomy; decentralized and personalized formulations already appear as alternatives, and if they become canonical, the relative weight of the four challenges would shift.
- The interaction of secure aggregation, compression, and differential privacy is an open design space the survey implicitly maps but does not explore; studying these mechanisms jointly could yield new accuracy-privacy-communication trade-offs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of federated learning. It formulates the canonical problem as minimizing a weighted sum of local objective functions over a star communication topology (Eq. (1)), identifies four core challenges—expensive communication, systems heterogeneity, statistical heterogeneity, and privacy—and reviews classical and recent methods for each in Section 2 before outlining future research directions in Section 3. The paper's central thesis is that these challenges make federated learning fundamentally different from traditional distributed optimization and privacy-preserving data analysis.
Significance. The survey is well organized and extensively referenced, and it provides a genuinely useful taxonomy for a nascent field. Its strengths include a broad and balanced coverage of algorithmic, systems, and privacy topics, explicit attention to limitations of existing methods, and pointers to practical tools such as LEAF and TensorFlow Federated. Descriptive claims are consistently attributed to the literature, and the paper does not overstate what is known. Because the contribution is organizational and bibliographic rather than technical, the main risk lies in the framing of the central claim: the assertion that federated learning requires a 'fundamental departure' from standard approaches is not operationalized, and the survey's own Section 2 shows that each individual challenge has classical precedents. This is a fixable support gap rather than a technical error, but it should be addressed before publication.
major comments (2)
- [Abstract, §1.2, §2 (opening)] The central assertion that federated learning 'differs significantly' from traditional distributed environments and 'requires a fundamental departure' from standard approaches is never given a precise comparison class or criterion. The survey itself shows that each individual challenge has classical precedents—local-updating SGD, compression schemes, asynchronous parameter servers, fault tolerance, and differential privacy are all discussed as prior work—so the novelty must reside in the scale or in the combination of these challenges. As written, the claim is unfalsifiable: a reader cannot determine what evidence would count against it. Please either provide a systematic comparison against a well-defined baseline (for example, a table of assumptions that federated learning violates relative to data-center distributed learning) or soften the claim to refer to the combination of challenges at scale. The material for such a comparison is already present in Section 2.
- [§1.1 and §2.1.3] The paper defines the canonical federated learning problem via Eq. (1) and the star-network topology, and the entire four-challenge taxonomy is built on this choice. However, the choice is not defended beyond the statement that the star network is 'predominant.' If decentralized topologies or personalized multi-task objectives become the primary use cases, the taxonomy would need substantial reordering. Please state this scope restriction more prominently—ideally in the abstract and introduction—and either justify the predominance claim with citations or explicitly frame the survey as covering the star-network, single-global-model variant of federated learning. The passing acknowledgments in Section 1.1 and Section 2.1.3 are not sufficient given the generality of the paper's language elsewhere.
minor comments (5)
- [§2.2.2] The introductory sentence of Section 2.2 lists the three directions as '(i) asynchronous communication, (ii) active device sampling, and (ii) fault tolerance'; the second '(ii)' should be '(iii)'.
- [§4 (Conclusion)] The phrase 'we have outlined out a handful of open problems' contains a typo; 'out' should be deleted.
- [§2.4.1] Describing differential privacy as having 'strong information theoretic guarantees' is potentially misleading; the guarantees are probabilistic and compositional, not information-theoretic in the Shannon sense. Consider rewording to 'rigorous probabilistic guarantees.'
- [References] Several reference entries contain an erroneous space in author initials (for example, 'V . Chen', 'V . Ivanov', 'P . Richtárik'). These should be cleaned up for consistency.
- [§1 (Introduction)] The sentence about 'mobile user modeling and personalization [60, 90]' cites XNOR-Net [90], which is a compressed-inference paper and does not directly support the claim about mobile user modeling; please replace or add a more directly relevant reference.
Circularity Check
No circularity: the paper is a survey with no derivation chain whose claims reduce to their inputs.
full rationale
This paper is an expository survey, not a derivation or prediction pipeline. Its central claim that federated learning 'differs significantly' from traditional distributed environments is a framing assertion supported by a broad literature review, not by a chain of equations or fitted parameters. The survey does cite the authors' own prior work (MOCHA, FedProx, q-FFL, LEAF) as examples of current approaches, but these citations function as pointers to independent published methods, not as load-bearing evidence in an argument that reduces to self-citation. There is no objective function fitted to a subset of data and then renamed as a prediction, no uniqueness theorem imported from the authors' own prior work to force a modeling choice, and no ansatz smuggled in via citation. Each of the four challenges is explicitly tied to classical predecessors in Section 2, which further undercuts any claim that the paper's novelty depends on a circular definition. Because the survey makes no quantitative derivations and its organizational claims are not verified or refuted through the paper's own equations, none of the enumerated circularity patterns applies.
Assumptions & free parameters
assumptions (1)
- domain assumption The canonical federated learning problem is the weighted-sum objective in Eq. (1) under a star topology with periodic central-server communication.
Cite this review
Pith. "Pith review of Federated Learning: Challenges, Methods, and Future Directions." pith.science (2026). https://pith.science/paper/X7L7I6NU
@misc{pith2026190807873,
author = {Pith},
title = {Pith review of: Federated Learning: Challenges, Methods, and Future Directions},
year = {2026},
howpublished = {\url{https://pith.science/paper/X7L7I6NU}},
note = {Machine review of arXiv:1908.07873}
}
read the original abstract
Federated learning involves training statistical models over remote devices or siloed data centers, such as mobile phones or hospitals, while keeping data localized. Training in heterogeneous and potentially massive networks introduces novel challenges that require a fundamental departure from standard approaches for large-scale machine learning, distributed optimization, and privacy-preserving data analysis. In this article, we discuss the unique characteristics and challenges of federated learning, provide a broad overview of current approaches, and outline several directions of future work that are relevant to a wide range of research communities.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Gradient Compression and Correlation Driven Federated Learning for Wireless Traffic Prediction
A federated learning method that compresses gradients and weights model aggregation by gradient correlation attains state-of-the-art traffic prediction at roughly one-fortieth of the communication cost.
Reference graph
Works this paper leans on
-
[1]
URL https://www.tensorflow.org/federated
Tensorflow federated: Machine learning on decentralized data. URL https://www.tensorflow.org/federated
-
[2]
Abadi, A
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang. Deep learning with differential privacy. In Conference on Computer and Communications Security, 2016
2016
-
[3]
Agarwal, A
N. Agarwal, A. T. Suresh, F. X. X. Yu, S. Kumar, and B. McMahan. cpSGD: Communication-efficient and differentially-private distributed sgd. In Advances in Neural Information Processing Systems, 2018
2018
-
[4]
Agrawal and R
R. Agrawal and R. Srikant. Privacy-preserving data mining. In International Conference on Management of Data, 2000
2000
-
[5]
M. Ammad-ud din, E. Ivannikova, S. A. Khan, W. Oyomno, Q. Fu, K. E. Tan, and A. Flanagan. Federated collab- orative filtering for privacy-preserving personalized recommendation system. arXiv preprint arXiv:1901.09888, 2019
arXiv 1901
-
[6]
Anguita, A
D. Anguita, A. Ghio, L. Oneto, X. Parra, and J. L. Reyes-Ortiz. A public domain dataset for human activity recognition using smartphones. In European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning, 2013
2013
-
[7]
Bassily, A
R. Bassily, A. Smith, and A. Thakurta. Private empirical risk minimization: Efficient algorithms and tight error bounds. In Foundations of Computer Science, 2014
2014
-
[8]
A. Bhowmick, J. Duchi, J. Freudiger, G. Kapoor, and R. Rogers. Protection against reconstruction and its applications in private federated learning. arXiv preprint arXiv:1812.00984, 2018
arXiv 2018
Show all 141 references
-
[9]
Bolukbasi, J
T. Bolukbasi, J. Wang, O. Dekel, and V . Saligrama. Adaptive neural networks for efficient inference. InInternational Conference on Machine Learning, 2017
2017
-
[10]
Bonawitz, V
K. Bonawitz, V . Ivanov, B. Kreuter, A. Marcedone, H. B. McMahan, S. Patel, D. Ramage, A. Segal, and K. Seth. Prac- tical secure aggregation for privacy-preserving machine learning. In Conference on Computer and Communications Security, 2017
2017
-
[11]
Bonawitz, H
K. Bonawitz, H. Eichner, W. Grieskamp, D. Huba, A. Ingerman, V . Ivanov, C. Kiddon, J. Konecny, S. Mazzocchi, H. B. McMahan, T. V . Overveldt, D. Petrou, D. Ramage, and J. Roselander. Towards federated learning at scale: system design. In Conference on Systems and Machine Lear...
2019
-
[12]
Bonomi, R
F. Bonomi, R. Milito, J. Zhu, and S. Addepalli. Fog computing and its role in the internet of things. In SIGCOMM Workshop on Mobile Cloud Computing, 2012. 14
2012
-
[13]
R. Bost, R. A. Popa, S. Tu, and S. Goldwasser. Machine learning classification over encrypted data. In Network and Distributed System Security Symposium, 2015
2015
-
[14]
S. Boyd, N. Parikh, E. Chu, B. Peleato, and J. Eckstein. Distributed optimization and statistical learning via the alternating direction method of multipliers. Foundations and Trends R© in Machine Learning, 3:1–122, 2011
2011
-
[15]
Caldas, J
S. Caldas, J. Koneˇ cny, H. B. McMahan, and A. Talwalkar. Expanding the reach of federated learning by reducing client resource requirements. arXiv preprint arXiv:1812.07210, 2018
2018 arXiv
-
[16]
Caldas, P
S. Caldas, P . Wu, T. Li, J. Koneˇ cn`y, H. B. McMahan, V . Smith, and A. Talwalkar. Leaf: A benchmark for federated settings. arXiv preprint arXiv:1812.01097, 2018
2018 arXiv
-
[17]
Carlini, C
N. Carlini, C. Liu, J. Kos, Ú. Erlingsson, and D. Song. The secret sharer: Measuring unintended neural network memorization & extracting secrets. arXiv preprint arXiv:1802.08232, 2018
2018 arXiv
-
[18]
R. Caruana. Multitask learning. Machine Learning, 28:41–75, 1997
1997
-
[19]
Castro, B
M. Castro, B. Liskov, et al. Practical byzantine fault tolerance. In Operating Systems Design and Implementation, 1999
1999
-
[20]
Charles and D
Z. Charles and D. Papailiopoulos. Gradient coding using the stochastic block model. In International Symposium on Information Theory, 2018
2018
-
[21]
Z. B. Charles, D. S. Papailiopoulos, and J. Ellenberg. Approximate gradient coding via sparse random graphs. arXiv preprint arXiv:1711.0677, 2017
2017
-
[22]
Chaudhuri, C
K. Chaudhuri, C. Monteleoni, and A. D. Sarwate. Differentially private empirical risk minimization. Journal of Machine Learning Research, 12:1069–1109, 2011
2011
-
[23]
D. Chaum. The dining cryptographers problem: Unconditional sender and recipient untraceability. Journal of Cryptology, 1:65–75, 1988
1988
-
[24]
F. Chen, Z. Dong, Z. Li, and X. He. Federated meta-learning for recommendation. arXiv preprint arXiv:1802.07876, 2018
2018 arXiv
-
[25]
V . Chen, V . Pastro, and M. Raykova. Secure computation for machine learning with spdz. arXiv preprint arXiv:1901.00329, 2019
1901 arXiv
-
[26]
Corinzia and J
L. Corinzia and J. M. Buhmann. Variational federated multi-task learning. arXiv preprint arXiv:1906.06268, 2019
1906 arXiv
-
[27]
W. Dai, A. Kumar, J. Wei, Q. Ho, G. Gibson, and E. P . Xing. High-performance distributed ML at scale through parameter server consistency models. In AAAI Conference on Artificial Intelligence, 2015
2015
-
[28]
Dekel, R
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao. Optimal distributed online prediction using mini-batches. Journal of Machine Learning Research, 13:165–202, 2012
2012
-
[29]
Deshpande, C
A. Deshpande, C. Guestrin, S. R. Madden, J. M. Hellerstein, and W. Hong. Model-based approximate querying in sensor networks. The VLDB Journal, 14:417–443, 2005
2005
-
[30]
Duchi, M
J. Duchi, M. I. Jordan, and B. McMahan. Estimation, optimization, and parallelism when data is sparse. In Advances in Neural Information Processing Systems, 2013
2013
-
[31]
J. C. Duchi, M. I. Jordan, and M. J. Wainwright. Privacy aware learning. In Advances in Neural Information Processing Systems, 2012
2012
-
[32]
C. Dwork. A firm foundation for private data analysis. Communications of the ACM, 54:86–95, 2011
2011
-
[33]
Dwork and A
C. Dwork and A. Roth. The algorithmic foundations of differential privacy. Foundations and Trends in Theoretical Computer Science, 9:211–407, 2014. 15
2014
-
[34]
Dwork, F
C. Dwork, F. McSherry, K. Nissim, and A. Smith. Calibrating noise to sensitivity in private data analysis. In Theory of Cryptography Conference, 2006
2006
-
[35]
Eichner, T
H. Eichner, T. Koren, H. B. McMahan, N. Srebro, and K. Talwar. Semi-cyclic stochastic gradient descent. In International Conference on Machine Learning, 2019
2019
-
[36]
El Emam and F
K. El Emam and F. K. Dankar. Protecting privacy using k-anonymity. Journal of the American Medical Informatics Association, 15:627–637, 2008
2008
-
[37]
Evgeniou and M
T. Evgeniou and M. Pontil. Regularized multi–task learning. In Conference on Knowledge Discovery and Data Mining, 2004
2004
-
[38]
Feldman, I
V . Feldman, I. Mironov, K. Talwar, and A. Thakurta. Privacy amplification by iteration. In Foundations of Computer Science, 2018
2018
-
[39]
Fredrikson, S
M. Fredrikson, S. Jha, and T. Ristenpart. Model inversion attacks that exploit confidence information and basic countermeasures. In Conference on Computer and Communications Security, 2015
2015
-
[40]
Garcia Lopez, A
P . Garcia Lopez, A. Montresor, D. Epema, A. Datta, T. Higashino, A. Iamnitchi, M. Barcellos, P . Felber, and E. Riviere. Edge-centric computing: Vision and challenges. SIGCOMM Computer Communication Review, 45:37–42, 2015
2015
-
[41]
R. C. Geyer, T. Klein, and M. Nabi. Differentially private federated learning: A client level perspective. arXiv preprint arXiv:1712.07557, 2017
2017 arXiv
-
[42]
Ghazi, R
B. Ghazi, R. Pagh, and A. Velingker. Scalable and differentially private distributed aggregation in the shuffled model. arXiv preprint arXiv:1906.08320, 2019
1906 arXiv
-
[43]
Goryczka and L
S. Goryczka and L. Xiong. A comprehensive comparison of multiparty secure additions with differential privacy. IEEE Transactions on Dependable and Secure Computing, 14:463–477, 2015
2015
-
[44]
Guha and V
N. Guha and V . Smith. Model aggregation via good-enough model spaces. arXiv preprint arXiv:1805.07782, 2018
2018 arXiv
-
[45]
N. Guha, A. Talwalkar, and V . Smith. One-shot federated learning. arXiv preprint arXiv:1902.11175, 2019
1902 arXiv
-
[46]
A. Hard, K. Rao, R. Mathews, F. Beaufays, S. Augenstein, H. Eichner, C. Kiddon, and D. Ramage. Federated learning for mobile keyboard prediction. arXiv preprint arXiv:1811.03604, 2018
2018 arXiv
-
[47]
L. He, A. Bian, and M. Jaggi. Cola: Decentralized linear learning. In Advances in Neural Information Processing Systems, 2018
2018
-
[48]
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P . B. Gibbons, G. A. Gibson, G. Ganger, and E. P . Xing. More effective distributed ML via a stale synchronous parallel parameter server. In Advances in Neural Information Processing Systems, 2013
2013
-
[49]
K. Hong, D. Lillethun, U. Ramachandran, B. Ottenwälder, and B. Koldehofe. Mobile fog: A programming model for large-scale applications on the internet of things. In SIGCOMM Workshop on Mobile Cloud Computing, 2013
2013
-
[50]
Huang, F
J. Huang, F. Qian, Y. Guo, Y. Zhou, Q. Xu, Z. M. Mao, S. Sen, and O. Spatscheck. An in-depth study of lte: effect of network protocol and application behavior on performance. SIGCOMM Computer Communication Review, 43: 363–374, 2013
2013
-
[51]
Huang and D
L. Huang and D. Liu. Patient clustering improves efficiency of federated machine learning to predict mortality and hospital stay time using distributed electronic medical records. arXiv preprint arXiv:1903.09296, 2019
1903 arXiv
-
[52]
Huang, Y
L. Huang, Y. Yin, Z. Fu, S. Zhang, H. Deng, and D. Liu. Loadaboost: Loss-based adaboost federated machine learning on medical data. arXiv preprint arXiv:1811.12629, 2018
2018 arXiv
-
[53]
Iyengar, J
R. Iyengar, J. P . Near, D. Song, O. Thakkar, A. Thakurta, and L. Wang. Towards practical differentially private convex optimization. In Conference on Computer and Communications Security, 2019. 16
2019
-
[54]
Jaggi, V
M. Jaggi, V . Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan. Communication-efficient distributed dual coordinate ascent. In Advances in Neural Information Processing Systems, 2014
2014
-
[55]
Jeong, S
E. Jeong, S. Oh, H. Kim, J. Park, M. Bennis, and S.-L. Kim. Communication-efficient on-device machine learning: Federated distillation and augmentation under non-iid private data. arXiv preprint arXiv:1811.11479, 2018
2018 arXiv
-
[56]
Jiang and G
P . Jiang and G. Agrawal. A linear speedup analysis of distributed deep learning with sparse and quantized communication. In Advances in Neural Information Processing Systems, 2018
2018
-
[57]
J. Kang, Z. Xiong, D. Niyato, H. Yu, Y.-C. Liang, and D. I. Kim. Incentive design for efficient federated learning in mobile networks: A contract theory approach. arXiv preprint arXiv:1905.07479, 2019
1905 arXiv
-
[58]
Khodak, M.-F
M. Khodak, M.-F. Balcan, and A. Talwalkar. Adaptive gradient-based meta-learning methods. arXiv preprint arXiv:1906.02717, 2019
1906 arXiv
-
[59]
Koneˇ cn`y, H
J. Koneˇ cn`y, H. B. McMahan, F. X. Yu, P . Richtárik, A. T. Suresh, and D. Bacon. Federated learning: strategies for improving communication efficiency. arXiv preprint arXiv:1610.05492, 2016
2016 arXiv
-
[60]
Kuflik, J
T. Kuflik, J. Kay, and B. Kummerfeld. Challenges and solutions of ubiquitous user modeling. In Ubiquitous Display Environments. 2012
2012
-
[61]
Lalitha, X
A. Lalitha, X. Wang, O. Kilinc, Y. Lu, T. Javidi, and F. Koushanfar. Decentralized bayesian learning over graphs. arXiv preprint arXiv:1905.10466, 2019
1905 arXiv
-
[62]
Lee and D
C.-P . Lee and D. Roth. Distributed box-constrained quadratic optimization for dual linear svm. In International Conference on Machine Learning, 2015
2015
-
[63]
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran. Speeding up distributed machine learning using codes. IEEE Transactions on Information Theory, 64:1514–1529, 2017
2017
-
[64]
J. Li, M. Khodak, S. Caldas, and A. Talwalkar. Differentially-private gradient-based meta-learning. Technical Report, 2019
2019
-
[65]
T. Li, A. K. Sahu, M. Sanjabi, M. Zaheer, A. Talwalkar, and V . Smith. Federated optimization for heterogeneous networks. arXiv preprint arXiv:1812.06127, 2018
2018 arXiv
-
[66]
T. Li, M. Sanjabi, and V . Smith. Fair resource allocation in federated learning. arXiv preprint arXiv:1905.10497, 2019
1905 arXiv
-
[67]
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu. Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent. In Advances in Neural Information Processing Systems, 2017
2017
-
[68]
T. Lin, S. U. Stich, and M. Jaggi. Don’t use large mini-batches, use local sgd. arXiv preprint arXiv:1808.07217, 2018
2018 arXiv
-
[69]
Lindell and B
Y. Lindell and B. Pinkas. Privacy preserving data mining. In Advances in Cryptology, 2000
2000
-
[70]
L. Liu, J. Zhang, S. Song, and K. B. Letaief. Edge-assisted hierarchical federated learning with non-iid data. arXiv preprint arXiv:1905.06641, 2019
1905 arXiv
-
[71]
Y. Liu, J. K. Muppala, M. Veeraraghavan, D. Lin, and M. Hamdi. Data center networks: Topologies, architectures and fault-tolerance characteristics. Springer Science & Business Media, 2013
2013
-
[72]
C. Ma, V . Smith, M. Jaggi, M. I. Jordan, P . Richtárik, and M. Takᡠc. Adding vs. averaging in distributed primal-dual optimization. In International Conference on Machine Learning, 2015
2015
-
[73]
L. W. Mackey, M. I. Jordan, and A. Talwalkar. Divide-and-conquer matrix factorization. In Advances in Neural Information Processing Systems, 2011
2011
-
[74]
S. R. Madden, M. J. Franklin, J. M. Hellerstein, and W. Hong. Tinydb: an acquisitional query processing system for sensor networks. Transactions on Database Systems, 30:122–173, 2005. 17
2005
-
[75]
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas. Communication-efficient learning of deep networks from decentralized data. In Conference on Artificial Intelligence and Statistics , 2017
2017
-
[76]
H. B. McMahan, D. Ramage, K. Talwar, and L. Zhang. Learning differentially private recurrent language models. In International Conference on Learning Representations, 2018
2018
-
[77]
Melis, G
L. Melis, G. Danezis, and E. D. Cristofaro. Efficient private statistics with succinct sketches. In Network and Distributed System Security Symposium, 2016
2016
-
[78]
Melis, C
L. Melis, C. Song, E. De Cristofaro, and V . Shmatikov. Exploiting unintended feature leakage in collaborative learning. In IEEE Symposium on Security & Privacy , 2019
2019
-
[79]
Mohassel and P
P . Mohassel and P . Rindal. Aby 3: a mixed protocol framework for machine learning. In Conference on Computer and Communications Security, 2018
2018
-
[80]
Mohri, G
M. Mohri, G. Sivek, and A. T. Suresh. Agnostic federated learning. In International Conference on Machine Learning, 2019
2019
-
[81]
M. E. Nergiz and C. Clifton. δ-presence without complete world knowledge. IEEE Transactions on Knowledge and Data Engineering, 22:868–883, 2010
2010
-
[82]
Nikolaenko, U
V . Nikolaenko, U. Weinsberg, S. Ioannidis, M. Joye, D. Boneh, and N. Taft. Privacy-preserving ridge regression on hundreds of millions of records. In Symposium on Security and Privacy, 2013
2013
-
[83]
Nishio and R
T. Nishio and R. Yonetani. Client selection for federated learning with heterogeneous resources in mobile edge. In International Conference on Communications, 2019
2019
-
[84]
Pantelopoulos and N
A. Pantelopoulos and N. G. Bourbakis. A survey on wearable sensor-based systems for health monitoring and prognosis. IEEE Transactions on Systems, Man, and Cybernetics, 40:1–12, 2010
2010
-
[85]
Papernot, M
N. Papernot, M. Abadi, U. Erlingsson, I. Goodfellow, and K. Talwar. Semi-supervised knowledge transfer for deep learning from private training data. In International Conference on Learning Representations, 2017
2017
-
[86]
Papernot, S
N. Papernot, S. Song, I. Mironov, A. Raghunathan, K. Talwar, and Ú. Erlingsson. Scalable private learning with pate. In International Conference on Learning Representations, 2018
2018
-
[87]
A. Qiao, B. Aragam, B. Zhang, and E. Xing. Fault tolerance in iterative-convergent machine learning. In International Conference on Machine Learning, 2019
2019
-
[88]
Z. Qu, P . Richtárik, and T. Zhang. Quartz: Randomized dual coordinate ascent with arbitrary sampling. In Advances in Neural Information Processing Systems, 2015
2015
-
[89]
Ramaswamy, R
S. Ramaswamy, R. Mathews, K. Rao, and F. Beaufays. Federated learning for emoji prediction in a mobile keyboard. arXiv preprint arXiv:1906.04329, 2019
1906 arXiv
-
[90]
Rastegari, V
M. Rastegari, V . Ordonez, J. Redmon, and A. Farhadi. Xnor-net: Imagenet classification using binary convolutional neural networks. In European Conference on Computer Vision, 2016
2016
-
[91]
Ratner et al
A. Ratner et al. SysML: The new frontier of machine learning systems. arXiv preprint arXiv:1904.03257, 2019
1904 arXiv
-
[92]
Recht, C
B. Recht, C. Re, S. Wright, and F. Niu. Hogwild: A lock-free approach to parallelizing stochastic gradient descent. In Advances in Neural Information Processing Systems, 2011
2011
-
[93]
S. J. Reddi, J. Koneˇ cn`y, P . Richtárik, B. Póczós, and A. Smola. Aide: Fast and communication efficient distributed optimization. arXiv preprint arXiv:1608.06879, 2016
2016 arXiv
-
[94]
Reisizadeh, S
A. Reisizadeh, S. Prakash, R. Pedarsani, and A. S. Avestimehr. Coded computation over heterogeneous clusters. IEEE Transactions on Information Theory, 65:4227–4242, 2019. 18
2019
-
[95]
M. S. Riazi, C. Weinert, O. Tkachenko, E. M. Songhori, T. Schneider, and F. Koushanfar. Chameleon: A hybrid secure computation framework for machine learning applications. In Asia Conference on Computer and Communications Security, 2018
2018
-
[96]
Richtárik and M
P . Richtárik and M. Takᡠc. Distributed coordinate descent method for learning with big data.Journal of Machine Learning Research, 17:2657–2681, 2016
2016
-
[97]
B. D. Rouhani, M. S. Riazi, and F. Koushanfar. Deepsecure: Scalable provably-secure deep learning. In Design Automation Conference, 2018
2018
-
[98]
Samarakoon, M
S. Samarakoon, M. Bennis, W. Saad, and M. Debbah. Federated learning for ultra-reliable low-latency v2v communications. In Global Communications Conference, 2018
2018
-
[99]
Sattler, S
F. Sattler, S. Wiedemann, K.-R. Müller, and W. Samek. Robust and communication-efficient federated learning from non-iid data. arXiv preprint arXiv:1903.02891, 2019
1903 arXiv
-
[100]
Schmidt and N
M. Schmidt and N. L. Roux. Fast convergence of stochastic gradient descent under a strong growth condition. arXiv preprint arXiv:1308.6370, 2013
2013 arXiv
-
[101]
Seide, H
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu. 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In International Speech Communication Association, 2014
2014
-
[102]
Shalev-Shwartz and T
S. Shalev-Shwartz and T. Zhang. Accelerated mini-batch stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, 2013
2013
-
[103]
Shamir and N
O. Shamir and N. Srebro. Distributed stochastic optimization and learning. InAllerton Conference on Communication, Control, and Computing, 2014
2014
-
[104]
Shamir, N
O. Shamir, N. Srebro, and T. Zhang. Communication-efficient distributed optimization using an approximate newton-type method. In International Conference on Machine Learning, 2014
2014
-
[105]
Silva, B
S. Silva, B. Gutman, E. Romero, P . M. Thompson, A. Altmann, and M. Lorenzi. Federated learning in distributed medical databases: Meta-analysis of large-scale subcortical brain data. arXiv preprint arXiv:1810.08553, 2018
2018 arXiv
-
[106]
Smith, C.-K
V . Smith, C.-K. Chiang, M. Sanjabi, and A. Talwalkar. Federated multi-task learning. In Advances in Neural Information Processing Systems, 2017
2017
-
[107]
Smith, S
V . Smith, S. Forte, C. Ma, M. Takac, M. I. Jordan, and M. Jaggi. Cocoa: a general framework for communication- efficient distributed optimization. Journal of Machine Learning Research, 18:1–47, 2018
2018
-
[108]
S. U. Stich. Local sgd converges fast and communicates little. In International Conference on Learning Representations, 2019
2019
-
[109]
Tandon, Q
R. Tandon, Q. Lei, A. G. Dimakis, and N. Karampatziakis. Gradient coding: Avoiding stragglers in distributed learning. In International Conference on Machine Learning, 2017
2017
-
[110]
A. S. Tanenbaum and M. Van Steen. Distributed systems: principles and paradigms . Prentice-Hall, 2007
2007
-
[111]
H. Tang, S. Gan, C. Zhang, T. Zhang, and J. Liu. Communication compression for decentralized training. In Advances in Neural Information Processing Systems, 2018
2018
-
[112]
H. Tang, C. Yu, C. Renggli, S. Kassing, A. Singla, D. Alistarh, J. Liu, and C. Zhang. Distributed learning over unreliable networks. In International Conference on Machine Learning, 2019
2019
-
[113]
Thakkar, G
O. Thakkar, G. Andrew, and H. B. McMahan. Differentially private learning with adaptive clipping. arXiv preprint arXiv:1905.03871, 2019
1905 arXiv
-
[114]
Thrun and L
S. Thrun and L. Pratt. Learning to learn. Springer Science & Business Media, 2012
2012
-
[115]
Van Berkel
C. Van Berkel. Multi-core for mobile phones. In Conference on Design, Automation and Test in Europe, 2009. 19
2009
-
[116]
Vaswani, F
S. Vaswani, F. Bach, and M. Schmidt. Fast and faster convergence of sgd for over-parameterized models (and an accelerated perceptron). In Conference on Artificial Intelligence and Statistics , 2019
2019
-
[117]
Vepakomma, O
P . Vepakomma, O. Gupta, A. Dubey, and R. Raskar. Reducing leakage in distributed deep learning for sensitive health data. arXiv preprint arXiv:1812.00564, 2019
2019 arXiv
-
[118]
Wagner and D
I. Wagner and D. Eckhoff. Technical privacy metrics: a systematic survey. ACM Computing Surveys, 51:57, 2018
2018
-
[119]
H. Wang, S. Sievert, S. Liu, Z. Charles, D. Papailiopoulos, and S. Wright. Atomo: Communication-efficient learning via atomic sparsification. In Advances in Neural Information Processing Systems, 2018
2018
-
[120]
Wang and G
J. Wang and G. Joshi. Cooperative sgd: A unified framework for the design and analysis of communication- efficient sgd algorithms. arXiv preprint arXiv:1808.07576, 2018
2018 arXiv
-
[121]
Wang and G
J. Wang and G. Joshi. Adaptive communication strategies to achieve the best error-runtime trade-off in local- update sgd. In Conference on Systems and Machine Learning, 2019
2019
-
[122]
S. Wang, F. Roosta-Khorasani, P . Xu, and M. W. Mahoney. Giant: Globally improved approximate newton method for distributed optimization. In Advances in Neural Information Processing Systems, 2018
2018
-
[123]
S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and K. Chan. Adaptive federated learning in resource constrained edge computing systems. Journal on Selected Areas in Communications, 37:1205–1221, 2019
2019
-
[124]
Federated learning white paper v1.0
WeBank AI Group. Federated learning white paper v1.0. 2018
2018
-
[125]
Woodworth, J
B. Woodworth, J. Wang, A. Smith, B. McMahan, and N. Srebro. Graph oracle models, lower bounds, and gaps for parallel stochastic optimization. In Advances in Neural Information Processing Systems, 2018
2018
-
[126]
X. Wu, F. Li, A. Kumar, K. Chaudhuri, S. Jha, and J. Naughton. Bolt-on differential privacy for scalable stochastic gradient descent-based analytics. In International Conference on Management of Data, 2017
2017
-
[127]
Q. Yang, Y. Liu, T. Chen, and Y. Tong. Federated machine learning: Concept and applications. ACM Transactions on Intelligent Systems and Technology, 10:12, 2019
2019
-
[128]
T. Yang. Trading computation for communication: Distributed stochastic dual coordinate ascent. In Advances in Neural Information Processing Systems, 2013
2013
-
[129]
Y. Yao, L. Rosasco, and A. Caponnetto. On early stopping in gradient descent learning. Constructive Approximation, 26:289–315, 2007
2007
-
[130]
D. Yin, A. Pananjady, M. Lam, D. Papailiopoulos, K. Ramchandran, and P . Bartlett. Gradient diversity: a key ingredient for scalable distributed learning. In Conference on Artificial Intelligence and Statistics , pages 1998–2007, 2018
1998
-
[131]
H. Yu, S. Yang, and S. Zhu. Parallel restarted sgd for non-convex optimization with faster convergence and less communication. In AAAI Conference on Artificial Intelligence, 2018
2018
-
[132]
H. Yu, R. Jin, and S. Yang. On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization. In International Conference on Machine Learning, 2019
2019
-
[133]
Yuan and S
J. Yuan and S. Yu. Privacy preserving back-propagation neural network learning made practical with cloud computing. IEEE Transactions on Parallel and Distributed Systems, 25:212–221, 2013
2013
-
[134]
Yurochkin, M
M. Yurochkin, M. Agarwal, S. Ghosh, K. Greenewald, T. N. Hoang, and Y. Khazaeni. Bayesian nonparametric federated learning of neural networks. In International Conference on Machine Learning, 2019
2019
-
[135]
Zhang, J
H. Zhang, J. Li, K. Kara, D. Alistarh, J. Liu, and C. Zhang. ZipML: Training linear models with end-to-end low precision, and a little bit of deep learning. In International Conference on Machine Learning, 2017
2017
-
[136]
Zhang, A
S. Zhang, A. E. Choromanska, and Y. LeCun. Deep learning with elastic averaging sgd. In Advances in Neural Information Processing Systems, 2015. 20
2015
-
[137]
Zhang, J
Y. Zhang, J. Duchi, and M. Wainwright. Divide and conquer kernel ridge regression: A distributed algorithm with minimax optimal rates. Journal of Machine Learning Research, 16:3299–3340, 2015
2015
-
[138]
Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V . Chandra. Federated learning with non-iid data. arXiv preprint arXiv:1806.00582, 2018
2018 arXiv
-
[139]
Y. Zhao, J. Zhao, L. Jiang, R. Tan, and D. Niyato. Mobile edge computing, blockchain and reputation-based crowdsourcing iot federated learning: A secure, decentralized and privacy-preserving system. arXiv preprint arXiv:1906.10893, 2019
1906 arXiv
-
[140]
Zhou and G
F. Zhou and G. Cong. On the convergence properties of a k-step averaging stochastic gradient descent algorithm for nonconvex optimization. In International Joint Conference on Artificial Intelligence , 2018
2018
-
[141]
Zinkevich, M
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola. Parallelized stochastic gradient descent. In Advances in Neural Information Processing Systems, 2010. 21
2010
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.