Pith. sign in

REVIEW 2 major objections 2 minor 50 references

DPDL perturbs cross-gradients with Gaussian noise and calibrates them via cosine similarity to achieve differential privacy while preserving linear speedup in decentralized learning on non-IID data.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-28 07:40 UTC pith:R6AJFFME

load-bearing objection The paper adds cosine-similarity calibration after Gaussian noise on cross-gradients for DP in decentralized non-IID learning, but the post-processing step likely weakens both the privacy bound and the linear-speedup claim. the 2 major comments →

arxiv 2606.04399 v1 pith:R6AJFFME submitted 2026-06-03 cs.LG cs.CR

DPDL: Towards Differential Privacy Preservation in Decentralized Stochastic Learning on Non-IID Data

classification cs.LG cs.CR
keywords differential privacydecentralized learningnon-IID datastochastic optimizationcosine similaritygradient aggregationprivacy preservationlinear speedup
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper seeks to protect individual data privacy during decentralized model training where agents exchange gradient information without any central coordinator and where local datasets follow different distributions. It introduces DPDL, which adds calibrated Gaussian noise to the cross-gradients that each agent computes from its neighbors' models on its own data, then uses cosine similarity to adjust the noisy values before a momentum-style update. The analysis supplies the smallest noise magnitude that meets a target privacy level and shows that convergence still scales linearly with the number of agents even under arbitrary non-IID partitions. A reader would care because many real distributed applications must share model information yet cannot tolerate privacy leaks or slow training when data are heterogeneous.

Core claim

DPDL uses the Gaussian noise mechanism to perturb cross-gradients before sharing and then applies cosine similarity calibration to the perturbed values so that their aggregation updates the local model in a momentum-like manner. The analysis determines the minimum noise level for a given privacy guarantee and proves that linear speedup in training is retained even with arbitrary non-IID data partitions.

What carries the argument

cosine similarity calibration applied to noisy cross-gradients before momentum-style aggregation

Load-bearing premise

The cosine similarity calibration step can be applied to noisy cross-gradients without violating the differential privacy guarantee or destroying the convergence properties that produce linear speedup under arbitrary non-IID partitions.

What would settle it

An experiment that either measures a privacy leakage exceeding the bound implied by the added noise after calibration or records sub-linear speedup on non-IID data partitions would falsify the central claims.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The minimum noise level required to reach a chosen privacy level is identified by the analysis.
  • Linear speedup in training time is retained despite non-IID data partitions.
  • The calibrated updates defend against privacy attacks while still producing accurate models.
  • The approach works in fully decentralized settings without any central server.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same calibration step could be examined for use with other noise mechanisms or aggregation rules.
  • The momentum-style update suggests the method may combine naturally with existing optimizers that already employ momentum.
  • Testing on a wider range of network topologies would reveal whether the linear speedup holds beyond the topologies studied.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes DPDL, a decentralized stochastic learning algorithm for non-IID data that perturbs cross-gradients with Gaussian noise to enforce (ε,δ)-DP and then applies cosine-similarity calibration before momentum-style aggregation; it claims a rigorous derivation of the minimum noise level needed for a target privacy budget together with a proof that linear speedup is retained under arbitrary non-IID partitions, supported by experiments on real-world datasets.

Significance. If the privacy and convergence claims hold after the calibration step, the work would supply a concrete, analyzable mechanism for trading privacy against convergence rate in fully decentralized non-IID settings, an area where existing DP-SGD analyses do not directly apply.

major comments (2)
  1. [§4] §4 (Theoretical Analysis): the derivation of the minimum noise variance σ^{2} for (ε,δ)-DP must be shown to remain valid after the deterministic cosine-similarity rescaling of the noisy cross-gradients; standard DP composition theorems do not automatically cover a subsequent data-dependent post-processing step whose output is fed into the momentum update.
  2. [§4] §4, Theorem on linear speedup: the convergence bound assumes that the calibrated updates remain unbiased (or have controlled bias) under arbitrary non-IID partitions; the cosine rescaling can introduce distribution-dependent scaling factors that violate this assumption, and the proof must explicitly bound the resulting bias term.
minor comments (2)
  1. [Experiments] The experimental section should report the exact baselines, privacy budgets, and statistical significance tests used to claim superiority over prior decentralized DP methods.
  2. [Algorithm 1] Notation for the calibrated gradient ̂g and the momentum coefficient should be introduced before the first use in the algorithm description.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments on the theoretical analysis in Section 4. We address each major comment below and indicate the corresponding revisions.

read point-by-point responses
  1. Referee: [§4] §4 (Theoretical Analysis): the derivation of the minimum noise variance σ^{2} for (ε,δ)-DP must be shown to remain valid after the deterministic cosine-similarity rescaling of the noisy cross-gradients; standard DP composition theorems do not automatically cover a subsequent data-dependent post-processing step whose output is fed into the momentum update.

    Authors: We appreciate this clarification. The Gaussian perturbation mechanism satisfies (ε,δ)-DP by construction. The subsequent cosine-similarity calibration is a deterministic post-processing function applied to the mechanism's output. By the post-processing property of differential privacy, any function of an (ε,δ)-DP output (data-dependent or otherwise) remains (ε,δ)-DP. Consequently, the minimum noise variance derived for the perturbation step continues to guarantee the target privacy budget for the calibrated cross-gradients. We will add an explicit remark invoking the post-processing theorem in Section 4 to make this connection clear. revision: yes

  2. Referee: [§4] §4, Theorem on linear speedup: the convergence bound assumes that the calibrated updates remain unbiased (or have controlled bias) under arbitrary non-IID partitions; the cosine rescaling can introduce distribution-dependent scaling factors that violate this assumption, and the proof must explicitly bound the resulting bias term.

    Authors: We agree that the cosine rescaling introduces a data-dependent scaling factor whose effect on bias must be controlled. In the existing proof we bound the deviation of the calibrated direction from the true cross-gradient using the fact that cosine similarity lies in [0,1] and the bounded gradient assumption; however, an explicit bias term arising from the scaling was not isolated. We will revise the proof of the linear-speedup theorem to introduce and bound this additional bias term, showing that it remains controlled under arbitrary non-IID partitions and does not alter the linear speedup rate. revision: yes

Circularity Check

0 steps flagged

No circularity detected; theoretical claims presented as independent analysis

full rationale

The provided abstract and context contain no equations, self-citations, or fitted parameters that reduce the claimed minimum noise level or linear speedup result to inputs by construction. The DP noise derivation and convergence bound are described as outputs of rigorous analysis rather than tautological renamings or post-hoc fits. The cosine calibration step is introduced as a deterministic post-processing technique whose interaction with DP and convergence is asserted to hold, but no load-bearing self-citation chain or self-definitional loop is visible. This is the normal case of a self-contained theoretical claim.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Only the abstract is available; no explicit free parameters, axioms, or invented entities are stated in the provided text.

pith-pipeline@v0.9.1-grok · 5798 in / 1201 out tokens · 19498 ms · 2026-06-28T07:40:58.924106+00:00 · methodology

0 comments
read the original abstract

In the paradigm of decentralized learning, a group of agents collaborate to train a global model using distributed datasets without a central server. Although the power of collaboration has been verified by many state-of-the-art studies, it entails extensive gradient information exchanging among the agents and thus induces high risk of privacy leakage for the individual agents. Moreover, in real-world applications, the training data are usually non-identically and independently distributed across the agents, inducing more challenges to enable privacy-preserved decentralized learning. To address these issues, we propose a privacy-preserved decentralized learning algorithm with non-IID data, DPDL, which leverages the notion of Differential Privacy (DP) in cross-gradient aggregation through a similarity-based calibration technique. Specifically, in each round, each agent perturbs the cross-gradients (i.e., the derivatives of its neighbors' local model in its private local data) by Gaussian noise mechanism before sharing them with its neighbors; it then adopt cosine similarity to calibrate the received perturbed cross-gradients such that the aggregation of the calibrated cross-gradients can be utilized to effectively update local model in a momentum-like manner. Our rigorous theoretical analysis not only reveals the minimum noise level required to achieve a specific level of privacy preservation, but also illustrates that our algorithm still achieves a linear speedup in training with non-IID data. We finally conduct extensive experiments on real-world dataset to validate the effectiveness of our algorithm in defending privacy attacks and in training accurate models.

Figures

Figures reproduced from arXiv: 2606.04399 by Feng Li, Lina Wang, Xue Xiao, Yunsheng Yuan.

Figure 1
Figure 1. Figure 1: Experiment results about convergence on MNIST dataset over bipartite graphs. [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Experiment results about convergence on MNIST dataset over fully connected graphs. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 1
Figure 1. Figure 1: The results indicate that increasing the number of [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Experiment results about convergence on CIFAR-10 dataset over bipartite graphs. [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Experiment results about convergence on CIFAR-10 dataset over fully connected graphs. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Images reconstructed by gradient inversion attacks on MNIST dataset. [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Images reconstructed by gradient inversion attacks on CIFAR-10 dataset. [PITH_FULL_IMAGE:figures/full_fig_p014_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Our roadmap to prove Theorem 2. Recall that g˜ i t denotes the aggregated gradient of agent i in round t, and 1 B P b g ii t,b represents its batch-averaged self-gradient. As mentioned above, we first give Lemma 6 and Lemma 7 as follows. They reveal the upper bounds of E  [PITH_FULL_IMAGE:figures/full_fig_p024_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Experiment results about convergence on MNIST dataset over ring graphs. [PITH_FULL_IMAGE:figures/full_fig_p035_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Experiment results about convergence on CIFAR-10 dataset over ring graphs. [PITH_FULL_IMAGE:figures/full_fig_p036_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 3 canonical work pages · 1 internal anchor

  1. [1]

    A Survey on Distributed Machine Learning,

    J. Verbraeken, M. Wolting, J. Katzy, J. Kloppenburg, T. Verbelen, and J. Rellermeyer, “A Survey on Distributed Machine Learning,”ACM Computing Surveys, vol. 53, no. 2, pp. 1–33, 2020. 14 (a1) Ground Truth in the 200th round (a2) Ground Truth in the 400th round (a3) Ground Truth in the 800th round (b1) ND-CG in the 200th round (b2) ND-CG in the 400th round...

  2. [2]

    A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection,

    Q. Li, Z. Wen, Z. Wu, S. Hu, N. Wang, Y . Li, X. Liu, and B. He, “A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection,”IEEE Trans. on Knowledge and Data Engineering, vol. 35, no. 4, pp. 3347–3366, 2023

  3. [3]

    Decentralized Federated Learning: Funda- mentals, State of the Art, Frameworks, Trends, and Challenges,

    E. Beltr ´an, M. P ´erez, P. S ´anchez, S. Bernal, G. Bovet, M. P ´erez, G. P ´erez, and A. Celdr ´an, “Decentralized Federated Learning: Funda- mentals, State of the Art, Frameworks, Trends, and Challenges,”IEEE Communications Surveys & Tutorials, vol. 25, no. 4, pp. 2983–3013, 15 TABLE III THE RESULTS OF GRADIENT INVERSION ATTACKS ONMNISTDATASET. Metric...

  4. [4]

    FedQV: Leveraging Quadratic V oting in Federated Learning,

    T. Chu and N. Laoutaris, “FedQV: Leveraging Quadratic V oting in Federated Learning,”Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 8, no. 2, pp. 22:1–22:36, 2024

  5. [5]

    Communication-Efficient Learning of Deep Networks from Decentral- ized Data,

    B. McMahan, E. Moore, D.Ramage, S. Hampson, and B. Arcas, “Communication-Efficient Learning of Deep Networks from Decentral- ized Data,” inProc. of the 20th AISTATS, vol. 54, 2017, pp. 1273–1282

  6. [6]

    Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,

    X. Lian, C. Zhang, H. Zhang, C. Hsieh, W. Zhang, and J. Liu, “Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,” inProc. of the 31st NIPS, 2017, p. 5336–5346

  7. [7]

    Network Topology and Communication-computation Tradeoffs in Decentralized Optimization,

    A. Nedi, A. Olshevsky, and M. Rabbat, “Network Topology and Communication-computation Tradeoffs in Decentralized Optimization,” Proceedings of the IEEE, vol. 106, no. 5, pp. 953–976, 2018

  8. [8]

    Asynchronous Decentralized Parallel Stochastic Gradient Descent,

    X. Lian, W. Zhang, C. Zhang, and J. Liu, “Asynchronous Decentralized Parallel Stochastic Gradient Descent,” inProc. of the 35th ICML, 2018, pp. 3043–3052

  9. [9]

    Quasi-global Mo- mentum: Accelerating Decentralized Deep Learning on Heterogeneous Data,

    T. Lin, S. Karimireddy, S. Stich, and M. Jaggi, “Quasi-global Mo- mentum: Accelerating Decentralized Deep Learning on Heterogeneous Data,” inProc. of the 38th ICML, 2021, pp. 6654–6665

  10. [10]

    Federated Learning on Non-iid Data Silos: An Experimental Study,

    Q. Li, Y . Diao, Q. Chen, and B. He, “Federated Learning on Non-iid Data Silos: An Experimental Study,” inProc. of the 38th IEEE ICDE, 2022, pp. 965–978

  11. [11]

    Decentralized Federated Learning via Mutual Knowledge Transfer,

    C. Li, G. Li, and P. Varshney, “Decentralized Federated Learning via Mutual Knowledge Transfer,”IEEE Internet of Things Journal, vol. 9, no. 2, pp. 1136–1147, 2021

  12. [12]

    Im- proving the Model Consistency of Decentralized Federated Learning,

    Y . Shi, L. Shen, K. Wei, Y . Sun, B. Yuan, X. Wang, and D. Tao, “Im- proving the Model Consistency of Decentralized Federated Learning,” inProc. of the 40th ICML, vol. 202, 2023, pp. 31 269–31 291

  13. [13]

    Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous Data,

    S. Aketi, A. Hashemi, and K. Roy, “Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous Data,” inProc. of the 37th NeurIPS, 2024

  14. [14]

    NET- FLEET: Achieving Linear Convergence Speedup for Fully Decentralized Federated Learning with Heterogeneous Data,

    X. Zhang, M. Fang, Z. Liu, H. Yang, J. Liu, and Z. Zhu, “NET- FLEET: Achieving Linear Convergence Speedup for Fully Decentralized Federated Learning with Heterogeneous Data,” inProc. of the 23rd ACM MobiHoc, 2022, p. 71–80

  15. [15]

    Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data,

    Y . Esfandiari, S. Tan, Z. Jiang, A. Balu, E. Herron, C. Hegde, and S. Sarkar, “Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data,” inProc. of the 38th ICML, 2021, pp. 3036–3046

  16. [16]

    On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization,

    H. Yu, R. Jin, and S. Yang, “On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization,” inProc. of the 36th ICML, 2019, pp. 7184–7193

  17. [17]

    Deep leakage from gradients,

    L. Zhu, Z. Liu, and S. Han, “Deep leakage from gradients,” inProc. of the 33rd NIPS, 2019, pp. 14 747–14 756

  18. [18]

    iDLG: Improved Deep Leakage from Gradients

    B. Zhao, K. Mopuri, and H. Bilen, “iDLG: Improved Deep Leakage from Gradients,”arXiv preprint arXiv:2001.02610, 2020

  19. [19]

    More than Enough is Too Much: Adaptive Defenses against Gradient Leakage in Production Federated Learning,

    F. Wang, E. Hugh, and B. Li, “More than Enough is Too Much: Adaptive Defenses against Gradient Leakage in Production Federated Learning,” inProc. of the 42nd INFOCOM, 2023, pp. 1–10

  20. [20]

    Differential Privacy,

    C. Dwork, “Differential Privacy,” inProc. of the 33rd ICALP, 2006, pp. 1–12

  21. [21]

    The Algorithmic Foundations of Differential Privacy,

    C. Dwork and A. Roth, “The Algorithmic Foundations of Differential Privacy,”Foundations and Trends in Theoretical Computer Science, vol. 9, pp. 211–407, 2014

  22. [22]

    How to dp-fy ml: A practical tutorial to machine learning with differential privacy,

    N. Ponomareva, S. Vassilvitskii, Z. Xu, B. McMahan, A. Kurakin, and C. Zhang, “How to dp-fy ml: A practical tutorial to machine learning with differential privacy,” inProc. of the 29th ACM SIGKDD, 2023, p. 5823–5824

  23. [23]

    Privacy-Preserving Deep Learning,

    R. Shokri and V . Shmatikov, “Privacy-Preserving Deep Learning,” in Proc. of the 22nd ACM CCS, 2015, pp. 1310–1321

  24. [24]

    Heterogeneous Differential-Private Federated Learning: Trading Privacy for Utility Truthfully,

    X. Lin, J. Wu, J. Li, C. Sang, S. Hu, and M. J. Deen, “Heterogeneous Differential-Private Federated Learning: Trading Privacy for Utility Truthfully,”IEEE Trans. on Dependable and Secure Computing, vol. 20, no. 6, pp. 5113–5129, 2023

  25. [25]

    Differentially Private Federated Learning on Non-iid Data: Convergence Analysis and Adap- tive Optimization,

    L. Chen, X. Ding, Z. Bao, P. Zhou, and H. Jin, “Differentially Private Federated Learning on Non-iid Data: Convergence Analysis and Adap- tive Optimization,”IEEE Trans. on Knowledge and Data Engineering, vol. 36, no. 9, pp. 4567–4581, 2024. 16

  26. [26]

    A(DP)2SGD: Asynchronous Decen- tralized Parallel Stochastic Gradient Descent With Differential Privacy,

    J. Xu, W. Zhang, and F. Wang, “A(DP)2SGD: Asynchronous Decen- tralized Parallel Stochastic Gradient Descent With Differential Privacy,” IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 8036–8047, 2021

  27. [27]

    Muffliato: Peer-to- Peer Privacy Amplification for Decentralized Optimization and Averag- ing,

    E. Cyffers, M. Even, A. Bellet, and L. Massouli ´e, “Muffliato: Peer-to- Peer Privacy Amplification for Decentralized Optimization and Averag- ing,” inProc. of 36th NIPS, vol. 35, 2022, pp. 15 889–15 902

  28. [28]

    Between Privacy and Utility: On Differential Privacy in Theory and Practice,

    J. Seeman and D. Susser, “Between Privacy and Utility: On Differential Privacy in Theory and Practice,”ACM Journal on Responsible Comput- ing, vol. 1, no. 1, pp. 1–18, 2024

  29. [29]

    Randomized Gossip Algorithms,

    S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah, “Randomized Gossip Algorithms,”IEEE Trans. on Information Theory, vol. 52, no. 6, pp. 2508–2530, 2006

  30. [30]

    Optimal Algorithms for Smooth and Strongly Convex Distributed Optimization in Networks,

    K. Scaman, F. Bach, S. Bubeck, Y . Lee, and L. Massouli ´e, “Optimal Algorithms for Smooth and Strongly Convex Distributed Optimization in Networks,” inProc. of the 34th ICML, 2017, pp. 3027–3036

  31. [31]

    Optimal Algorithms for Non-smooth Distributed Optimization in Networks,

    K. Scaman, F. Bach, S. Bubeck, Y . Lee, and L. Massouli ´e, “Optimal Algorithms for Non-smooth Distributed Optimization in Networks,” in Proc. of the 32nd NIPS, 2018, p. 2745–2754

  32. [32]

    Distributed Stochastic Gradient Tracking Methods,

    S. Pu and A. Nedi, “Distributed Stochastic Gradient Tracking Methods,” Mathematical Programming, vol. 187, no. 1, pp. 409–457, 2021

  33. [33]

    Momentum Tracking: Momentum Acceleration for Decentralized Deep Learning on Heterogeneous Data,

    Y . Takezawa, H. Bao, K. Niwa, R. Sato, and M. Yamada, “Momentum Tracking: Momentum Acceleration for Decentralized Deep Learning on Heterogeneous Data,”Trans. on Machine Learning Research, vol. 2023, 2023

  34. [34]

    Neighborhood Gradient Mean: An Ef- ficient Decentralized Learning Method for Non-IID Data Distributions,

    S. Aketi, S. Kodge, and K. Roy, “Neighborhood Gradient Mean: An Ef- ficient Decentralized Learning Method for Non-IID Data Distributions,” Trans. on Machine Learning Research, 2023

  35. [35]

    Differentially Private ADMM for Convex Dis- tributed Learning: Improved Accuracy via Multi-step Approximation,

    Z. Huang and Y . Gong, “Differentially Private ADMM for Convex Dis- tributed Learning: Improved Accuracy via Multi-step Approximation,” arXiv preprint arXiv:2005.07890, 2020

  36. [36]

    Pdsl: Privacy-preserved decentralized stochastic learning with heterogeneous data distribution,

    L. Wang, Y . Yuan, C. Wang, and F. Li, “Pdsl: Privacy-preserved decentralized stochastic learning with heterogeneous data distribution,” in2025 IEEE 45th ICDCS. IEEE Computer Society, 2025, pp. 736– 746

  37. [37]

    Understanding Clip- ping for Federated Learning: Convergence and Client-Level Differential Privacy,

    X. Zhang, X. Chen, M. Hong, S. Wu, and J. Yi, “Understanding Clip- ping for Federated Learning: Convergence and Client-Level Differential Privacy,” inProc. of the 39th ICML, vol. 162, 2022, pp. 26 048–26 067

  38. [38]

    What Can We Learn Privately?

    S. Kasiviswanathan, H. Lee, K. Nissim, S. Raskhodnikova, and A. Smith, “What Can We Learn Privately?” inProc. of the 49th IEEE FOCS, 2008, p. 531–540

  39. [39]

    Deep Learning with Differential Privacy,

    M. Abadi, A. Chu, I. Goodfellow, H. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep Learning with Differential Privacy,” inProc. of the 2016 ACM CCS, 2016, p. 308–318

  40. [40]

    Inverting gradients - how easy is it to break privacy in federated learning?

    J. Geiping, H. Bauermeister, H. Dr ¨oge, and M. Moeller, “Inverting gradients - how easy is it to break privacy in federated learning?” in Proceedings of the 34th NIPS, 2020

  41. [41]

    Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures,

    M. Fredrikson, S. Jha, and T. Ristenpart, “Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures,” inProc. of the 22nd ACM CCS, 2015, p. 1322–1333

  42. [42]

    On the momentum term in gradient descent learning algo- rithms,

    N. Qian, “On the momentum term in gradient descent learning algo- rithms,”Neural networks, vol. 12, no. 1, pp. 145–151, 1999

  43. [43]

    Gradient-based learning applied to document recognition,

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”Proc. of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998

  44. [44]

    Learning multiple layers of features from tiny images,

    A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009

  45. [45]

    Very Deep Convolutional Networks for Large-Scale Image Recognition

    K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,”arXiv preprint arXiv:1409.1556, 2014

  46. [46]

    Data-Free Knowledge Distillation for Heterogeneous Federated Learning,

    Z. Zhu, J. Hong, and J. Zhou, “Data-Free Knowledge Distillation for Heterogeneous Federated Learning,” inProc. of the 38th ICML, vol. 139, 2021, pp. 12 878–12 889

  47. [47]

    Federated Learning With Differential Privacy: Algorithms and Performance Analysis,

    K. Wei, J. Li, M. Ding, C. Ma, H. Yang, F. Farokhi, S. Jin, T. Quek, and H. V . Poor, “Federated Learning With Differential Privacy: Algorithms and Performance Analysis,”IEEE Trans. on Information Forensics and Security, vol. 15, pp. 3454–3469, 2020

  48. [48]

    Decentralized Wireless Federated Learning With Differential Privacy,

    S. Chen, D. Yu, Y . Zou, J. Yu, and X. Cheng, “Decentralized Wireless Federated Learning With Differential Privacy,”IEEE Trans. on Industrial Informatics, vol. 18, no. 9, pp. 6273–6282, 2022

  49. [49]

    Decen- tralized parallel sgd with privacy preservation in vehicular networks,

    D. Yu, Z. Zou, S. Chen, Y . Tao, B. Tian, W. Lv, and X. Cheng, “Decen- tralized parallel sgd with privacy preservation in vehicular networks,” IEEE Transactions on Vehicular Technology, vol. 70, no. 6, pp. 5211– 5220, 2021

  50. [50]

    Multiscale Structural Similarity for Image Quality Assessment,

    Z. Wang, E. Simoncelli, and A. Bovik, “Multiscale Structural Similarity for Image Quality Assessment,” inProc. of the 37th ACSSC, vol. 2, 2003, pp. 1398–1402. 17 APPENDIXA PROOF OFTHEOREM1 In this section, we prove the efficacy of our DPDL algorithm in preserving differential privacy. In the decentralized learning system, each agent collaborates with its ...