REVIEW 2 major objections 2 minor 50 references
DPDL perturbs cross-gradients with Gaussian noise and calibrates them via cosine similarity to achieve differential privacy while preserving linear speedup in decentralized learning on non-IID data.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-28 07:40 UTC pith:R6AJFFME
load-bearing objection The paper adds cosine-similarity calibration after Gaussian noise on cross-gradients for DP in decentralized non-IID learning, but the post-processing step likely weakens both the privacy bound and the linear-speedup claim. the 2 major comments →
DPDL: Towards Differential Privacy Preservation in Decentralized Stochastic Learning on Non-IID Data
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
DPDL uses the Gaussian noise mechanism to perturb cross-gradients before sharing and then applies cosine similarity calibration to the perturbed values so that their aggregation updates the local model in a momentum-like manner. The analysis determines the minimum noise level for a given privacy guarantee and proves that linear speedup in training is retained even with arbitrary non-IID data partitions.
What carries the argument
cosine similarity calibration applied to noisy cross-gradients before momentum-style aggregation
Load-bearing premise
The cosine similarity calibration step can be applied to noisy cross-gradients without violating the differential privacy guarantee or destroying the convergence properties that produce linear speedup under arbitrary non-IID partitions.
What would settle it
An experiment that either measures a privacy leakage exceeding the bound implied by the added noise after calibration or records sub-linear speedup on non-IID data partitions would falsify the central claims.
If this is right
- The minimum noise level required to reach a chosen privacy level is identified by the analysis.
- Linear speedup in training time is retained despite non-IID data partitions.
- The calibrated updates defend against privacy attacks while still producing accurate models.
- The approach works in fully decentralized settings without any central server.
Where Pith is reading between the lines
- The same calibration step could be examined for use with other noise mechanisms or aggregation rules.
- The momentum-style update suggests the method may combine naturally with existing optimizers that already employ momentum.
- Testing on a wider range of network topologies would reveal whether the linear speedup holds beyond the topologies studied.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DPDL, a decentralized stochastic learning algorithm for non-IID data that perturbs cross-gradients with Gaussian noise to enforce (ε,δ)-DP and then applies cosine-similarity calibration before momentum-style aggregation; it claims a rigorous derivation of the minimum noise level needed for a target privacy budget together with a proof that linear speedup is retained under arbitrary non-IID partitions, supported by experiments on real-world datasets.
Significance. If the privacy and convergence claims hold after the calibration step, the work would supply a concrete, analyzable mechanism for trading privacy against convergence rate in fully decentralized non-IID settings, an area where existing DP-SGD analyses do not directly apply.
major comments (2)
- [§4] §4 (Theoretical Analysis): the derivation of the minimum noise variance σ^{2} for (ε,δ)-DP must be shown to remain valid after the deterministic cosine-similarity rescaling of the noisy cross-gradients; standard DP composition theorems do not automatically cover a subsequent data-dependent post-processing step whose output is fed into the momentum update.
- [§4] §4, Theorem on linear speedup: the convergence bound assumes that the calibrated updates remain unbiased (or have controlled bias) under arbitrary non-IID partitions; the cosine rescaling can introduce distribution-dependent scaling factors that violate this assumption, and the proof must explicitly bound the resulting bias term.
minor comments (2)
- [Experiments] The experimental section should report the exact baselines, privacy budgets, and statistical significance tests used to claim superiority over prior decentralized DP methods.
- [Algorithm 1] Notation for the calibrated gradient ̂g and the momentum coefficient should be introduced before the first use in the algorithm description.
Simulated Author's Rebuttal
We thank the referee for the constructive comments on the theoretical analysis in Section 4. We address each major comment below and indicate the corresponding revisions.
read point-by-point responses
-
Referee: [§4] §4 (Theoretical Analysis): the derivation of the minimum noise variance σ^{2} for (ε,δ)-DP must be shown to remain valid after the deterministic cosine-similarity rescaling of the noisy cross-gradients; standard DP composition theorems do not automatically cover a subsequent data-dependent post-processing step whose output is fed into the momentum update.
Authors: We appreciate this clarification. The Gaussian perturbation mechanism satisfies (ε,δ)-DP by construction. The subsequent cosine-similarity calibration is a deterministic post-processing function applied to the mechanism's output. By the post-processing property of differential privacy, any function of an (ε,δ)-DP output (data-dependent or otherwise) remains (ε,δ)-DP. Consequently, the minimum noise variance derived for the perturbation step continues to guarantee the target privacy budget for the calibrated cross-gradients. We will add an explicit remark invoking the post-processing theorem in Section 4 to make this connection clear. revision: yes
-
Referee: [§4] §4, Theorem on linear speedup: the convergence bound assumes that the calibrated updates remain unbiased (or have controlled bias) under arbitrary non-IID partitions; the cosine rescaling can introduce distribution-dependent scaling factors that violate this assumption, and the proof must explicitly bound the resulting bias term.
Authors: We agree that the cosine rescaling introduces a data-dependent scaling factor whose effect on bias must be controlled. In the existing proof we bound the deviation of the calibrated direction from the true cross-gradient using the fact that cosine similarity lies in [0,1] and the bounded gradient assumption; however, an explicit bias term arising from the scaling was not isolated. We will revise the proof of the linear-speedup theorem to introduce and bound this additional bias term, showing that it remains controlled under arbitrary non-IID partitions and does not alter the linear speedup rate. revision: yes
Circularity Check
No circularity detected; theoretical claims presented as independent analysis
full rationale
The provided abstract and context contain no equations, self-citations, or fitted parameters that reduce the claimed minimum noise level or linear speedup result to inputs by construction. The DP noise derivation and convergence bound are described as outputs of rigorous analysis rather than tautological renamings or post-hoc fits. The cosine calibration step is introduced as a deterministic post-processing technique whose interaction with DP and convergence is asserted to hold, but no load-bearing self-citation chain or self-definitional loop is visible. This is the normal case of a self-contained theoretical claim.
Axiom & Free-Parameter Ledger
read the original abstract
In the paradigm of decentralized learning, a group of agents collaborate to train a global model using distributed datasets without a central server. Although the power of collaboration has been verified by many state-of-the-art studies, it entails extensive gradient information exchanging among the agents and thus induces high risk of privacy leakage for the individual agents. Moreover, in real-world applications, the training data are usually non-identically and independently distributed across the agents, inducing more challenges to enable privacy-preserved decentralized learning. To address these issues, we propose a privacy-preserved decentralized learning algorithm with non-IID data, DPDL, which leverages the notion of Differential Privacy (DP) in cross-gradient aggregation through a similarity-based calibration technique. Specifically, in each round, each agent perturbs the cross-gradients (i.e., the derivatives of its neighbors' local model in its private local data) by Gaussian noise mechanism before sharing them with its neighbors; it then adopt cosine similarity to calibrate the received perturbed cross-gradients such that the aggregation of the calibrated cross-gradients can be utilized to effectively update local model in a momentum-like manner. Our rigorous theoretical analysis not only reveals the minimum noise level required to achieve a specific level of privacy preservation, but also illustrates that our algorithm still achieves a linear speedup in training with non-IID data. We finally conduct extensive experiments on real-world dataset to validate the effectiveness of our algorithm in defending privacy attacks and in training accurate models.
Figures
Reference graph
Works this paper leans on
-
[1]
A Survey on Distributed Machine Learning,
J. Verbraeken, M. Wolting, J. Katzy, J. Kloppenburg, T. Verbelen, and J. Rellermeyer, “A Survey on Distributed Machine Learning,”ACM Computing Surveys, vol. 53, no. 2, pp. 1–33, 2020. 14 (a1) Ground Truth in the 200th round (a2) Ground Truth in the 400th round (a3) Ground Truth in the 800th round (b1) ND-CG in the 200th round (b2) ND-CG in the 400th round...
2020
-
[2]
A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection,
Q. Li, Z. Wen, Z. Wu, S. Hu, N. Wang, Y . Li, X. Liu, and B. He, “A Survey on Federated Learning Systems: Vision, Hype and Reality for Data Privacy and Protection,”IEEE Trans. on Knowledge and Data Engineering, vol. 35, no. 4, pp. 3347–3366, 2023
2023
-
[3]
Decentralized Federated Learning: Funda- mentals, State of the Art, Frameworks, Trends, and Challenges,
E. Beltr ´an, M. P ´erez, P. S ´anchez, S. Bernal, G. Bovet, M. P ´erez, G. P ´erez, and A. Celdr ´an, “Decentralized Federated Learning: Funda- mentals, State of the Art, Frameworks, Trends, and Challenges,”IEEE Communications Surveys & Tutorials, vol. 25, no. 4, pp. 2983–3013, 15 TABLE III THE RESULTS OF GRADIENT INVERSION ATTACKS ONMNISTDATASET. Metric...
2023
-
[4]
FedQV: Leveraging Quadratic V oting in Federated Learning,
T. Chu and N. Laoutaris, “FedQV: Leveraging Quadratic V oting in Federated Learning,”Proceedings of the ACM on Measurement and Analysis of Computing Systems, vol. 8, no. 2, pp. 22:1–22:36, 2024
2024
-
[5]
Communication-Efficient Learning of Deep Networks from Decentral- ized Data,
B. McMahan, E. Moore, D.Ramage, S. Hampson, and B. Arcas, “Communication-Efficient Learning of Deep Networks from Decentral- ized Data,” inProc. of the 20th AISTATS, vol. 54, 2017, pp. 1273–1282
2017
-
[6]
Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,
X. Lian, C. Zhang, H. Zhang, C. Hsieh, W. Zhang, and J. Liu, “Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent,” inProc. of the 31st NIPS, 2017, p. 5336–5346
2017
-
[7]
Network Topology and Communication-computation Tradeoffs in Decentralized Optimization,
A. Nedi, A. Olshevsky, and M. Rabbat, “Network Topology and Communication-computation Tradeoffs in Decentralized Optimization,” Proceedings of the IEEE, vol. 106, no. 5, pp. 953–976, 2018
2018
-
[8]
Asynchronous Decentralized Parallel Stochastic Gradient Descent,
X. Lian, W. Zhang, C. Zhang, and J. Liu, “Asynchronous Decentralized Parallel Stochastic Gradient Descent,” inProc. of the 35th ICML, 2018, pp. 3043–3052
2018
-
[9]
Quasi-global Mo- mentum: Accelerating Decentralized Deep Learning on Heterogeneous Data,
T. Lin, S. Karimireddy, S. Stich, and M. Jaggi, “Quasi-global Mo- mentum: Accelerating Decentralized Deep Learning on Heterogeneous Data,” inProc. of the 38th ICML, 2021, pp. 6654–6665
2021
-
[10]
Federated Learning on Non-iid Data Silos: An Experimental Study,
Q. Li, Y . Diao, Q. Chen, and B. He, “Federated Learning on Non-iid Data Silos: An Experimental Study,” inProc. of the 38th IEEE ICDE, 2022, pp. 965–978
2022
-
[11]
Decentralized Federated Learning via Mutual Knowledge Transfer,
C. Li, G. Li, and P. Varshney, “Decentralized Federated Learning via Mutual Knowledge Transfer,”IEEE Internet of Things Journal, vol. 9, no. 2, pp. 1136–1147, 2021
2021
-
[12]
Im- proving the Model Consistency of Decentralized Federated Learning,
Y . Shi, L. Shen, K. Wei, Y . Sun, B. Yuan, X. Wang, and D. Tao, “Im- proving the Model Consistency of Decentralized Federated Learning,” inProc. of the 40th ICML, vol. 202, 2023, pp. 31 269–31 291
2023
-
[13]
Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous Data,
S. Aketi, A. Hashemi, and K. Roy, “Global Update Tracking: A Decentralized Learning Algorithm for Heterogeneous Data,” inProc. of the 37th NeurIPS, 2024
2024
-
[14]
NET- FLEET: Achieving Linear Convergence Speedup for Fully Decentralized Federated Learning with Heterogeneous Data,
X. Zhang, M. Fang, Z. Liu, H. Yang, J. Liu, and Z. Zhu, “NET- FLEET: Achieving Linear Convergence Speedup for Fully Decentralized Federated Learning with Heterogeneous Data,” inProc. of the 23rd ACM MobiHoc, 2022, p. 71–80
2022
-
[15]
Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data,
Y . Esfandiari, S. Tan, Z. Jiang, A. Balu, E. Herron, C. Hegde, and S. Sarkar, “Cross-Gradient Aggregation for Decentralized Learning from Non-IID Data,” inProc. of the 38th ICML, 2021, pp. 3036–3046
2021
-
[16]
On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization,
H. Yu, R. Jin, and S. Yang, “On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization,” inProc. of the 36th ICML, 2019, pp. 7184–7193
2019
-
[17]
Deep leakage from gradients,
L. Zhu, Z. Liu, and S. Han, “Deep leakage from gradients,” inProc. of the 33rd NIPS, 2019, pp. 14 747–14 756
2019
-
[18]
iDLG: Improved Deep Leakage from Gradients
B. Zhao, K. Mopuri, and H. Bilen, “iDLG: Improved Deep Leakage from Gradients,”arXiv preprint arXiv:2001.02610, 2020
work page Pith review arXiv 2001
-
[19]
More than Enough is Too Much: Adaptive Defenses against Gradient Leakage in Production Federated Learning,
F. Wang, E. Hugh, and B. Li, “More than Enough is Too Much: Adaptive Defenses against Gradient Leakage in Production Federated Learning,” inProc. of the 42nd INFOCOM, 2023, pp. 1–10
2023
-
[20]
Differential Privacy,
C. Dwork, “Differential Privacy,” inProc. of the 33rd ICALP, 2006, pp. 1–12
2006
-
[21]
The Algorithmic Foundations of Differential Privacy,
C. Dwork and A. Roth, “The Algorithmic Foundations of Differential Privacy,”Foundations and Trends in Theoretical Computer Science, vol. 9, pp. 211–407, 2014
2014
-
[22]
How to dp-fy ml: A practical tutorial to machine learning with differential privacy,
N. Ponomareva, S. Vassilvitskii, Z. Xu, B. McMahan, A. Kurakin, and C. Zhang, “How to dp-fy ml: A practical tutorial to machine learning with differential privacy,” inProc. of the 29th ACM SIGKDD, 2023, p. 5823–5824
2023
-
[23]
Privacy-Preserving Deep Learning,
R. Shokri and V . Shmatikov, “Privacy-Preserving Deep Learning,” in Proc. of the 22nd ACM CCS, 2015, pp. 1310–1321
2015
-
[24]
Heterogeneous Differential-Private Federated Learning: Trading Privacy for Utility Truthfully,
X. Lin, J. Wu, J. Li, C. Sang, S. Hu, and M. J. Deen, “Heterogeneous Differential-Private Federated Learning: Trading Privacy for Utility Truthfully,”IEEE Trans. on Dependable and Secure Computing, vol. 20, no. 6, pp. 5113–5129, 2023
2023
-
[25]
Differentially Private Federated Learning on Non-iid Data: Convergence Analysis and Adap- tive Optimization,
L. Chen, X. Ding, Z. Bao, P. Zhou, and H. Jin, “Differentially Private Federated Learning on Non-iid Data: Convergence Analysis and Adap- tive Optimization,”IEEE Trans. on Knowledge and Data Engineering, vol. 36, no. 9, pp. 4567–4581, 2024. 16
2024
-
[26]
A(DP)2SGD: Asynchronous Decen- tralized Parallel Stochastic Gradient Descent With Differential Privacy,
J. Xu, W. Zhang, and F. Wang, “A(DP)2SGD: Asynchronous Decen- tralized Parallel Stochastic Gradient Descent With Differential Privacy,” IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, pp. 8036–8047, 2021
2021
-
[27]
Muffliato: Peer-to- Peer Privacy Amplification for Decentralized Optimization and Averag- ing,
E. Cyffers, M. Even, A. Bellet, and L. Massouli ´e, “Muffliato: Peer-to- Peer Privacy Amplification for Decentralized Optimization and Averag- ing,” inProc. of 36th NIPS, vol. 35, 2022, pp. 15 889–15 902
2022
-
[28]
Between Privacy and Utility: On Differential Privacy in Theory and Practice,
J. Seeman and D. Susser, “Between Privacy and Utility: On Differential Privacy in Theory and Practice,”ACM Journal on Responsible Comput- ing, vol. 1, no. 1, pp. 1–18, 2024
2024
-
[29]
Randomized Gossip Algorithms,
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah, “Randomized Gossip Algorithms,”IEEE Trans. on Information Theory, vol. 52, no. 6, pp. 2508–2530, 2006
2006
-
[30]
Optimal Algorithms for Smooth and Strongly Convex Distributed Optimization in Networks,
K. Scaman, F. Bach, S. Bubeck, Y . Lee, and L. Massouli ´e, “Optimal Algorithms for Smooth and Strongly Convex Distributed Optimization in Networks,” inProc. of the 34th ICML, 2017, pp. 3027–3036
2017
-
[31]
Optimal Algorithms for Non-smooth Distributed Optimization in Networks,
K. Scaman, F. Bach, S. Bubeck, Y . Lee, and L. Massouli ´e, “Optimal Algorithms for Non-smooth Distributed Optimization in Networks,” in Proc. of the 32nd NIPS, 2018, p. 2745–2754
2018
-
[32]
Distributed Stochastic Gradient Tracking Methods,
S. Pu and A. Nedi, “Distributed Stochastic Gradient Tracking Methods,” Mathematical Programming, vol. 187, no. 1, pp. 409–457, 2021
2021
-
[33]
Momentum Tracking: Momentum Acceleration for Decentralized Deep Learning on Heterogeneous Data,
Y . Takezawa, H. Bao, K. Niwa, R. Sato, and M. Yamada, “Momentum Tracking: Momentum Acceleration for Decentralized Deep Learning on Heterogeneous Data,”Trans. on Machine Learning Research, vol. 2023, 2023
2023
-
[34]
Neighborhood Gradient Mean: An Ef- ficient Decentralized Learning Method for Non-IID Data Distributions,
S. Aketi, S. Kodge, and K. Roy, “Neighborhood Gradient Mean: An Ef- ficient Decentralized Learning Method for Non-IID Data Distributions,” Trans. on Machine Learning Research, 2023
2023
-
[35]
Z. Huang and Y . Gong, “Differentially Private ADMM for Convex Dis- tributed Learning: Improved Accuracy via Multi-step Approximation,” arXiv preprint arXiv:2005.07890, 2020
-
[36]
Pdsl: Privacy-preserved decentralized stochastic learning with heterogeneous data distribution,
L. Wang, Y . Yuan, C. Wang, and F. Li, “Pdsl: Privacy-preserved decentralized stochastic learning with heterogeneous data distribution,” in2025 IEEE 45th ICDCS. IEEE Computer Society, 2025, pp. 736– 746
2025
-
[37]
Understanding Clip- ping for Federated Learning: Convergence and Client-Level Differential Privacy,
X. Zhang, X. Chen, M. Hong, S. Wu, and J. Yi, “Understanding Clip- ping for Federated Learning: Convergence and Client-Level Differential Privacy,” inProc. of the 39th ICML, vol. 162, 2022, pp. 26 048–26 067
2022
-
[38]
What Can We Learn Privately?
S. Kasiviswanathan, H. Lee, K. Nissim, S. Raskhodnikova, and A. Smith, “What Can We Learn Privately?” inProc. of the 49th IEEE FOCS, 2008, p. 531–540
2008
-
[39]
Deep Learning with Differential Privacy,
M. Abadi, A. Chu, I. Goodfellow, H. McMahan, I. Mironov, K. Talwar, and L. Zhang, “Deep Learning with Differential Privacy,” inProc. of the 2016 ACM CCS, 2016, p. 308–318
2016
-
[40]
Inverting gradients - how easy is it to break privacy in federated learning?
J. Geiping, H. Bauermeister, H. Dr ¨oge, and M. Moeller, “Inverting gradients - how easy is it to break privacy in federated learning?” in Proceedings of the 34th NIPS, 2020
2020
-
[41]
Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures,
M. Fredrikson, S. Jha, and T. Ristenpart, “Model Inversion Attacks that Exploit Confidence Information and Basic Countermeasures,” inProc. of the 22nd ACM CCS, 2015, p. 1322–1333
2015
-
[42]
On the momentum term in gradient descent learning algo- rithms,
N. Qian, “On the momentum term in gradient descent learning algo- rithms,”Neural networks, vol. 12, no. 1, pp. 145–151, 1999
1999
-
[43]
Gradient-based learning applied to document recognition,
Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”Proc. of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998
1998
-
[44]
Learning multiple layers of features from tiny images,
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009
2009
-
[45]
Very Deep Convolutional Networks for Large-Scale Image Recognition
K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,”arXiv preprint arXiv:1409.1556, 2014
work page internal anchor Pith review Pith/arXiv arXiv 2014
-
[46]
Data-Free Knowledge Distillation for Heterogeneous Federated Learning,
Z. Zhu, J. Hong, and J. Zhou, “Data-Free Knowledge Distillation for Heterogeneous Federated Learning,” inProc. of the 38th ICML, vol. 139, 2021, pp. 12 878–12 889
2021
-
[47]
Federated Learning With Differential Privacy: Algorithms and Performance Analysis,
K. Wei, J. Li, M. Ding, C. Ma, H. Yang, F. Farokhi, S. Jin, T. Quek, and H. V . Poor, “Federated Learning With Differential Privacy: Algorithms and Performance Analysis,”IEEE Trans. on Information Forensics and Security, vol. 15, pp. 3454–3469, 2020
2020
-
[48]
Decentralized Wireless Federated Learning With Differential Privacy,
S. Chen, D. Yu, Y . Zou, J. Yu, and X. Cheng, “Decentralized Wireless Federated Learning With Differential Privacy,”IEEE Trans. on Industrial Informatics, vol. 18, no. 9, pp. 6273–6282, 2022
2022
-
[49]
Decen- tralized parallel sgd with privacy preservation in vehicular networks,
D. Yu, Z. Zou, S. Chen, Y . Tao, B. Tian, W. Lv, and X. Cheng, “Decen- tralized parallel sgd with privacy preservation in vehicular networks,” IEEE Transactions on Vehicular Technology, vol. 70, no. 6, pp. 5211– 5220, 2021
2021
-
[50]
Multiscale Structural Similarity for Image Quality Assessment,
Z. Wang, E. Simoncelli, and A. Bovik, “Multiscale Structural Similarity for Image Quality Assessment,” inProc. of the 37th ACSSC, vol. 2, 2003, pp. 1398–1402. 17 APPENDIXA PROOF OFTHEOREM1 In this section, we prove the efficacy of our DPDL algorithm in preserving differential privacy. In the decentralized learning system, each agent collaborates with its ...
2003
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.