REVIEW 3 major objections 4 minor 68 references
The paper claims that replacing the Gaussian mechanism in DP-SGD with a two-parameter randomized-scale Laplace noise can improve model accuracy in the high-privacy regime, reporting gains of roughly 10 to 40 accuracy points at epsilon aroun
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 23:54 UTC pith:HGRWZJFK
load-bearing objection The mechanism is clever and the experiments are impressive, but the multivariate privacy bound in Theorem 4.9 is not valid, so the headline epsilon values are unsupported. the 3 major comments →
PLRV-O: Advancing Differentially Private Deep Learning via Privacy Loss Random Variable Optimization
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
PLRV-O claims that the privacy loss random variable, not the noise distribution itself, is the right target for mechanism design in DP-SGD. It introduces a search space of randomized-scale Laplace mechanisms, where the inverse scale u=1/b follows a gamma distribution with shape k and scale theta. For this family, Theorem 4.9 bounds the subsampled multivariate moments accountant by a sum of per-coordinate terms evaluated on the majorization set x_i = C(sqrt(i)-sqrt(i-1)), with each term depending on the moment generating function of u; because that MGF is (1 - t theta)^{-k}, the account is closed-form and can support thousands of moment orders. This turns noise design into a constrained optim
What carries the argument
Two interlocking pieces. The first is the PLRV noise itself: a multivariate Laplace mechanism whose scale b is a random draw from a seed density, so the noise density is a mixture of Laplace densities. The second is a privacy accountant for it: after l2 clipping, the coordinate-wise subsampled moments are Schur-convex, and the majorization set x_i = C(sqrt(i)-sqrt(i-1)) dominates the clipped gradient magnitudes in weak majorization order; therefore the total moments bound is the sum of the univariate bounds, with each univariate bound expressed through the MGF M_u(t)=(1-t theta)^{-k} for gamma seeds. The separation of k (which mostly controls entropy and utility) and theta (which mostly cont
Load-bearing premise
The whole privacy accounting assumes that the privacy loss of one training step can be bounded by the sum of per-coordinate privacy-loss bounds, which requires the subsampling events and the random noise scale to be independent across coordinates; in the actual algorithm they are shared, so correlated coordinate losses could make the true loss exceed the bound.
What would settle it
Run one iteration of the actual PLRV-O algorithm (shared batch mask and a single sampled scale b for all coordinates) on a small model, compute the full multivariate privacy-loss distribution numerically by exact or PLD-style accounting, and compare the resulting epsilon with the sum-of-coordinate bound; if the direct epsilon is larger than the reported bound, the privacy guarantee is understated.
If this is right
- At strictly small epsilon, PLRV-O reports accuracy gains over Gaussian DP-SGD large enough to change deployment choices: e.g., 94.03% vs 83.93% on CIFAR-10 ViT at epsilon about 0.5, and 92.20% vs 50.25% on SST-2 RoBERTa-large at epsilon about 0.2.
- Because the accountant runs per coordinate and sums via a majorization set, the method avoids the sqrt(n) noise inflation that made Laplace noise impractical for deep nets, so it applies to models with tens of millions of parameters.
- The same noise mechanism can be dropped into other DP algorithms: the paper reports consistent gains when PLRV-O noise replaces Gaussian noise in DP-FTRL on MNIST and CIFAR-10.
- The framework supports tailoring noise to task properties such as model size, number of steps, batch sampling rate, and clipping threshold; the reported runs also show faster convergence at the same epsilon.
- Any seed distribution with a known moment generating function can be used in the same accounting, so the gamma family is an instance rather than an endpoint.
Where Pith is reading between the lines
- The central reported epsilons inherit a risk: because one scale b is sampled per update and shared by all parameters, the per-coordinate independence assumption in the moments accountant may understate true privacy loss; a numerical PLD audit of the exact multivariate mechanism would settle this.
- If the multivariate bound fails, the strong-accuracy results would need to be re-read at their true (larger) epsilon; the Gaussian comparisons might still favor PLRV-O, but the margin would shrink.
- The optimization objective C(k-1)theta is a signal-to-noise surrogate, not an accuracy guarantee; the paper does not prove that maximizing this objective maximizes test accuracy, so the reported gains are empirical rather than optimality-based.
- A natural next step is to test the same randomized-scale construction with other distortion measures beyond l1 error, such as l2 or per-coordinate heterogeneous costs, since the accountant only needs the MGF of the reciprocal scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PLRV-O, a framework for DP-SGD in which the noise distribution is a multivariate Laplace mechanism whose scale parameter is itself randomized (Definition 4.6). The scale is drawn from a Gamma distribution, giving two tunable parameters (k, theta) in addition to the clipping threshold C. The authors derive a moments accountant for this mechanism (Theorem 4.9), apply majorization to support ℓ2 clipping, and then optimize the parameters with a constrained nonlinear solver. The empirical sections report large utility improvements at very small ε: e.g., 94.03% CIFAR-10 ViT accuracy at ε≈0.5 and 92.20% SST-2 RoBERTa-large accuracy at ε≈0.2, compared to Gaussian baselines. The central claim is that PLRV-O achieves these results under valid (ε, δ)-DP guarantees with a tight accountant.
Significance. If the privacy accounting were sound, this would be a significant contribution: it proposes a genuinely non-Gaussian DP-SGD noise family, gives a closed-form moments accountant, and demonstrates substantial utility gains at strong privacy levels across both vision and language tasks. The paper includes source code and extensive experiments, and the attempt to optimize the privacy loss random variable itself is an interesting angle. However, the central privacy bound is not a valid upper bound on the true joint moments accountant. Since the headline ε values and all comparisons depend on that bound, the main scientific claim is currently unsupported.
major comments (3)
- [Theorem 4.5, Eq. (16)] Eq. (19) bounds the multivariate PLRV mechanism's moments accountant by the sum of per-coordinate univariate moment bounds. This is not a valid upper bound, because the PLRV mechanism in Definitions 4.6–4.7 draws a single scale b for the whole vector, and the DP-SGD subsampling event is shared across all coordinates. For one iteration, after the binomial expansion the true subsampled moment contains a term of the form E_b[∏_i F_i(b)] (where the product is over coordinates and F_i depends on the coordinate's majorized sensitivity x_i and the same random b), while the proof uses ∏_i E_b[F_i(b)]. The functions F_i(b) are comonotone (decreasing) in b, so E_b[∏_i F_i(b)] ≥ ∏_i E_b[F_i(b)]. Thus Eq. (19) underestimates the privacy loss. The 'composition over coordinates' argument in §4.1 would require releasing coordinates with independent randomness; it does not apply to a one-shot vector-val
- The same flaw already appears in Theorem 4.5 for the standard multivariate Laplace mechanism: Eq. (15) sets the total MAF to Σ_i α_{g_i}(λ), and Eq. (16) bounds it by summing per-coordinate univariate subsampled Laplace bounds. Conditional on the scale b, coordinates are independent, but the batch-inclusion event is common to all coordinates, so the mixture induced by subsampling does not factor across coordinates. Even for fixed b, the joint moment is log E_{(I,z)}[(1−ζ+ζ r(z))^{λ+1}] with r(z) = exp(Σ_i (|z_i| − |z_i−g_i|)/b), not the sum of per-coordinate log-moments. Consequently the resulting bound is not an upper bound. Since Theorem 4.9 inherits this structure from Theorem 4.5, the central privacy guarantee of the paper is unsupported.
- [§6.4] The privacy audit reports only empirical lower bounds on ε under specific attacks (ClipBKD), with a finite number of trials. Such an audit cannot certify the claimed (ε, δ)-DP guarantees. The validation of the headline privacy numbers must come from the moments accountant, which is exactly the step that fails. The audit therefore does not mitigate the flaw in Theorem 4.9.
minor comments (4)
- [§1] The phrase 'Mironov et al. [42] and Sander et al. [42, 51]' appears to have a citation error: reference [42] is listed twice. Please correct.
- [Appendix B.2] The proof refers to 'Proof B.2 (Theorem 3.2)', but there is no Theorem 3.2 in the paper; the intended reference is likely Theorem 4.2. Similar internal cross-reference issues appear elsewhere.
- [Algorithm 3] The algorithm header includes 'φ2' but the algorithm signature uses 'φ1'; this appears to be an artifact. Also, line 3 computes T = ⌈E/q⌉ while the text earlier uses T = ⌈E·N/B⌉; please reconcile the notation.
- [Figures 10–11] The y-axis in the audit plots is labeled 'Estimate' but the caption says 'empirical ε'. Clarify whether these are lower bounds on ε and how the confidence level 0.01 maps to the displayed values.
Circularity Check
No significant circularity: PLRV-O's privacy bound is derived from the mechanism, and the utility numbers are empirical outcomes, not refits of the bound.
full rationale
The paper's central derivation chain is: Definitions 4.6/4.7 define PLRV noise/mechanism; Theorem 4.8 derives a univariate subsampled PLRV moments accountant by applying the law of total expectation to the fixed-scale Laplace bound (Theorem 4.2) and using the MGF of the reciprocal scale; Theorem 4.9 extends this to n coordinates via a Schur-convexity/majorization argument (Theorem 4.3, Lemma 4.4). None of these steps defines the conclusion in terms of the inputs: the MAF bound is an analytic upper bound in (k, theta, C, zeta, T), and the reported epsilon values come from plugging chosen parameters into that bound, not from fitting epsilon to the observed accuracies. The optimization (Algorithm 1) maximizes a surrogate J = C(k-1)theta and imposes heuristic constraints (Cmin, rho=2, distortion cap 10); these are tuning heuristics and do not make the privacy guarantee a renamed fit. The citations to the authors' prior work (e.g., [43], [21], [22]) are contextual or for orthogonal applications, not load-bearing for the privacy proof; the load-bearing references (Mironov et al. [42], majorization [41,56]) are external. Footnote 6 flags that the multivariate bound sums per-coordinate accountants 'similar to how composability traditionally applies over iterations'; this is a potential validity gap (the shared scale b and shared subsampling event make coordinate privacy losses positively correlated, so the sum may under-estimate the true joint MAF), but a correctness gap is not a circularity. The privacy audit in Sec. 6.4 gives empirical lower bounds and is not used to calibrate the theoretical epsilon. Hence no step in the paper's derivation reduces to its own inputs by construction.
Axiom & Free-Parameter Ledger
free parameters (6)
- k (Gamma shape) =
Examples: 141.06, 5242.4, 60000, 10000
- theta (Gamma scale) =
Examples: 8.32e-4, 2.08e-5, 1e-5
- C (clipping threshold) =
Examples: 0.1, 0.3, 0.5, 5.0, 10.0, 15.0
- rho (clip-span ratio) =
approximately 2
- Distortion cap =
10
- lambda_max (max moment order) =
Large, e.g., 10^3
axioms (5)
- standard math Standard moments accountant and Rényi DP composition theorems (Mironov et al.)
- domain assumption Poisson subsampling with rate zeta
- standard math Post-processing: DP guarantee of a mechanism that outputs (b,z) transfers to marginal z
- standard math Schur-convexity of the univariate Laplace MAF and the majorization set bound (Lemma 4.4)
- ad hoc to paper Heuristic constraints c0-c4 in Algorithm 1
Cite this review
Pith. "Pith review of PLRV-O: Advancing Differentially Private Deep Learning via Privacy Loss Random Variable Optimization." pith.science (2026). https://pith.science/paper/HGRWZJFK
@misc{pith2026250906264,
author = {Pith},
title = {Pith review of: PLRV-O: Advancing Differentially Private Deep Learning via Privacy Loss Random Variable Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/HGRWZJFK}},
note = {Machine review of arXiv:2509.06264}
}
read the original abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is a standard method for enforcing privacy in deep learning, typically using the Gaussian mechanism to perturb gradient updates. However, conventional mechanisms such as Gaussian and Laplacian noise are parameterized only by variance or scale. This single degree of freedom ties the magnitude of noise directly to both privacy loss and utility degradation, preventing independent control of these two factors. The problem becomes more pronounced when the number of composition rounds T and batch size B vary across tasks, as these variations induce task-dependent shifts in the privacy-utility trade-off, where small changes in noise parameters can disproportionately affect model accuracy. To address this limitation, we introduce PLRV-O, a framework that defines a broad search space of parameterized DP-SGD noise distributions, where privacy loss moments are tightly characterized yet can be optimized more independently with respect to utility loss. This formulation enables systematic adaptation of noise to task-specific requirements, including (i) model size, (ii) training duration, (iii) batch sampling strategies, and (iv) clipping thresholds under both training and fine-tuning settings. Empirical results demonstrate that PLRV-O substantially improves utility under strict privacy constraints. On CIFAR-10, a fine-tuned ViT achieves 94.03% accuracy at epsilon approximately 0.5, compared to 83.93% with Gaussian noise. On SST-2, RoBERTa-large reaches 92.20% accuracy at epsilon approximately 0.2, versus 50.25% with Gaussian.
Figures
Reference graph
Works this paper leans on
-
[1]
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016. Deep learning with differential privacy. In CCS. 308–318
work page 2016
-
[2]
Borja Balle, Gilles Barthe, Marco Gaboardi, Justin Hsu, and Tetsuya Sato. 2020. Hypothesis Testing Interpretations and Rényi Differential Privacy. InAISTATS
work page 2020
-
[3]
Borja Balle and Yu-Xiang Wang. 2018. Improving the Gaussian mechanism for differential privacy: Analytical calibration and optimal denoising. InICML
work page 2018
-
[4]
Jeremy Bernstein, Jiawei Zhao, Kamyar Azizzadenesheli, and Anima Anandkumar
-
[5]
Paul T Boggs and Jon W Tolle. 1995. Sequential quadratic programming.Acta numerica4 (1995), 1–51
work page 1995
-
[6]
Zhiqi Bu, Sivakanth Gopi, Janardhan Kulkarni, Yin Tat Lee, Hanwen Shen, and Uthaipon Tantipongpipat. 2021. Fast and Memory Efficient Eifferentially Private- SGD via JL projections.NeurIPS34 (2021), 19680–19691
work page 2021
-
[7]
Zhiqi Bu, Jialin Mao, and Shiyun Xu. 2022. Scalable and efficient training of large convolutional neural networks with differential privacy.NeurIPS(2022)
work page 2022
-
[8]
Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, and George Karypis. 2022. Differen- tially private bias-term only fine-tuning of foundation models. InNeurIPS 2022 Workshop on Trustworthy and Socially Responsible Machine Learning (TSRML)
work page 2022
-
[9]
Mark Bun, Cynthia Dwork, Guy N. Rothblum, and Thomas Steinke. 2018. Com- posable and Versatile Privacy via Truncated CDP. InSTOC. 74–86
work page 2018
-
[10]
Richard H Byrd, Jean Charles Gilbert, and Jorge Nocedal. 2000. A trust region method based on interior point techniques for nonlinear programming.Mathe- matical programming89, 1 (2000), 149–185
work page 2000
-
[11]
Richard H Byrd, Mary E Hribar, and Jorge Nocedal. 1999. An interior point algorithm for large-scale nonlinear programming.SIAM Journal on Optimization 9, 4 (1999), 877–900
work page 1999
-
[12]
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jeremy Kos, and Dawn Song. 2019. The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks. InUSENIX Security. 267–284
work page 2019
-
[13]
Giovanni Cherubin, Boris Köpf, Andrew Paverd, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin. 2024. Closed-Form Bounds for DP-SGD against Record-level Inference Attacks. InUSENIX Security. 4819–4836
work page 2024
-
[14]
Erik Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen, Matthias Bauer, and Philipp Hennig. 2021. Laplace redux-effortless bayesian deep learning.NeurIPS34 (2021), 20089–20103
work page 2021
-
[15]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT. 4171–4186
work page 2019
-
[16]
Minxin Du, Xiang Yue, Sherman SM Chow, Tianhao Wang, Chenyu Huang, and Huan Sun. 2023. Dp-forward: Fine-tuning and inference on language models with differential privacy in forward pass. InCCS. 2665–2679
work page 2023
-
[17]
Cynthia Dwork. 2006. Differential Privacy. InICALP. 1–12
work page 2006
-
[18]
Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor. 2006. Our data, ourselves: Privacy via distributed noise generation. InAdvances in Cryptology—EUROCRYPT. 486–503
work page 2006
-
[19]
Cynthia Dwork, Aaron Roth, et al. 2014. The algorithmic foundations of differ- ential privacy.Foundations and Trends®in Theoretical Computer Science(2014)
work page 2014
-
[20]
Cynthia Dwork, Guy N Rothblum, and Salil Vadhan. 2010. Boosting and differen- tial privacy. InFOCS. 51–60
work page 2010
-
[21]
Shuya Feng, Meisam Mohammady, Hanbin Hong, Shenao Yan, Ashish Kundu, Binghui Wang, and Yuan Hong. 2025. Harmonizing Differential Privacy Mecha- nisms for Federated Learning: Boosting Accuracy and Convergence. InCODASPY
work page 2025
-
[22]
Shuya Feng, Meisam Mohammady, Han Wang, Xiaochen Li, Zhan Qin, and Yuan Hong. 2024. DPI: Ensuring Strict Differential Privacy for Infinite Data Streaming. InIEEE Symposium on Security and Privacy. 1009–1027
work page 2024
-
[23]
Jie Fu, Qingqing Ye, Haibo Hu, Zhili Chen, Lulu Wang, Kuncan Wang, and Xun Ran. 2024. DPSUR: accelerating differentially private stochastic gradient descent using selective update and release.VLDB(2024)
work page 2024
-
[24]
Jonas Geiping, Hartmut Bauermeister, Hannah Dröge, and Michael Moeller. 2020. Inverting gradients - how easy is it to break privacy in federated learning?. In NeurIPS
work page 2020
-
[25]
Quan Geng and Pramod Viswanath. 2014. The optimal mechanism in differential privacy. InISIT. 2371–2375
work page 2014
-
[26]
Quan Geng and Pramod Viswanath. 2015. The optimal noise-adding mechanism in differential privacy.IEEE Transactions on Information Theory(2015)
work page 2015
-
[27]
Sivakanth Gopi, Yin Tat Lee, and Lukas Wutschitz. 2021. Numerical composition of differential privacy.NeurIPS34 (2021), 11631–11642
work page 2021
-
[28]
Godfrey H Hardy. 1929. Some simple inequalities satisfied by convex functions. Messenger Math.58 (1929), 145–152
work page 1929
-
[29]
Naoise Holohan, Spiros Antonatos, Stefano Braghin, and Pól Mac Aonghusa
-
[30]
Hanbin Hong, Binghui Wang, and Yuan Hong. 2022. UniCR: Universally Approx- imated Certified Robustness via Randomized Smoothing. InECCV
work page 2022
-
[31]
Hanbin Hong, Xinyu Zhang, Binghui Wang, Zhongjie Ba, and Yuan Hong. 2024. Certifiable Black-Box Attacks with Randomized Adversarial Examples: Breaking Defenses with Provable Confidence. InCCS. 600–614
work page 2024
-
[32]
Matthew Jagielski, Jonathan Ullman, and Alina Oprea. 2020. Auditing differen- tially private machine learning: How private is private SGD?NeurIPS(2020)
work page 2020
-
[33]
Jonggyu Jang, Seongjin Hwang, and Hyun Jong Yang. 2024. Rethinking DP-SGD in Discrete Domain: Exploring Logistic Distribution in the Realm of signSGD. In ICML
work page 2024
-
[34]
Peter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar, Abhradeep Thakurta, and Zheng Xu. 2021. Practical and private (deep) learning without sampling or shuffling. InICML. 5213–5225
work page 2021
-
[35]
Peter Kairouz, Sewoong Oh, and Pramod Viswanath. 2015. The composition theorem for differential privacy. InICML. 1376–1385
work page 2015
-
[36]
Peter Kairouz, Sewoong Oh, and Pramod Viswanath. 2017. The Composition Theorem for Differential Privacy.IEEE Information Theory(2017), 4037–4049
work page 2017
-
[37]
Alex Krizhevsky, Geoffrey Hinton, et al. 2009. Learning multiple layers of features from tiny images. (2009)
2009
-
[38]
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient- based learning applied to document recognition.Proc. IEEE86, 11 (1998)
work page 1998
-
[39]
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2022. Large Language Models Can Be Strong Differentially Private Learners. InICLR
work page 2022
-
[40]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020. RoBERTa: A Robustly Optimized BERT Pretraining Approach. InICLR
work page 2020
-
[41]
Albert W Marshall, Ingram Olkin, and Barry C Arnold. 1979. Inequalities: theory of majorization and its applications. (1979)
work page 1979
-
[42]
Ilya Mironov, Kunal Talwar, and Li Zhang. 2019. Rényi Differential Privacy of the Sampled Gaussian Mechanism. InNeurIPS
work page 2019
-
[43]
Meisam Mohammady, Shangyu Xie, Yuan Hong, Mengyuan Zhang, Lingyu Wang, Makan Pourzandi, and Mourad Debbabi. 2020. R2dp: A universal and automated approach to optimizing the randomization mechanisms of differential privacy for utility metrics with no known optimal distributions. InCCS
work page 2020
-
[44]
Nima Naderloui, Shenao Yan, Binghui Wang, Jie Fu, Wendy Hui Wang, Weiran Liu, and Yuan Hong. 2025. Rectifying Privacy and Efficacy Measurements in Machine Unlearning: A New Inference Attack Perspective. InUSENIX Security
work page 2025
-
[45]
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangx- iaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani...
work page 2021
-
[46]
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017. The E2E Dataset: New Challenges For End-to-End Generation. (2017)
work page 2017
-
[47]
Constrained Nonlinear Optimization. 2025. https://www.mathworks.com/help/ optim/ug/constrained-nonlinear-optimization-algorithms.html. (2025)
work page 2025
-
[48]
1992.Convex functions, partial orderings, and statistical applications
Josip E Peajcariaac and Yung Liang Tong. 1992.Convex functions, partial orderings, and statistical applications. Academic Press
work page 1992
-
[49]
Gege Qi, YueFeng Chen, Xiaofeng Mao, Binyuan Hui, Xiaodan Li, Rong Zhang, and Hui Xue. 2023. Model Inversion Attack via Dynamic Memory Learning. In MM. 5614–5622
work page 2023
-
[50]
Ahmed Salem, Apratim Bhattacharya, Michael Backes, Mario Fritz, and Yang Zhang. 2020. Updates-Leak: Data Set Inference and Reconstruction Attacks in Online Learning. InUSENIX Security 20. 1291–1308
work page 2020
-
[51]
Tom Sander, Pierre Stock, and Alexandre Sablayrolles. 2023. TAN without a burn: scaling laws of DP-SGD. InICML
work page 2023
-
[52]
Issai Schur. 1923. Uber eine Klasse von Mittelbildungen mit Anwendungen auf die Determinantentheorie.Sitzungsberichte der Berliner Mathematischen Gesellschaft 22, 9-20 (1923), 51
work page 1923
-
[53]
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017. Mem- bership inference attacks against machine learning models. In2017 IEEE Sympo- sium on Security and Privacy (SP). IEEE, 3–18
work page 2017
-
[54]
David Sommer, Sebastian Meiser, and Esfandiar Mohammadi. 2018. Privacy loss classes: The central limit theorem in differential privacy.Cryptology ePrint Archive(2018)
work page 2018
-
[55]
J. Michael Steele. 2004.The Cauchy-Schwarz Master Class: An Introduction to the Art of Mathematical Inequalities. Cambridge University Press
work page 2004
-
[56]
Michel Talagrand. 1996. Majorizing Measures: The Generic Chaining.The Annals of Probability24, 3 (1996), 1049–1103
work page 1996
-
[57]
Sasha Targ, Diogo Almeida, and Kevin Lyman. 2016. Resnet in Resnet: Generaliz- ing Residual Architectures. InICLR Workshop
work page 2016
-
[58]
Adrian M. Walker. 1965. Probability Theory and Mathematical Statistics.The Mathematical Gazette49 (1965), 109 – 112
work page 1965
-
[59]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. InICLR
work page 2019
-
[60]
Chendi Wang, Yuqing Zhu, Weijie J Su, and Yu-Xiang Wang. 2024. Neural Collapse meets Differential Privacy: Curious behaviors of NoisyGD with Near-Perfect Representation Learning. InICML, Vol. 235. 52334–52360
work page 2024
-
[61]
Yu-Xiang Wang, Borja Balle, and Shiva Prasad Kasiviswanathan. 2019. Subsam- pled rényi differential privacy and analytical moments accountant. InAISTATS
work page 2019
-
[62]
Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms.arXiv preprint arXiv:1708.07747(2017)
Pith/arXiv arXiv 2017
-
[63]
Da Yu, Saurabh Naik, et al. 2021. Differentially Private Fine-tuning of Language Models. InICLR
work page 2021
-
[64]
Yaodong Yu, Maziar Sanjabi, Yi Ma, Kamalika Chaudhuri, and Chuan Guo. 2023. Vip: A differentially private foundation model for computer vision.arXiv preprint arXiv:2306.08842(2023)
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[65]
Xinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang, Binghui Wang, Zhongjie Ba, and Kui Ren. 2024. Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks. InIEEE Symposium on Security and Privacy
work page 2024
-
[66]
Ligeng Zhu, Zhijian Liu, and Song Han. 2019.Deep leakage from gradients. Curran Associates Inc., Red Hook, NY, USA. A Probability Theory Backgrounds A.1 Moment Generating Function (MGF) Definition A.1.Let 𝑓(𝑥) be the probability density function (PDF) of a random variable𝑋 . The moment generating function of𝑋 , if it exists, is defined as: ℳ𝑋(𝑡)=E[𝑒 𝑡𝑋]= ...
work page 2019
-
[2018]
The bounded laplace mechanism in differential privacy.arXiv preprint arXiv:1808.10410(2018)
Pith/arXiv arXiv 2018
-
[2019]
InInternational Conference on Learning Representations (ICLR)
signSGD with Majority Vote is Communication Efficient and Fault Tolerant. InInternational Conference on Learning Representations (ICLR)
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.