Pith. sign in

REVIEW 1 cited by

Coupling public and private gradient provably helps optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01304 v1 pith:TTMYVVVN submitted 2023-10-02 cs.LG

classification cs.LG
keywords datacouplingoptimalprivatepubliccoefficientgradientoptimization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of large neural networks is crucially determined by the availability of data. It has been observed that training only on a small amount of public data, or privately on the abundant private data can lead to undesirable degradation of accuracy. In this work, we leverage both private and public data to improve the optimization, by coupling their gradients via a weighted linear combination. We formulate an optimal solution for the optimal weight in the convex setting to indicate that the weighting coefficient should be hyperparameter-dependent. Then, we prove the acceleration in the convergence of non-convex loss and the effects of hyper-parameters such as privacy budget, number of iterations, batch size, and model size on the choice of the weighting coefficient. We support our analysis with empirical experiments across language and vision benchmarks, and provide a guideline for choosing the optimal weight of the gradient coupling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Locally Private Sampling with Public Data

    cs.LG 2024-11 conditional novelty 5.0 of 10

    For finite discrete distributions, a recursive mechanism that preserves the public prior is exactly minimax optimal for every f-divergence, and reduces to randomized response when the prior is uniform.

Pith tools