Pith. sign in

REVIEW 1 cited by

Privately Customizing Prefinetuning to Better Match User Data in Federated Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.09042 v2 pith:MSUT2FXC submitted 2023-02-17 cs.LG cs.AIcs.DC

classification cs.LGcs.AIcs.DC
keywords datasetfederatedprefinetuningprivatedataprivatelydistancefred
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In Federated Learning (FL), accessing private client data incurs communication and privacy costs. As a result, FL deployments commonly prefinetune pretrained foundation models on a (large, possibly public) dataset that is held by the central server; they then FL-finetune the model on a private, federated dataset held by clients. Evaluating prefinetuning dataset quality reliably and privately is therefore of high importance. To this end, we propose FreD (Federated Private Fr\'echet Distance) -- a privately computed distance between a prefinetuning dataset and federated datasets. Intuitively, it privately computes and compares a Fr\'echet distance between embeddings generated by a large language model on both the central (public) dataset and the federated private client data. To make this computation privacy-preserving, we use distributed, differentially-private mean and covariance estimators. We show empirically that FreD accurately predicts the best prefinetuning dataset at minimal privacy cost. Altogether, using FreD we demonstrate a proof-of-concept for a new approach in private FL training: (1) customize a prefinetuning dataset to better match user data (2) prefinetune (3) perform FL-finetuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Distributed, communication-efficient, and differentially private estimation of KL divergence

    cs.LG 2024-11 reject novelty 5.0 of 10

    PRIEST-KLD is a family of differentially private, communication-efficient estimators of KL divergence for federated data, with three trust models; however, the unbiasedness and privacy proofs have load-bearing gaps.

Pith tools