Pith. sign in

REVIEW 2 cited by

Communication-Efficient Federated Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.00632 v1 pith:3TPTJS7W submitted 2020-12-01 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords federatedcommunicationdistillationlearningwhencompareddatamagnitude
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Communication constraints are one of the major challenges preventing the wide-spread adoption of Federated Learning systems. Recently, Federated Distillation (FD), a new algorithmic paradigm for Federated Learning with fundamentally different communication properties, emerged. FD methods leverage ensemble distillation techniques and exchange model outputs, presented as soft labels on an unlabeled public data set, between the central server and the participating clients. While for conventional Federated Learning algorithms, like Federated Averaging (FA), communication scales with the size of the jointly trained model, in FD communication scales with the distillation data set size, resulting in advantageous communication properties, especially when large models are trained. In this work, we investigate FD from the perspective of communication efficiency by analyzing the effects of active distillation-data curation, soft-label quantization and delta-coding techniques. Based on the insights gathered from this analysis, we present Compressed Federated Distillation (CFD), an efficient Federated Distillation method. Extensive experiments on Federated image classification and language modeling problems demonstrate that our method can reduce the amount of communication necessary to achieve fixed performance targets by more than two orders of magnitude, when compared to FD and by more than four orders of magnitude when compared with FA.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Provably Near-Optimal Federated Ensemble Distillation with Negligible Overhead

    cs.LG 2025-02 conditional novelty 6.0 of 10

    FedGO weights each client's prediction by the estimated density ratio of that client's data using GAN discriminators, and proves this weighting makes the ensemble at least as good as the best single model under convex loss.

  2. Hypernetworks for Model-Heterogeneous Personalized Federated Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A server-side multi-head hypernetwork generates personalized parameters for clients with heterogeneous model architectures, plus an optional global-model distillation variant, and beats several pFL baselines on four b...

Pith tools