Pith. sign in

REVIEW 1 cited by

Federated Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.02367 v1 pith:VCPYWRA7 submitted 2020-11-04 cs.LG cs.DCcs.ITcs.NImath.IT

classification cs.LGcs.DCcs.ITcs.NImath.IT
keywords learningmodelcommunicationdistillationdistributedfederatedneuralpart
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Distributed learning frameworks often rely on exchanging model parameters across workers, instead of revealing their raw data. A prime example is federated learning that exchanges the gradients or weights of each neural network model. Under limited communication resources, however, such a method becomes extremely costly particularly for modern deep neural networks having a huge number of model parameters. In this regard, federated distillation (FD) is a compelling distributed learning solution that only exchanges the model outputs whose dimensions are commonly much smaller than the model sizes (e.g., 10 labels in the MNIST dataset). The goal of this chapter is to provide a deep understanding of FD while demonstrating its communication efficiency and applicability to a variety of tasks. To this end, towards demystifying the operational principle of FD, the first part of this chapter provides a novel asymptotic analysis for two foundational algorithms of FD, namely knowledge distillation (KD) and co-distillation (CD), by exploiting the theory of neural tangent kernel (NTK). Next, the second part elaborates on a baseline implementation of FD for a classification task, and illustrates its performance in terms of accuracy and communication efficiency compared to FL. Lastly, to demonstrate the applicability of FD to various distributed learning tasks and environments, the third part presents two selected applications, namely FD over asymmetric uplink-and-downlink wireless channels and FD for reinforcement learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tackling Data Heterogeneity in Federated Learning through Knowledge Distillation with Inequitable Aggregation

    cs.LG 2025-06 conditional novelty 5.0 of 10

    KDIA uses a triFreqs-weighted all-client teacher model plus knowledge distillation and a conditional generator to improve accuracy and convergence in large-client, low-participation heterogeneous federated learning.

Pith tools