Pith. sign in

REVIEW 8 cited by

Teacher-Student Architecture for Knowledge Distillation: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.04268 v1 pith:E34DGV35 submitted 2023-08-08 cs.LG cs.AI

classification cs.LGcs.AI
keywords knowledgeteacher-studentarchitecturesdistillationobjectivessurveynetworkslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although Deep neural networks (DNNs) have shown a strong capacity to solve large-scale problems in many areas, such DNNs are hard to be deployed in real-world systems due to their voluminous parameters. To tackle this issue, Teacher-Student architectures were proposed, where simple student networks with a few parameters can achieve comparable performance to deep teacher networks with many parameters. Recently, Teacher-Student architectures have been effectively and widely embraced on various knowledge distillation (KD) objectives, including knowledge compression, knowledge expansion, knowledge adaptation, and knowledge enhancement. With the help of Teacher-Student architectures, current studies are able to achieve multiple distillation objectives through lightweight and generalized student networks. Different from existing KD surveys that primarily focus on knowledge compression, this survey first explores Teacher-Student architectures across multiple distillation objectives. This survey presents an introduction to various knowledge representations and their corresponding optimization objectives. Additionally, we provide a systematic overview of Teacher-Student architectures with representative learning algorithms and effective distillation schemes. This survey also summarizes recent applications of Teacher-Student architectures across multiple purposes, including classification, recognition, generation, ranking, and regression. Lastly, potential research directions in KD are investigated, focusing on architecture design, knowledge quality, and theoretical studies of regression-based learning, respectively. Through this comprehensive survey, industry practitioners and the academic community can gain valuable insights and guidelines for effectively designing, learning, and applying Teacher-Student architectures on various distillation objectives.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets

    cs.CE 2026-04 conditional novelty 7.0 of 10

    In a simulated economy of 25 LLM agents, a programmable market testbed shows that market rules and agent configuration reshape trade, quality, and wealth, with transparency and honesty norms backfiring.

  2. AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    AnyBody distills a privileged teacher tracker into a latent unit-sphere representation and uses a masked transformer to drive humanoid control from arbitrary keypoint subsets.

  3. KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning

    cs.AI 2026-04 unverdicted novelty 6.0 of 10

    KD-MARL distills both actions and coordination structure from expert MARL policies into heterogeneous lightweight students, retaining over 90% performance while cutting FLOPs by up to 28.6 times on SMAC and MPE benchmarks.

  4. Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets

    cs.CE 2026-04 unverdicted novelty 6.0 of 10

    DIAGON simulation shows agent markets produce 3.2 times more wealth than isolated agents, but institutional choices like transparency and competitive selection can reduce rather than increase performance.

  5. Doppler Prompting for Stable mmWave-based Human Pose Estimation

    cs.HC 2026-05 unverdicted novelty 5.0 of 10

    PULSE stabilizes mmWave human pose estimation by screening Doppler motion prompts before injecting them into spatial magnitude reasoning.

  6. Diagon: A Programmable Testbed for AI-Agent Cognitive Labor Markets

    cs.CE 2026-04 unverdicted novelty 5.0 of 10

    Market exchange among AI agents can raise productivity over self-sufficient agents, but institutional rules such as identity transparency and stronger selection can degrade those gains.

  7. Accelerating Reinforcement Learning Algorithms Convergence using Pre-trained Large Language Models as Tutors With Advice Reusing

    cs.LG 2025-09 conditional novelty 4.0 of 10

    LLM tutoring modestly accelerates RL convergence on average, with advice reuse saving wall-clock time but reducing stability.

  8. A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification

    cs.SD 2026-07 unverdicted novelty 3.0 of 10

    A challenge submission system using expanded training data, feature-specific branches, and post-processing achieves up to 81.25% hierarchical F1 on BSD10k-v1.2.

Pith tools