Pith. sign in

REVIEW 5 cited by

Self-Knowledge Distillation in Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.01851 v1 pith:O5CJU4SZ submitted 2019-08-02 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords deepdistillationmethodknowledgelanguagelearningproposedsoft
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Since deep learning became a key player in natural language processing (NLP), many deep learning models have been showing remarkable performances in a variety of NLP tasks, and in some cases, they are even outperforming humans. Such high performance can be explained by efficient knowledge representation of deep learning models. While many methods have been proposed to learn more efficient representation, knowledge distillation from pretrained deep networks suggest that we can use more information from the soft target probability to train other neural networks. In this paper, we propose a new knowledge distillation method self-knowledge distillation, based on the soft target probabilities of the training model itself, where multimode information is distilled from the word embedding space right below the softmax layer. Due to the time complexity, our method approximates the soft target probabilities. In experiments, we applied the proposed method to two different and fundamental NLP tasks: language model and neural machine translation. The experiment results show that our proposed method improves performance on the tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Active Data Curation Effectively Distills Large-Scale Multimodal Models

    cs.CV 2024-11 conditional novelty 7.0 of 10

    Selecting training data by a reference model's loss acts as an implicit distillation, and combining it with explicit distillation yields more FLOP-efficient vision-language models that beat prior SoTA on 27 benchmarks.

  2. Dynamic Contrastive Knowledge Distillation for Efficient Image Restoration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Dynamic contrastive knowledge distillation with EMA-generated negatives and VQGAN codebook distribution alignment improves compact image restoration students.

  3. From Collapse to Stability: A Knowledge-Driven Ensemble Framework for Scaling Up Click-Through Rate Prediction Models

    cs.IR 2024-11 conditional novelty 6.0 of 10

    KDEF combines knowledge distillation and deep mutual learning with adaptive exam-score weighting so that CTR ensembles with up to ten sub-networks improve instead of collapse.

  4. Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models

    cs.CL 2024-11 conditional novelty 5.0 of 10

    DynSDPB fine-tunes small language models by self-distilling soft labels from the previous mini-batch, with dynamic per-sample temperature and loss weighting.

  5. Metric Learning with Progressive Self-Distillation for Audio-Visual Embedding Learning

    cs.SD 2025-01 conditional novelty 4.0 of 10

    A self-distillation training scheme for audio-visual embeddings progressively replaces labeled triplets with model-generated soft alignments, improving cross-modal retrieval MAP by roughly 2 percent on AVE and VEGAS.

Pith tools