Pith. sign in

REVIEW 1 cited by

DistilCSE: Effective Knowledge Distillation For Contrastive Sentence Embeddings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.05638 v2 pith:54TCGTOB submitted 2021-12-10 cs.AI

classification cs.AI
keywords distillationknowledgecontrastivemodelstudentmodelsdatadistilcse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale contrastive learning models can learn very informative sentence embeddings, but are hard to serve online due to the huge model size. Therefore, they often play the role of "teacher", transferring abilities to small "student" models through knowledge distillation. However, knowledge distillation inevitably brings some drop in embedding effect. To tackle that, we propose an effective knowledge distillation framework for contrastive sentence embeddings, termed DistilCSE. It first applies knowledge distillation on a large amount of unlabeled data, and then fine-tunes student models through contrastive learning on limited labeled data. To achieve better distillation results, we further propose Contrastive Knowledge Distillation (CKD). CKD uses InfoNCE as the loss function in knowledge distillation, enhancing the objective consistency among teacher model training, knowledge distillation, and student model fine-tuning. Extensive experiments show that student models trained with the proposed DistilCSE and CKD suffer from little or even no performance decrease and consistently outperform the corresponding counterparts of the same parameter size. Impressively, our 110M student model outperforms the latest state-of-the-art model, i.e., Sentence-T5 (11B), with only 1% parameters and 0.25% unlabeled data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantic Compression for Word and Sentence Embeddings using Discrete Wavelet Transform

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Keeping only the low-frequency DWT coefficients of word and sentence embeddings preserves most of their semantic quality at 50 to 93 percent fewer dimensions.

Pith tools