SAGE adaptively augments high-loss embedding regions with UMAP-generated synthetic vectors and claims competitive distillation, though its average GLUE score (78.6) is below DistilBERT and MiniLM (79.4).
Do deep nets really need to be deep?Advances in Neural Information Processing Systems (NeurIPS), 2014
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
REJECT 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Synthetic Adaptive Guided Embeddings (SAGE): A Novel Knowledge Distillation Method
SAGE adaptively augments high-loss embedding regions with UMAP-generated synthetic vectors and claims competitive distillation, though its average GLUE score (78.6) is below DistilBERT and MiniLM (79.4).