Empirical scaling laws for task-specific LLM distillation in quantitative finance indicate that chain-of-thought supervision recovers general knowledge lost during iterative pruning while in-domain performance degrades predictably.
On the efficacy of knowledge distillation.ICCV 2019, 10 2019
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
TALAS is a knowledge distillation method that selectively aligns upper student layers to teacher sentence embeddings, propagates knowledge top-down via relational constraints in lower layers, and uses ASAM to seek flatter minima.
citing papers explorer
-
Scaling Laws for Task-Specific LLM Distillation
Empirical scaling laws for task-specific LLM distillation in quantitative finance indicate that chain-of-thought supervision recovers general knowledge lost during iterative pruning while in-domain performance degrades predictably.
-
TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding Distillation
TALAS is a knowledge distillation method that selectively aligns upper student layers to teacher sentence embeddings, propagates knowledge top-down via relational constraints in lower layers, and uses ASAM to seek flatter minima.