s-step self-distillation is optimal among spectral shrinkage estimators for s-spiked covariance matrices and necessary for optimality.
Towards a theory of model distillation
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
InstructMPC uses an LLM plus tunable last layer to map operational context to disturbance trajectories for MPC, proving an O(sqrt(T log T)) regret bound for linear systems and showing lower grid costs on the OpenCEM microgrid.
Distilling CoT from DeepSeek-R1 to Qwen2.5-7B on competition problems yields 4.76 pp accuracy gain to 69.43% and 73.1% on MATH-500, with accuracy falling as response length decreases.
citing papers explorer
-
Self-Distillation is Optimal Among Spectral Shrinkage Estimators in Spiked Covariance Models
s-step self-distillation is optimal among spectral shrinkage estimators for s-spiked covariance matrices and necessary for optimality.
-
Context-Aware Model Predictive Control for Microgrid Energy Management via LLMs
InstructMPC uses an LLM plus tunable last layer to map operational context to disturbance trajectories for MPC, proving an O(sqrt(T log T)) regret bound for linear systems and showing lower grid costs on the OpenCEM microgrid.
-
Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition
Distilling CoT from DeepSeek-R1 to Qwen2.5-7B on competition problems yields 4.76 pp accuracy gain to 69.43% and 73.1% on MATH-500, with accuracy falling as response length decreases.