MiniLLM distills large language models into smaller ones via reverse KL divergence and on-policy optimization, yielding higher-quality responses with lower exposure bias than standard KD baselines.
Light- paff: A two-stage distillation framework for pre-training and fine-tuning
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
CLPD improves LLM distillation for reasoning by combining explicit data curriculum with progressive teacher scheduling of increasing capacity.
PRIDE distills empathetic reasoning from large teacher LLMs to smaller students via an empathy prompt, multi-source attention, and dual-alignment loss using privileged information available only at training time.
citing papers explorer
-
MiniLLM: On-Policy Distillation of Large Language Models
MiniLLM distills large language models into smaller ones via reverse KL divergence and on-policy optimization, yielding higher-quality responses with lower exposure bias than standard KD baselines.
-
Curriculum Learning-Guided Progressive Distillation in Large Language Models
CLPD improves LLM distillation for reasoning by combining explicit data curriculum with progressive teacher scheduling of increasing capacity.
-
PRIDE: Privileged Information-enhanced Distillation for Empathetic Dialogue Generation
PRIDE distills empathetic reasoning from large teacher LLMs to smaller students via an empathy prompt, multi-source attention, and dual-alignment loss using privileged information available only at training time.
- Hybrid Policy Distillation for LLMs