Pith. sign in

Knowledge distillation of black-box large language models, 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

baseline 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

baseline 1

polarities

baseline 1

representative citing papers

Online Knowledge Distillation with Reward Guidance

cs.LG · 2025-05-25 · conditional · novelty 6.0

A preference-based knowledge distillation framework uses a confidence-set reward model in a min-max imitation game, with offline, online, and white-box variants, and outperforms prior KD baselines on LLM benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Online Knowledge Distillation with Reward Guidance cs.LG · 2025-05-25 · conditional · none · ref 9

    A preference-based knowledge distillation framework uses a confidence-set reward model in a min-max imitation game, with offline, online, and white-box variants, and outperforms prior KD baselines on LLM benchmarks.