Pith. sign in

Gpt-4 is openai’s most advanced system, producing safer and more useful responses, 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

other 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

other 1

polarities

unclear 1

representative citing papers

Online Knowledge Distillation with Reward Guidance

cs.LG · 2025-05-25 · conditional · novelty 6.0

A preference-based knowledge distillation framework uses a confidence-set reward model in a min-max imitation game, with offline, online, and white-box variants, and outperforms prior KD baselines on LLM benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Online Knowledge Distillation with Reward Guidance cs.LG · 2025-05-25 · conditional · none · ref 7

    A preference-based knowledge distillation framework uses a confidence-set reward model in a min-max imitation game, with offline, online, and white-box variants, and outperforms prior KD baselines on LLM benchmarks.