TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.
Lixin Wu, Na Cai, Qiao Cheng, Jiachen Wang, and Yitao Duan
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
GanitLLM improves Bengali math accuracy by 6-8 points over its base model by training on a difficulty-tagged corpus with Curriculum-GRPO that boosts Bengali reasoning tokens from 14% to 88%.
citing papers explorer
-
Trust Region On-Policy Distillation
TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.
-
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
GanitLLM improves Bengali math accuracy by 6-8 points over its base model by training on a difficulty-tagged corpus with Curriculum-GRPO that boosts Bengali reasoning tokens from 14% to 88%.