QLPO resamples GRPO training groups to favor short correct and long incorrect responses, cutting reasoning length substantially while keeping accuracy roughly unchanged.
Large-Scale Self- and Semi-Supervised Learning for Speech Translation , booktitle =
1 Pith paper cite this work, alongside 15 external citations. Polarity classification is still indexing.
1
Pith paper citing it
15
external citations · OpenAlex
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization
QLPO resamples GRPO training groups to favor short correct and long incorrect responses, cutting reasoning length substantially while keeping accuracy roughly unchanged.