Using an LLM judge to reward semantic conciseness during reinforcement learning makes 1.5B reasoning models produce much shorter traces with comparable or better accuracy.
Check \( G > B \): 26 > 9, which is true
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
Using an LLM judge to reward semantic conciseness during reinforcement learning makes 1.5B reasoning models produce much shorter traces with comparable or better accuracy.