Training a single model on both concise and long-chain-of-thought math data creates a trade-off: more concise-answer supervision degrades long-reasoning accuracy, and the optimal training order depends on the data ratio.
Calc- X and Calcformers: Empowering Arithmetical Chain-of-Thought through Interaction with Symbolic Systems
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Fusion Training for Mathematical Generalization in Large Language Models
Training a single model on both concise and long-chain-of-thought math data creates a trade-off: more concise-answer supervision degrades long-reasoning accuracy, and the optimal training order depends on the data ratio.