In matched Bangla math training, chain-of-thought supervision does not beat answer-only training in-domain for strong backbones, but wins out-of-domain by 20 to 28 points, improving language adherence and auditable reasoning more than reasoning validity.
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
In matched Bangla math training, chain-of-thought supervision does not beat answer-only training in-domain for strong backbones, but wins out-of-domain by 20 to 28 points, improving language adherence and auditable reasoning more than reasoning validity.