REVIEW 4 cited by
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) demonstrate exceptional performance on complex reasoning tasks. However, despite their strong reasoning capabilities in high-resource languages (e.g., English and Chinese), a significant performance gap persists in other languages. To investigate this gap in Korean, we introduce HRM8K, a benchmark comprising 8,011 English-Korean parallel bilingual math problems. Through systematic analysis of model behaviors, we identify a key finding: these performance disparities stem primarily from difficulties in comprehending non-English inputs, rather than limitations in reasoning capabilities. Based on these findings, we propose UST (Understand, Solve, and Translate), a method that strategically uses English as an anchor for reasoning and solution generation. By fine-tuning the model on 130k synthetically generated data points, UST achieves a 10.91% improvement on the HRM8K benchmark and reduces the multilingual performance gap from 11.6% to 0.7%. Additionally, we show that improvements from UST generalize effectively to different Korean domains, demonstrating that capabilities acquired from machine-verifiable content can be generalized to other areas. We publicly release the benchmark, training dataset, and models.
Forward citations
Cited by 4 Pith papers
-
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
An integrated survey organizing AI mathematical reasoning into informal, formal, discovery, and technique axes while cataloging benchmarks and assessing failure modes.
-
EfficientXLang: Towards Improving Token Efficiency Through Cross-Lingual Reasoning
Reasoning in non-English languages reduces thinking tokens by 20-40% while largely preserving math accuracy, with savings persisting after translation to English.
-
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback
A large LLM-to-LLM math tutoring simulation across 11 languages shows English-language hints often yield the largest accuracy gains for student models, but the low-resource-language results lack statistical support.
-
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.
Discussion (0). Sign in to comment.