DATG framework diagnoses that non-English reasoning in Qwen3 models shows reduced mathematical anchor coverage and dependency fidelity, with Loop-Retry and Formula-Retry improving target-language accuracy.
arXiv preprint arXiv:2508.14828
4 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 4years
2026 4verdicts
UNVERDICTED 4representative citing papers
Fine-tuning on matched native and English-pivoted multilingual reasoning datasets across six languages reduces the native reasoning gap to 1.9-3.5%; layer swap of English mid-layers largely closes the remaining gap while preserving target-language CoT.
UL-XCoT maintains competitive accuracy on multilingual benchmarks while cutting decoding tokens by over 50% through per-query language selection and logic-space trajectory pruning.
HiMed releases a Hindi medical reasoning corpus and benchmark and shows that training an 8B LLM with decaying scaffolding reward improves Hindi performance and narrows the English-Hindi accuracy gap.
citing papers explorer
-
Beyond Input Understanding: Diagnosing Multilingual Mathematical Reasoning with Directed Acyclic Trace Graphs
DATG framework diagnoses that non-English reasoning in Qwen3 models shows reduced mathematical anchor coverage and dependency fidelity, with Loop-Retry and Formula-Retry improving target-language accuracy.
-
Rethinking the Multilingual Reasoning Gap with Layer Swap
Fine-tuning on matched native and English-pivoted multilingual reasoning datasets across six languages reduces the native reasoning gap to 1.9-3.5%; layer swap of English mid-layers largely closes the remaining gap while preserving target-language CoT.
-
Less Languages, Less Tokens: An Efficient Unified Logic Cross-lingual Chain-of-Thought Reasoning Framework
UL-XCoT maintains competitive accuracy on multilingual benchmarks while cutting decoding tokens by over 50% through per-query language selection and logic-space trajectory pruning.
-
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
HiMed releases a Hindi medical reasoning corpus and benchmark and shows that training an 8B LLM with decaying scaffolding reward improves Hindi performance and narrows the English-Hindi accuracy gap.