Pith. sign in

REVIEW 3 major objections 5 minor 51 references

Japanese reasoning traces can be forced at 100% with continual pretraining plus GRPO, but capability and culture gains are not free.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 14:11 UTC pith:URLUZRY4

load-bearing objection Solid 8B Japanese-reasoning case study: CPT + two-stage GRPO gets 100% Japanese traces (even on English prompts) at modest capability cost, with no free cultural win; the hiragana metric is soft but the ablations still land. the 3 major comments →

arxiv 2607.10114 v1 pith:URLUZRY4 submitted 2026-07-11 cs.CL cs.AIcs.LG

Cost of Reasoning in non-English Languages: A Case Study on Japanese

classification cs.CL cs.AIcs.LG
keywords reasoning language modelsJapanese reasoning tracesGRPOcontinual pretraininglanguage rewardmultilingual reasoningcultural benchmarks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Reasoning models usually think in English because that is where most reasoning training data lives. The paper asks whether a model can be made to reason reliably in Japanese instead, without collapsing on coding, math, and science tasks. Using a Japanese continually pretrained 8B base, a short supervised warm-start on Japanese-format traces, and two-stage GRPO that first rewards format-plus-Japanese-plus-correctness and then tightens to strict correctness, the authors produce models that emit Japanese inside every think block—even on fully English prompts. Performance stays competitive with English-reasoning baselines on several tasks, but does not beat them, and drops on Japanese cultural and morality benchmarks. Continual Japanese pretraining turns out to be load-bearing: the same recipe applied to the English base never leaves English reasoning. The practical upshot is that language-of-thought control is possible, yet it carries a capability cost and does not automatically improve cultural alignment.

Core claim

Reasoning-language control in Japanese is feasible for an 8B model when Japanese continual pretraining is combined with a language reward under GRPO: the resulting models reach 100% Japanese reasoning traces on every evaluated benchmark, including English-instruction tasks, while remaining at best on par with strong English-reasoning baselines on coding, math, and science, and underperforming the same baselines on Japanese cultural benchmarks.

What carries the argument

Two-stage GRPO after Japanese continual pretraining and a short SFT warm-start: a permissive dense reward (format + Japanese language + answer correctness) first supplies learning signal, then a strict sparse reward enforces boxed correctness; without the Japanese pretraining substrate the language reward finds nothing to reinforce.

Load-bearing premise

That Japanese continual pretraining is jointly necessary because an English-centric base of the same size has no Japanese-capable region of policy that the language reward can reinforce.

What would settle it

Apply the identical SFT-plus-two-stage-GRPO recipe to another English-centric 8B base that has never seen Japanese continual pretraining and measure whether spontaneous Japanese reasoning traces ever exceed a few percent after many thousands of steps.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies whether an 8B model can be trained to produce Japanese reasoning traces while retaining competitive capability. Starting from Qwen-3-Swallow-8B (Japanese continual pretraining of Qwen-3-8B), the authors apply SFT warm-start on translated Japanese reasoning traces, then two-stage GRPO (permissive dense reward for format/language/answer, then strict sparse reward). They report that the resulting CAT-Thinking model yields 100% Japanese <think> blocks on every evaluated benchmark (including English-instruction tasks; Table 2), is competitive with English-reasoning Qwen-3 and Swallow on several coding/math/science sets (Figure 3), and underperforms Swallow on Japanese cultural/morality benchmarks (Table 3). Ablations claim Japanese CPT is necessary (Table 1: Qwen-3-8B stays at 3.1% Japanese after 15k GRPO steps) and that SFT warm-start prevents reward-std collapse (Figure 2a). Logit-lens and qualitative analyses document late-layer non-English preference and an emergent permission-phrase RL artifact.

Significance. If the feasibility and cost claims hold under a stronger language metric, the work is a useful empirical baseline for non-English reasoning-language control at 8B scale: it documents a concrete recipe (Japanese CPT + SFT warm-start + two-stage GRPO), shows that control is not free on capability or cultural tasks, and releases models. Strengths include transparent training dynamics (Figure 2), explicit ablations of CPT and warm-start, public model collection, and honest negative results on cultural benchmarks. The contribution is primarily engineering/empirical rather than theoretical; its value for the community is as a reproducible case study and caution against assuming near-free language redirection for typologically distant languages.

major comments (3)
  1. [Section 3.1, Table 2] Section 3.1 / Table 2: The central feasibility claim (100% Japanese reasoning on every benchmark) rests on a single heuristic—presence of at least one hiragana character in a <think> block. The authors note CJK ambiguity and that they cannot distinguish Japanese from Chinese (Section 3.4, Limitations), and logit-lens shows Chinese tokens for Qwen-3/Swallow (Figure 9). A mixed or Chinese-dominant trace with one hiragana (or the permission phrase of Section 4.1) would still score 100%. Appendix D examples look fully Japanese but are selected. Please report a stronger metric on the full evaluation set (e.g., fraction of Japanese characters/tokens, language-ID of the full <think> block, or human audit of a random sample) so that the 100% claim and the CPT-necessity result (Table 1) are not overstated.
  2. [Section 2.1, Table 1] Section 2.1 / Table 1: The joint-necessity claim that Japanese continual pretraining is required for the language reward to work is supported only by one base (Qwen-3-8B) at one scale, with Japanese appearing only when the prompt explicitly requests it. Without additional bases or scales, the claim that CPT establishes a necessary Japanese substrate is under-supported for the paper’s broader framing. Either qualify the claim as specific to this family/scale or add at least one further ablation (e.g., another English-centric base or a different Japanese CPT checkpoint).
  3. [Section 3, Figure 3] Figure 3 and evaluation protocol: Several training datasets post-date the benchmarks, and the authors note possible contamination and that absolute scores are reference points only. The cost claim (“at best on par”) is therefore hard to interpret as a clean capability gap. Please either (i) document contamination checks / decontamination, (ii) report results on a clearly post-cutoff or held-out set, or (iii) reframe the comparison as relative under matched decoding rather than absolute capability cost.
minor comments (5)
  1. [Figure 3] Figure 3 label: “PoyMath (Ja)” is misspelled; should be PolyMath.
  2. [Section 3.1] Section 3.1: State the hiragana heuristic explicitly in the main text when first introducing Table 2, not only later in Limitations.
  3. [Table 3, Section 3.3] Table 3 / Section 3.3: Jubaku is listed in the table but not defined or cited in the main cultural-benchmark discussion; add a short description and reference.
  4. [Section 4.1] Section 4.1: The permission-phrase rate (84.8%) is useful; consider reporting whether ablating the min-length-10 format reward removes the artifact, even if only as a small diagnostic.
  5. [Appendix D] Appendix D: Generation examples are valuable; adding character-level Japanese fraction for those traces would strengthen the qualitative support for Table 2.

Circularity Check

0 steps flagged

Empirical training study with external held-out metrics; no derivation reduces to its inputs by construction.

full rationale

This paper is an empirical case study of SFT warm-start plus two-stage GRPO on a Japanese continually pretrained 8B base. Its central claims (100% Japanese reasoning traces by a hiragana heuristic; competitive coding/math scores; worse cultural scores) are measured on held-out benchmarks and are not algebraically forced by the training reward or by any self-cited uniqueness theorem. The language component of the permissive reward is an optimization pressure, not a definition of the evaluation metric; Table 2 and Figure 3 report independent measurements. The CPT necessity argument (Table 1) is a controlled comparison, not a self-definitional identity. There are no fitted parameters re-labeled as predictions, no load-bearing self-citation of an unverified uniqueness result, and no renaming of a known closed-form result. Weaknesses of the hiragana heuristic are measurement-validity concerns, not circularity. Score 0 is therefore the correct finding.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 2 invented entities

The central feasibility claim rests on standard RL and LLM post-training machinery plus a few domain choices: that GRPO with a language reward can steer reasoning language once a Japanese substrate exists, that a hiragana heuristic is a valid proxy for Japanese reasoning, and that translated English reasoning traces plus English-heavy task mixtures are adequate training signal. No new physical entities are postulated; free parameters are ordinary training hyperparameters and reward design choices.

free parameters (4)
  • permissive reward component weights (format / language / answer)
    Additive partial rewards are described but exact numeric weights are not fixed in the main text; they shape whether the model receives learning signal before strict correctness.
  • SFT learning rate 7e-6 and early-stop before convergence
    Chosen so GRPO can still reshape style; stopping point is hand-selected and affects how much the model imitates translated English traces.
  • GRPO LRs 3e-6 then 5e-6, generations per prompt 4 then 8, max completion ~4096–4200
    Standard but consequential hyperparameters that control exploration, length, and the observed concise Japanese style.
  • minimum reasoning-trace length of 10 tokens in permissive stage
    Authors hypothesize this length threshold produced the emergent 'permission phrase' artifact in 84.8% of traces.
axioms (4)
  • domain assumption Japanese continual pretraining creates a Japanese-capable policy region that a language reward can reinforce; without it the language reward fails.
    Stated as joint necessity in Section 2.1 from the Qwen-3-8B control (Table 1).
  • ad hoc to paper Presence of at least one hiragana character in a <think> block is a valid measure of Japanese reasoning.
    Section 3.1 heuristic; cannot distinguish Japanese from mixed CJK and ignores fully kanji Japanese.
  • domain assumption GRPO with format+language+correctness rewards is a valid way to install a target reasoning language.
    Standard RLHF/GRPO assumption used throughout Sections 2.2–2.3; shared with related multilingual reasoning work.
  • domain assumption gpt-oss-120b translations of English reasoning traces and tasks are adequate supervision for Japanese reasoning style.
    Used for SFT warm-start and Japanese task counterparts; quality of translation is not independently validated.
invented entities (2)
  • CAT-Thinking / CAT-Paws Japanese-reasoning models independent evidence
    purpose: Concrete artifacts that always emit Japanese reasoning traces under the trained policy.
    New model checkpoints produced by the pipeline; independent evidence is the public HF collection and reported benchmark numbers.
  • Permission-phrase RL artifact no independent evidence
    purpose: Documents an emergent Japanese preamble reinforced by the length/format reward.
    Descriptive finding (84.8% of traces), not a postulated mechanism required for the feasibility claim.

pith-pipeline@v1.1.0-grok45 · 25870 in / 3365 out tokens · 39817 ms · 2026-07-14T14:11:36.282187+00:00 · methodology

0 comments
read the original abstract

Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented training data is most abundant. However, reasoning trace is a clue for model interpretability and safety, and useful in practice for both the model users and for model developers. Thus, it is desirable to be able to develop a model that reasons in a language of the user's choice, while still maintaining strong reasoning performance. To this end, we study the feasibility of training a model that reasons in Japanese. We develop a Japanese-reasoning variant of Qwen-3-Swallow-8B, which is a Japanese LLM continually pretrained from Qwen-3-8B, with GRPO and evaluate it across coding, math, and science benchmarks. The study shows that reasoning-language control is feasible by training a Japanese continually pretrained model with GRPO. However, its performance is at best on par with strong English-reasoning baselines on several benchmarks. We also evaluate the trained model on Japanese cultural benchmarks and observe that the model's performance is worse than the baseline models, suggesting that the reasoning in Japanese does not immediately improve performance on culturally relevant tasks for free.

Figures

Figures reproduced from arXiv: 2607.10114 by Yuu Jinnai.

Figure 1
Figure 1. Figure 1: Loss curves of the three models during SFT [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Training dynamics of the two GRPO stages. (a) Fraction of batches with zero reward standard deviation [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Evaluation on single-turn coding and math benchmarks. We compare our Japanese-reasoning models, [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Layerwise logit-lens Japanese-token prefer [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of the top-k tokens for CAT￾Thinking on the first instance of the GPQA Main (En). Japanese tokens appear to have high probability from the early layers. have similar curves in most of the layers except the last few layers, where CAT-Thinking shifts to non￾English tokens whereas Swallow shifts to English tokens. This suggests that the Japanese reasoning traces are not a surface-level effect, b… view at source ↗
Figure 6
Figure 6. Figure 6: Layerwise logit-lens Japanese-token prefer [PITH_FULL_IMAGE:figures/full_fig_p013_6.png] view at source ↗
Figure 8
Figure 8. Figure 8: Layerwise logit-lens Japanese-token prefer [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Visualization of the top-k tokens for Qwen-3 and Swallow on the first instance of the GPQA Main (En). 15 [PITH_FULL_IMAGE:figures/full_fig_p015_9.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

51 extracted references · 1 canonical work pages

  1. [1]

    Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. https://openreview.net/forum?id=3zKtaqxLhW On-policy distillation of language models: Learning from self-generated mistakes . In The Twelfth International Conference on Learning Representations

  2. [2]

    Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, and 106 others

    Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, and 106 others. 2025. gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925

  3. [3]

    Alon Albalak, Duy Phung, Nathan Lile, Rafael Rafailov, Kanishk Gandhi, Louis Castricato, Anikait Singh, Chase Blagden, Violet Xiang, Dakota Mahan, and Nick Haber. 2025. Big-math: A large-scale, high-quality math dataset for reinforcement learning in language models. arXiv preprint arXiv:2502.17387

  4. [4]

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. 2021. Program synthesis with large language models. arXiv preprint arXiv:2108.07732

  5. [5]

    Victor Barres, Honghua Dong, Soham Ray, Xujie Si, and Karthik R Narasimhan. 2026. https://openreview.net/forum?id=OC2z7iSQKa \ tau 2\ -bench: Evaluating conversational agents in a dual-control environment . In Forty-third International Conference on Machine Learning

  6. [6]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, and 39 others. 2021. Evaluating large language models trained on code. arXi...

  7. [7]

    Jasper Dekoninck, Nikola Jovanović, Tim Gehrunger, Kári Rögnvaldsson, Ivo Petrov, Chenhao Sun, and Martin Vechev. 2026. Beyond benchmarks: Matharena as an evaluation platform for mathematics with llms. arXiv preprint arXiv:2605.00674

  8. [8]

    Wei Fu, Jiaxuan Gao, Xujie Shen, Chen Zhu, Zhiyu Mei, Chuyi He, Shusheng Xu, Guo Wei, Jun Mei, WANG JIASHU, Tongkai Yang, Binhang Yuan, and Yi Wu. 2026. https://openreview.net/forum?id=X9diEuva9R AREAL : A large-scale asynchronous reinforcement learning system for language reasoning . In The Thirty-ninth Annual Conference on Neural Information Processing Systems

  9. [9]

    Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Hiroki Iida, Masanari Ohi, Kakeru Hattori, Hirai Shota, Sakae Mizuki, Rio Yokota, and Naoaki Okazaki. 2024. https://openreview.net/forum?id=TQdd1VhWbe Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities . In First Conference on Language Modeling

  10. [10]

    Schuck, Jo \ a o R

    Gabriel Lino Garcia, Andr \'e da F. Schuck, Jo \ a o R. R. Manesco, Pedro Henrique Paiola, Leandro A. Passos, and Jo \ a o Paulo Papa. 2026. https://aclanthology.org/2026.propor-1.95/ Think P ortuguese with bode reasoning . In Proceedings of the 17th International Conference on Computational Processing of P ortuguese ( PROPOR 2026) - Vol. 1 , pages 953--9...

  11. [11]

    Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, and Amelia Glaese. 2025. Deliberative alignment: Reasoning enables safer language models. arXiv preprint arXiv:2412.16339

  12. [12]

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z. F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao, Aixin Liu, and 175 others. 2025. Deep S eek- R 1: I ncentivizing R easoning C apability in LLM s via R einforcement L earning. arXiv preprint arX...

  13. [13]

    Daniil Gurgurov, Tom Röhr, Sebastian von Rohrscheidt, Josef van Genabith, Alexander Löser, and Simon Ostermann. 2026. ReasonXL : Shifting LLM reasoning language without sacrificing performance. arXiv preprint arXiv:2604.12378

  14. [14]

    Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021. https://openreview.net/forum?id=dNy_RKzJacY Aligning \ ai \ with shared human values . In International Conference on Learning Representations

  15. [15]

    u botter, Frederike L \

    Jonas H \"u botter, Frederike L \"u beck, Lejs Deen Behric, Anton Baumann, Marco Bagatella, Daniel Marta, Ido Hakimi, Idan Shenfeld, Thomas Kleine Buening, Carlos Guestrin, and Andreas Krause. 2026. https://openreview.net/forum?id=QkfkxyRizZ Reinforcement learning via self-distillation . In Forty-third International Conference on Machine Learning

  16. [16]

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. 2025. https://openreview.net/forum?id=chfJJYC3iL Livecodebench: Holistic and contamination free evaluation of large language models for code . In The Thirteenth International Conference on Learning Representations

  17. [17]

    Fengqing Jiang, Zhangchen Xu, Yuetai Li, Luyao Niu, Zhen Xiang, Bo Li, Bill Yuchen Lin, and Radha Poovendran. 2025. https://doi.org/10.18653/v1/2025.findings-acl.1197 S afe C hain: Safety of language models with long chain-of-thought reasoning capabilities . In Findings of the Association for Computational Linguistics: ACL 2025, pages 23303--23320, Vienna...

  18. [18]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles

  19. [19]

    Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, Karina Nguyen, Newton Cheng, Nicholas Joseph, Nicholas Schiefer, Oliver Rausch, Robin Larson, Sam McCandlish, Sandipan Kundu, and 11 others. 2023. Measuring faithfulness in chain-of-though...

  20. [20]

    Yifei Liu, Li Lyna Zhang, Yi Zhu, bingcheng dong, Xudong Zhou, Ning Shang, Fan Yang, Cheng Li, and Mao Yang. 2026. https://openreview.net/forum?id=NzPwDutzz8 rstar-coder: Scaling competitive code reasoning with a large-scale verified dataset . In The Thirty-ninth Annual Conference on Neural Information Processing Systems

  21. [21]

    LLM-jp, :, Akiko Aizawa, Eiji Aramaki, Bowen Chen, Fei Cheng, Hiroyuki Deguchi, Rintaro Enomoto, Kazuki Fujii, Kensuke Fukumoto, Takuya Fukushima, Namgi Han, Yuto Harada, Chikara Hashimoto, Tatsuya Hiraoka, Shohei Hisada, Sosuke Hosokawa, Lu Jie, Keisuke Kamata, and 64 others. 2024. Llm-jp: A cross-organizational project for the research and development o...

  22. [22]

    Youmi Ma, Sakae Mizuki, Kazuki Fujii, Taishi Nakamura, Masanari Ohi, Hinari Shimada, Taihei Shiotani, Koshiro Saito, Koki Maeda, Kakeru Hattori, Takumi Okamoto, Shigeki Ishida, Rio Yokota, Hiroya Takamura, and Naoaki Okazaki. 2025. Building instruction-tuning datasets from human-written instructions with open-weight large language models. In Second Confer...

  23. [24]

    Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Cand \`e s, and Tatsunori Hashimoto. 2025. https://doi.org/10.18653/v1/2025.emnlp-main.1025 s1: Simple test-time scaling . In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 2027...

  24. [25]

    nostalgebraist. 2020. interpreting gpt: the logit lens. https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens

  25. [26]

    Qwen Team . 2026. https://qwen.ai/blog?id=qwen3.6-27b Qwen3.6-27B : Flagship-level coding in a 27B dense model

  26. [27]

    Leonardo Ranaldi and Giulia Pucci. 2025. https://doi.org/10.18653/v1/2025.naacl-long.577 Multilingual reasoning via self-training . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 11566--11582, Albuquerque, New Mexico. ...

  27. [28]

    Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020. https://doi.org/10.1145/3394486.3406703 Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters . In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD '20, page 3505–3506, New York, NY, U...

  28. [29]

    David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024. https://openreview.net/forum?id=Ti67584b98 GPQA : A graduate-level google-proof q&a benchmark . In First Conference on Language Modeling

  29. [30]

    Narutatsu Ri, Abhishek Panigrahi, and Sanjeev Arora. 2026. Do thinking tokens help with safety? arXiv preprint arXiv:2606.25013

  30. [31]

    Alan Saji, Raj Dabre, Anoop Kunchukuttan, and Ratish Puduppully. 2026. https://doi.org/10.18653/v1/2026.eacl-short.25 The reasoning lingua franca: A double-edged sword for multilingual AI . In Proceedings of the 19th Conference of the E uropean Chapter of the A ssociation for C omputational L inguistics (Volume 2: Short Papers) , pages 329--344, Rabat, Mo...

  31. [32]

    Wonduk Seo, Wonseok Choi, Junseo Koh, Juhyeon Lee, Hyunjin An, Minhyeong Yu, Jian Park, Qingshan Zhou, Seunghyun Lee, and Yi Bu. 2026. Toward culturally aligned llms through ontology-guided multi-agent reasoning. arXiv preprint arXiv:2601.21700

  32. [33]

    Idan Shenfeld, Mehul Damani, Jonas H \"u botter, and Pulkit Agrawal. 2026. https://openreview.net/forum?id=qA6FgH0nnZ Self-distillation enables continual learning . In Forty-third International Conference on Machine Learning

  33. [34]

    Zayne Rea Sprague, Fangcong Yin, Juan Diego Rodriguez, Dongwei Jiang, Manya Wadhwa, Prasann Singhal, Xinyu Zhao, Xi Ye, Kyle Mahowald, and Greg Durrett. 2025. https://openreview.net/forum?id=w6nlcS8Kkn To cot or not to cot? chain-of-thought helps mainly on math and symbolic reasoning . In The Thirteenth International Conference on Learning Representations

  34. [35]

    Richard S Sutton and Andrew G Barto. 2018. Reinforcement learning: An introduction. MIT press

  35. [36]

    Sho Takase and Ukyo Honda. 2026. Toward LLM s beyond E nglish-centric development. arXiv preprint arXiv:2506.10910

  36. [37]

    Masashi Takeshita, Rafal Rzepka, and Kenji Araki. 2023. https://www.anlp.jp/proceedings/annual_meeting/2023/pdf_dir/D2-1.pdf Jcommonsensemorality: Japanese dataset for evaluating commonsense morality understanding . In In Proceedings of The Twenty Nineth Annual Meeting of The Association for Natural Language Processing (NLP2023), pages 357--362. In Japanese

  37. [38]

    GLM Team, Aohan Zeng, Xin Lv, Qinkai Zheng, Zhenyu Hou, Bin Chen, Chengxing Xie, Cunxiang Wang, Da Yin, Hao Zeng, Jiajie Zhang, Kedong Wang, Lucen Zhong, Mingdao Liu, Rui Lu, Shulin Cao, Xiaohan Zhang, Xuancheng Huang, Yao Wei, and 152 others. 2025. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models. arXiv preprint arXiv:2508.06471

  38. [39]

    The Microsoft AI Team. 2026. https://microsoft.ai/pdf/mai-thinking-1.pdf Mai-thinking-1: Building a hill-climbing machine . Technical report, Microsoft AI

  39. [40]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. https://openreview.net/forum?id=bzs4uPLXvi Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting . In Thirty-seventh Conference on Neural Information Processing Systems

  40. [41]

    Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec. 2020. https://github.com/huggingface/trl TRL: Transformers Reinforcement Learning

  41. [42]

    Weixuan Wang, Minghao Wu, Barry Haddow, and Alexandra Birch. 2025. https://doi.org/10.18653/v1/2025.findings-emnlp.519 Demystifying multilingual reasoning in process reward modeling . In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9775--9788, Suzhou, China. Association for Computational Linguistics

  42. [43]

    Yiming Wang, Pei Zhang, Jialong Tang, Hao-Ran Wei, Baosong Yang, Rui Wang, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang, Junyang Lin, Fei Huang, and Jingren Zhou. 2026. https://openreview.net/forum?id=B1vCImy6yI Polymath: Evaluating mathematical reasoning in multilingual contexts . In The Thirty-ninth Annual Con...

  43. [44]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903

  44. [45]

    Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024. https://doi.org/10.18653/v1/2024.findings-naacl.138 How interpretable are reasoning explanations from prompting large language models? In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2148--2164, Mexico City, Mexico. Association for Computational Linguistics

  45. [46]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others. 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6 Transformers: Sta...

  46. [47]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388

  47. [48]

    Farid Adilazuarda, Jonibek Mansurov, Ruochen Zhang, Niklas Muennighoff, Carsten Eickhoff, Genta Indra Winata, Julia Kreutzer, Stephen H

    Zheng-Xin Yong, M. Farid Adilazuarda, Jonibek Mansurov, Ruochen Zhang, Niklas Muennighoff, Carsten Eickhoff, Genta Indra Winata, Julia Kreutzer, Stephen H. Bach, and Alham Fikri Aji. 2025. Crosslingual reasoning through test-time scaling. arXiv preprint arXiv:2505.05408

  48. [49]

    Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, YuYue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, LingJun Liu, Xin Liu, Haibin Lin, Zhiqi Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang, Mofan Zhang, and 17 others. 2026. https://openreview.net/forum?id=2a36EMSSTp DAPO : An open-source LLM reinforcement learning system at scale...

  49. [50]

    Weixiang Zhao, Jiahe Guo, Yang Deng, Tongtong Wu, Wenxuan Zhang, Yulin Hu, Xingyu Sui, Yanyan Zhao, Wanxiang Che, Bing Qin, Tat-Seng Chua, and Ting Liu. 2025. https://openreview.net/forum?id=fleQlZ2VTx When less language is more: Language-reasoning disentanglement makes LLM s better multilingual reasoners . In The Thirty-ninth Annual Conference on Neural ...

  50. [51]

    Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, Alban Desmaison, Can Balioglu, Pritam Damania, Bernard Nguyen, Geeta Chauhan, Yuchen Hao, Ajit Mathews, and Shen Li. 2023. Pytorch fsdp: Experiences on scaling fully sharded data parallel. arXiv preprint arXiv:2304.11277

  51. [52]

    Wenhao Zhu, Shujian Huang, Fei Yuan, Cheng Chen, Jiajun Chen, and Alexandra Birch. 2024. The power of question alignment in multilingual reasoning: Broadened scope and deepened insights. arXiv preprint arXiv:2405.01345