{"work":{"id":"5d347bac-2a53-443b-b727-d4ed3ea9be1a","openalex_id":"https://openalex.org/W4416525378","doi":"10.48550/arxiv.2507.00432","arxiv_id":"2507.00432","raw_key":null,"title":"Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning","authors":null,"authors_text":"Maggie Huan, Yuetai Li, Tuney Zheng, Xiaoyu Xu, Seungone Kim, Minxin Du","year":2025,"venue":"cs.AI","abstract":"Math reasoning has become the poster child of progress in large language models (LLMs), with new models rapidly surpassing human-level performance on benchmarks like MATH and AIME. But as math leaderboards improve week by week, it is worth asking: do these gains reflect broader problem-solving ability or just narrow overfitting? To answer this question, we evaluate over 20 open-weight reasoning-tuned models across a broad suite of tasks, including math, scientific QA, agent planning, coding, and standard instruction-following. We surprisingly find that most models that succeed in math fail to transfer their gains to other domains. To rigorously study this phenomenon, we conduct controlled experiments on Qwen3-14B models using math-only data but different tuning methods. We find that reinforcement learning (RL)-tuned models generalize well across domains, while supervised fine-tuning (SFT)-tuned models often forget general capabilities. Latent-space representation and token-space distribution shift analyses reveal that SFT induces substantial representation and output drift, while RL preserves general-domain structure. Our results suggest a need to rethink standard post-training recipes, particularly the reliance on SFT-distilled data for advancing reasoning models.","external_url":"https://arxiv.org/abs/2507.00432","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2507.00432","created_at":"2026-05-10T05:25:55.283147+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning","render_title":"Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning"},"hub":{"state":{"work_id":"5d347bac-2a53-443b-b727-d4ed3ea9be1a","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":31,"external_cited_by_count":1,"distinct_field_count":4,"first_pith_cited_at":"2025-03-12T17:35:03+00:00","last_pith_cited_at":"2026-07-02T07:07:28+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T12:39:42.445997+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2},{"context_role":"method","n":2}],"polarity_counts":[{"context_polarity":"background","n":2},{"context_polarity":"use_method","n":2}],"runs":{},"summary":{},"graph":{},"authors":[]}}