Pith. sign in

REVIEW 6 cited by

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.02126 v1 pith:FFPE7BTQ submitted 2025-06-02 cs.CL

classification cs.CL
keywords reasoningknowledgemedicaldomainsmodelsaccuracydomainquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in reasoning-enhanced Large Language Models such as OpenAI-o1/3 and DeepSeek-R1 have significantly improved performance on complex tasks. However, the quality and transparency of their internal reasoning processes remain underexplored. This work moves beyond the final-answer accuracy and investigates step-by-step reasoning in the medical and mathematical domains by explicitly decomposing the thinking trajectories into two parts: knowledge and reasoning. Specifically, we introduce a fine-grained evaluation framework that judges: (1) the correctness of knowledge used (measured by Knowledge Index (KI)) and (2) the quality of reasoning (measured by Information Gain (InfoGain)). Using this framework, we study R1-distilled and base Qwen models trained with supervised fine-tuning (SFT) and/or reinforcement learning (RL) in the medical and math domains. Three intriguing findings emerge: (1) The general reasoning abilities in R1-distilled models do not transfer effectively to the medical domain through either SFT or RL. (2) SFT raises final-answer accuracy in both domains, but often at the cost of reasoning quality: InfoGain drops by 38.9% on average compared with untrained models; In the medical domain, however, SFT remains crucial because domain knowledge is indispensable. (3) RL enhances medical reasoning by pruning inaccurate or irrelevant knowledge from reasoning paths, thereby improving both reasoning accuracy and knowledge correctness.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    ClinSeekAgent automates active multimodal evidence seeking for clinical reasoning, improving LLM performance on raw EHR and CXR tasks while enabling distillation into smaller models.

  2. Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

    cs.CL 2026-05 unverdicted novelty 6.0 of 10

    Large reasoning models show measurable hidden-state dynamics that a new statistic can use to distinguish correct reasoning trajectories without labels.

  3. Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

    cs.LG 2026-02 conditional novelty 6.0 of 10

    For most benchmarks, the best achievable post-training accuracy follows a stable sigmoid curve in pre-training compute; math reasoning is the exception, with a boundary that keeps rising over time.

  4. Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach

    eess.SY 2025-08 unverdicted novelty 6.0 of 10

    The paper uses an LPV linearization and sparse RKHS estimators to detect the additive structure of continuous-time nonlinear models.

  5. CausalMix: Data Mixture as Causal Inference for Language Model Training

    cs.LG 2026-07 unverdicted novelty 5.0 of 10

    CausalMix fits a causal model on 512 runs of a 0.5B model to estimate CATE, then extrapolates optimal mixtures for an 800K data pool applied to 7B and 4B models, outperforming RegMix.

  6. LLM Parameters for Math Across Languages: Shared or Separate?

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Mechanistic analysis of LLMs finds partial overlap in math-associated parameters across languages, concentrated in middle layers, with systematic language-dependent differences.

Pith tools