Pith. sign in

REVIEW 10 cited by

Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13950 v4 pith:AAOOEENJ submitted 2024-02-21 cs.CL

classification cs.CL
keywords reasoningstepsfrodoanswercausalfinalintermediatellms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) have been shown to perform better when asked to reason step-by-step before answering a question. However, it is unclear to what degree the model's final answer is faithful to the stated reasoning steps. In this paper, we perform a causal mediation analysis on twelve LLMs to examine how intermediate reasoning steps generated by the LLM influence the final outcome and find that LLMs do not reliably use their intermediate reasoning steps when generating an answer. To address this issue, we introduce FRODO, a framework to tailor small-sized LMs to generate correct reasoning steps and robustly reason over these steps. FRODO consists of an inference module that learns to generate correct reasoning steps using an implicit causal reward function and a reasoning module that learns to faithfully reason over these intermediate inferences using a counterfactual and causal preference objective. Our experiments show that FRODO significantly outperforms four competitive baselines. Furthermore, FRODO improves the robustness and generalization ability of the reasoning LM, yielding higher performance on out-of-distribution test sets. Finally, we find that FRODO's rationales are more faithful to its final answer predictions than standard supervised fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training Large Language Models for Self-Explanation Faithfulness

    cs.LG 2026-07 conditional novelty 6.0 of 10

    RL fine-tuning with a counterfactual mention/influence reward raises LLM self-explanation faithfulness (Phi-CCT) from near zero to ~0.66 in-distribution for two 8B models, with partial transfer to held-out tasks.

  2. Training Language Models to Use Prolog as a Tool

    cs.CL 2025-12 unverdicted novelty 6.0 of 10

    GRPO can teach a 3B language model to emit executable Prolog, but the highest-accuracy models often hardcode answers instead of reasoning in Prolog, producing an accuracy–auditability trade-off.

  3. CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

    cs.CL 2025-10 conditional novelty 6.0 of 10

    A consistency reward parsed by a 7B LLM, plus a two-stage refine-then-monitor pipeline and reformulated easy questions, improves MCQ-RL accuracy-with-consistency in law and medicine (58.9 vs 51.4 average Acc+).

  4. OwkinZero: Accelerating Biological Discovery with AI

    cs.LG 2025-08 conditional novelty 6.0 of 10

    RL-trained 8-32B models outperform larger commercial LLMs on new biology QA benchmarks, with mixed and partly overstated cross-task generalization.

  5. Do Cognitively Interpretable Reasoning Traces Improve LLM Performance?

    cs.CL 2025-08 conditional novelty 6.0 of 10

    On CoTemp QA, supervised fine-tuning with raw R1 traces gave the best model accuracy while human raters found those traces least interpretable, showing model-useful traces and human-readable traces can diverge.

  6. Graph-Guided Textual Explanation Generation Framework

    cs.CL 2024-12 conditional novelty 6.0 of 10

    G-Tex injects attention-based highlight tokens into a language model through a graph neural network layer, improving faithfulness of generated explanations.

  7. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  8. Self-Reflective Generation at Test Time

    cs.CL 2025-10 conditional novelty 5.0 of 10

    SRGen improves LLM math reasoning by detecting high-entropy tokens and injecting a small corrected vector into the hidden state at those points during decoding, without training.

  9. MetaOpenFOAM 2.0: Large Language Model Driven Chain of Thought for Automating CFD Simulation and Post-Processing

    cs.AI 2025-02 conditional novelty 5.0 of 10

    MetaOpenFOAM 2.0 uses chain-of-thought decomposition and iterative verification to let users run OpenFOAM CFD simulations and post-processing from natural language, achieving 86.9% pass@1 on a new 13-task benchmark.

  10. Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation

    cs.IR 2025-11 conditional novelty 4.0 of 10

    A review and position paper proposing a six-dimension success framework and risk diagnostics for evaluating LLM-based music recommendation systems.

Pith tools