Pith. sign in

Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 6 2025 1

roles

background 1

polarities

background 1

representative citing papers

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold

cs.AI · 2026-04-15 · unverdicted · novelty 6.0 · 2 refs

ReSS extracts decision paths from trees as scaffolds to guide LLM reasoning generation, fine-tunes the LLM on the resulting dataset with scaffold-invariant augmentation, and reports up to 10% gains on medical and financial tabular benchmarks with new faithfulness metrics.

Training Language Models to Use Prolog as a Tool

cs.CL · 2025-12-08 · conditional · novelty 6.0

GRPO can teach a 3B language model to emit executable Prolog, but the highest-accuracy models often hardcode answers instead of reasoning in Prolog, producing an accuracy–auditability trade-off.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

cs.AI · 2026-05-23 · unverdicted · novelty 5.0 · 3 refs

Proposes a multi-dimensional behavioral framework with six dimensions (Correctness, Consistency, Robustness, Local Logical Coherence, Efficiency, Stability) plus deployment-aware aggregation to diagnose LLM reasoning beyond accuracy-based benchmarks.

citing papers explorer

Showing 7 of 7 citing papers.