Pith. sign in

REVIEW 3 major objections 6 minor 39 references

Behavioral detectors of unfaithful chain-of-thought mostly rediscover wrong answers, and go blind exactly where models fail.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 21:40 UTC pith:NBCQPJX4

load-bearing objection Correctness splits CoT unfaithfulness detection into two regimes; black-box signals mostly fail where models are wrong, and the audit is careful enough to take seriously. the 3 major comments →

arxiv 2607.23458 v1 pith:NBCQPJX4 submitted 2026-07-26 cs.CL

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

classification cs.CL
keywords chain-of-thoughtfaithfulnessunfaithfulness detectionanswer correctnessbehavioral detectorslinear probeshint-induced rationalizationstep-removal metrics
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Chain-of-thought explanations are useful for oversight only if the stated steps actually produced the answer. This paper audits black-box detectors against human faithfulness labels and finds that answer correctness organizes the whole problem. Simply knowing the answer is wrong already beats every purpose-built behavioral signal, because most annotated unfaithfulness sits on incorrect answers. Split the data by correctness and two regimes appear: on correct answers, surface signals can moderately separate genuine reasoning from post-hoc rationalization; on incorrect answers—the regime that holds most of the unfaithfulness—no tested behavioral signal rises above chance. The standard step-removal faithfulness metric even points the wrong way. Internals can decode each regime in some models, but the encodings do not share a common direction, and the usual way of manufacturing unfaithful traces by instruction fails to transfer to the human-labeled regimes.

Core claim

Answer correctness structures instance-level detection of unfaithful chain-of-thought. Incorrectness alone reaches AUROC 0.696 against human labels and outperforms every purpose-built black-box signal, because about 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness yields two regimes with opposite character: on correct answers, behavioral signals moderately separate faithful reasoning from post-hoc rationalization (roughly 0.63–0.67); on incorrect answers, where most unfaithfulness lives, no tested black-box signal is detectably above chance. The usual step-removal metric anti-correlates with human labels, and that inversion reproduces on the benchmark’s

What carries the argument

Two-regime stratification by answer correctness: the correct-answer regime (faithful reasoning versus post-hoc rationalization) versus the incorrect-answer regime (honest error versus unfaithful error). This split, not any single detector family, carries the argument—it explains aggregate detector failure, the step-removal metric inversion, and why internal probes and constructed testbeds transfer unevenly.

Load-bearing premise

The human faithfulness labels on this one benchmark are reliable enough ground truth for the two-regime split, even with small minority classes and some label noise.

What would settle it

On a larger, independently annotated set of incorrect-answer traces, show a black-box behavioral signal whose AUROC confidence interval clearly excludes chance for honest versus unfaithful error; or show that step-removal scores align positively with those human labels once correctness is controlled.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Black-box CoT auditing in deployment largely re-detects answer correctness and offers little help on the high-stakes incorrect-answer cases.
  • Scoring new detectors against step-removal metrics is circular and can invert the true faithfulness ranking.
  • Training or evaluating detectors on instructed answer-first rationalizations measures the wrong phenomenon; they do not transfer to human-annotated regimes.
  • Hint-induced unverbalized answer flips can share linearly decodable structure with annotated unfaithful errors, but only in model- and source-dependent ways.
  • White-box probes can recover regime distinctions behavior misses, yet no single shared linear direction covers both regimes across models.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Oversight pipelines that gate on surface CoT ‘suspicion’ will systematically under-scrutinize the very errors that matter most.
  • If honest and unfaithful errors are encoded along different internal directions, a one-probe monitor is unlikely to generalize across failure modes without regime conditioning.
  • Future faithfulness benchmarks may need forced balance across correct and incorrect answers, or the field will keep rediscovering accuracy.
  • The documentation–data label mismatch is a caution that any multi-axis annotation scheme should be re-derived from raw fields before secondary analyses.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The manuscript audits black-box (behavioral) detection of chain-of-thought unfaithfulness against FaithCoT-Bench's human annotations and argues that answer correctness structures the entire problem. Its claims: (i) answer incorrectness alone (oracle) outperforms every purpose-built behavioral signal (AUROC 0.696), since ~69% of annotated unfaithfulness sits on incorrect answers; (ii) stratifying by correctness yields two regimes — on correct answers, NLI step-support and (inverted) answer-tracing/prefix-instability signals reach 0.63–0.67, while on incorrect answers no tested signal is detectably above chance, replicated on benchmark-wide signals across all four models; (iii) the benchmark's step-removal metric anti-correlates with human labels (intended-direction AUROC ~0.35), an inversion reproduced from the release's own stored scores and replicated on hint-induced counterfactually labeled traces; (iv) linear probes decode the blind regime in Llama-3.1-8B and the correct regime in Qwen-2.5-7B with no cross-regime transfer; (v) instructed answer-first constructions (7 models) transfer to neither annotated regime, while hint-induced unverbalized flips transfer to the incorrect regime in a model- and source-dependent way. En route the authors verify and correct a documentation–data mismatch in the benchmark's label semantics, checked three independent ways against released fields. The empirical work is careful: bootstrap and cluster-bootstrap CIs, selection-corrected coupled

Significance. If the results hold, the paper delivers three things the field needs. First, a concrete, quantified demonstration that black-box CoT auditing largely re-detects answer correctness, with a stratified decomposition that explains why: the regime where detection partially works (correct answers) is the lower-stakes one. Second, a well-replicated inversion of the standard step-removal metric, established four ways (full benchmark, released-score recomputation, correct-regime stratification, and an independent constructed testbed with intervention-defined sample-level labels) — a falsifiable, reproducible finding with an immediate prescription (stop scoring detectors against answer-tracing metrics; the ρ=0.87 circularity demonstration alone is useful). Third, a transfer-validated caution that instructed rationalization — a common construction — shares no decodable structure with annotated unfaithfulness. Strengths worth naming: the label-semantics verification against released data (a genuine service to anyone using FaithCoT-Bench, and responsibly disclosed to maintainers), selection-corrected permutation testing with coupled nulls, honest reporting of underpowered cells as "not detected

major comments (3)
  1. [§5, Table 2; §3] The two-regime claim rests entirely on FaithCoT-Bench's human labels, whose validity is established only via internal consistency (95.7% binary/four-way agreement, verified code semantics), not as ground truth. Two confounds follow. (a) Correct regime: annotators judging 'post-hoc rationalization' on correct answers have only surface cues — non-entailing steps, loose reasoning-answer coupling — and the three signals that work there (NLI step-support, answer-tracing, prefix instability) measure nearly those same cues. The 0.63–0.67 AUROCs may partially quantify agreement with the annotators' decision rule rather than detection of unfaithfulness per se. (b) Blind regime: honest-vs-unfaithful error is arguably the hardest annotation call from surface features, so those labels may be the noisiest, which would flatten real behavioral signal toward chance and manufacture the null. Notably, the
  2. [§4, Table 1/Table 2; Appendix G] Two of the three signals that clear chance in the correct regime are reported in inverted directions (answer-tracing, prefix instability), against the statement that 'signal directions are fixed a priori by each family's faithfulness rationale.' For soft_faithfulness the inversion is itself a finding with a proposed mechanism, and it is independently corroborated — that part is sound. But prefix instability is a novel signal defined in this paper (Appendix G); reporting it 'in its empirically discriminative direction' is post-hoc direction selection, and it enters Table 2 as one of the three load-bearing correct-regime positives. Please either specify what a priori direction prefix instability has (and why), or validate its direction out-of-sample (e.g., split-half or per-domain direction confirmation), and separate a-priori-direction from inverted-direction results in the headline claim
  3. [§6, Table 3; Abstract] The 'different models for different regimes' asymmetry rests on one significant cell per model: Llama's incorrect-regime probe (p≤0.03) and Qwen's correct-regime probe (p=0.014), while Llama's correct-regime cell is nominally the largest CV AUROC in the table (0.759) yet n.s. (p=0.19, ft4 n=26). 'Significant in one model, not in the other' is not evidence of a model×regime interaction; with one minority class of n=26 the underpowered reading is at least as plausible. The Limitations section concedes this, but the abstract and §6 framing ('each regime is decodable in a different model') overstate it. Please either add a direct interaction/contrast analysis (e.g., a permutation test on the difference of AUROCs across models within a regime, with shared resampling) or soften the framing throughout to 'decodable in at least one model each, with model-dependence suggestive but unresolved.'
minor comments (6)
  1. [§3, Appendix A; Ethics Statement] All code, cross-tabulations, and the 898 hint-flip + 2,379 genuine traces are 'released upon publication,' so nothing in the pipeline was checkable at review time. Given that the label-semantics correction (§3, Appendix A) is a community-facing claim about another group's release, depositing at least the verification scripts and label_crosstabs.json with the submission would materially strengthen it.
  2. [§3 (Statistical standards); §9] Cluster bootstrap over only 8 model×domain cells is at the edge of what cluster-robust resampling supports; with both headline effects near chance on HLE-Bio (n=40/model), a wild cluster bootstrap or exact per-cell reporting alongside Figure 6 would be more convincing. At minimum, state the small-cluster caveat where the cluster-bootstrap CIs are quoted.
  3. [§4, Table 1; Abstract] Table 1 is restricted to the 633-trace complete-feature subset (2 open models), but the abstract's headline AUROCs (0.696, 0.63–0.67) read as benchmark-wide. The four-model replication for NLI/DAG appears only in §5's text. Please state the subset restriction in the abstract or add the four-model numbers to Table 1.
  4. [§4; Appendix G; Table 4] Prefix instability is used in Tables 1–2 but defined only in Appendix G; the relation to Lanham et al.-style prefix metrics deserves one sentence in §4. Also clarify the '0.005 means p≤1/201' convention where first used, and mark in Table 4 which statistic (nested vs. held-out) is primary.
  5. [Table 5; Appendix C, Figure 4] Table 5 mixes confirmatory (Holm-corrected) and exploratory cells; the Qwen LogiQA-source transfer (p=.046, uncorrected across three targets) is appropriately labeled exploratory in §7.3(2b) but not in the table itself. Add a marker. Also, Figure 4's depth-gradient claim (instructed L9 → hint L17 → annotated L29) is described as descriptive; the peak-location instability caveat should appear in the figure caption, not only the text.
  6. [Abstract; §3] PDF text extraction shows spacing artifacts in the abstract ('are faithful', 'auditing black-box', 'on incorrect answers'); please check the compiled version. The 69% figure (233/340) should state explicitly that the denominator is annotated unfaithfulness benchmark-wide, not the complete-feature subset.

Circularity Check

0 steps flagged

No significant circularity: empirical audit against external human labels, with explicit anti-circular controls.

full rationale

This paper is an empirical detection audit, not a first-principles derivation. Its load-bearing claims are measured AUROCs, regime stratifications, probe held-out scores, and cross-construction transfer statistics evaluated against FaithCoT-Bench’s external human annotations (and, separately, against intervention-defined hint labels). The authors explicitly refuse the common circular trap of scoring detectors against soft_faithfulness—itself a detector—and document that an apparent ρ=0.87 collapses to chance on human labels. Constructions (instructed answer-first; hint-induced flips) are transfer-tested onto held-out annotated regimes rather than only self-detected. Label-semantics verification is done against the release’s own stored parsed answers and gold labels, not by redefining the target. There is no self-definitional loop, no fitted parameter renamed as a prediction, no load-bearing self-citation uniqueness theorem, and no renaming of a known result presented as derivation. Ordinary dependence on one external annotated benchmark is a validity/generalization concern, not circularity of the claimed chain. Score 0; steps empty.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

This is an empirical measurement paper, not a derivation from first principles. Load-bearing commitments are standard statistical detection assumptions, trust in one human-annotated benchmark’s labels after semantic correction, and the operational definitions of behavioral signals and constructed unfaithfulness. No new physical entities; free parameters are ordinary ML analysis choices (PCA dimension, layer selection under correction, thresholds for hint mention filters).

free parameters (4)
  • PCA dimension (50) for linear probes = 50
    Chosen analysis hyperparameter for activation probes; affects decodability estimates though selection-corrected tests mitigate cherry-picking of layers.
  • Best-layer selection for probes = model- and task-dependent (e.g., Llama annotated incorrect ~L29)
    Layer chosen post hoc per task; paper applies selection-corrected permutation nulls, but finite-sample power still depends on this choice.
  • Hint-mention filter / deference keyword patterns = keyword/pattern audit; few exclusions (e.g., 3–4 traces)
    Operational rule deciding which flipped traces count as unverbalized hint-induced unfaithfulness vs honest deference.
  • Sampling temperature and decoding settings for constructions = T=0.7, top-p 0.9, rep-penalty 1.1
    T=0.7, top-p 0.9, rep penalty 1.1 define the sampled baseline/hint distributions and thus which traces enter the hint-induced set.
axioms (5)
  • domain assumption FaithCoT-Bench human binary and four-way faithfulness labels are valid enough ground truth once documentation–data code mapping is corrected to the data-side semantics.
    All annotated AUROC and regime claims are scored against these labels (§3, §4–6).
  • standard math AUROC against human labels with bootstrap/permutation tests is an appropriate measure of detector quality at instance level.
    Standard detection evaluation; class imbalance mild (13–43% minority) as stated.
  • domain assumption Linear decodability of final-position hidden states of the completed trace indicates information present in the representation (not that the model used it during generation).
    Explicitly scoped via Belinkov/Elazar caveats in §2 and §6; teacher-forced re-encoding of stored traces.
  • domain assumption Hint-induced unverbalized answer flips satisfy a counterfactual unfaithfulness criterion at sample level (baseline wrong → hint-correct, hint not mentioned).
    Operationalizes Turpin/Chen-style constructions as labeled testbeds (§7.2); not a problem-level inability claim.
  • ad hoc to paper Behavioral signal directions fixed a priori by each family’s faithfulness rationale; opposite-direction discrimination reported as inverted rather than flipped post hoc without disclosure.
    Analysis policy stated in §4; important for interpreting soft_faithfulness inversion.
invented entities (2)
  • Two regimes of CoT unfaithfulness (incorrect-answer honest vs unfaithful error; correct-answer faithful vs post-hoc) independent evidence
    purpose: Organize detection results and explain why aggregate black-box audits fail and why metrics invert.
    Not a new ontological object in nature, but a postulated partition of the detection problem induced by answer correctness and faithful_type codes; independent handle is whether other benchmarks/models reproduce the split.
  • Hint-induced counterfactual unfaithfulness testbed (sample-level unverbalized flips) independent evidence
    purpose: Scale labeled unfaithfulness beyond scarce human annotations and test transfer into annotated regimes.
    Construction rather than physical entity; falsifiable via transfer, mention filters, and strict-subset resampling checks provided in the paper.

pith-pipeline@v1.2.0-grok45-kimik3 · 22404 in / 3967 out tokens · 75678 ms · 2026-07-30T21:40:25.987411+00:00 · methodology

0 comments
read the original abstract

Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench's human annotations, we find answer correctness structures the problem at every level. Answer incorrectness alone (an oracle diagnostic, not a deployable detector) outperforms every purpose-built signal (AUROC 0.696), because 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness splits detection into two regimes: on correct answers, behavioral signals moderately separate faithful from post-hoc reasoning (0.63-0.67); on incorrect answers, where most unfaithfulness lives, no tested signal is detectably above chance (replicated on all four models for benchmark-wide signals). The standard step-removal metric anti-correlates with human labels; this inversion reproduces on the benchmark's released scores and on hint-dependent counterfactually labeled traces. Linear probes decode the behaviorally blind regime in Llama-3.1-8B and the correct-answer regime in Qwen-2.5-7B, with no shared, positively aligned direction detected across regimes; instructed answer-first traces (7 models) transfer to neither annotated regime, while hint-induced unverbalized answer flips do, in model- and source-dependent settings. We also independently verify and resolve a documentation-data mismatch in the benchmark's label semantics.

Figures

Figures reproduced from arXiv: 2607.23458 by Dikshant Aryal, Nick Rahimi, Suramya R. Angdembay.

Figure 1
Figure 1. Figure 1: A real trace pair from our hint-induced testbed (Llama-3.1-8B, AQuA-RAT). Unhinted, the model answers [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The audit in one view (complete-feature sub [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Instructed construction across 7 models: probe [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Per-layer probe AUROC for the three con [PITH_FULL_IMAGE:figures/full_fig_p011_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: The three-way bridge at a glance (Llama). Bor [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Per-cell (signal × model × domain) AUROC with bootstrap 95% CIs, by regime; NLI is available benchmark-wide (all four models), answer-tracing for the two open models. Cells with n<10 or a single class are omitted; per-cell n is small, so individual CIs are wide. The incorrect-regime cells scatter around chance across models and domains, including the closed models; the correct-regime cells lean above chanc… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

39 extracted references · 12 linked inside Pith

  1. [1]

    2026 , url =

    Shen, Xu and Wang, Song and Tan, Zhen and Yao, Laura and Zhao, Xinyu and Xu, Kaidi and Wang, Xin and Chen, Tianlong , booktitle =. 2026 , url =

  2. [2]

    Premise-Augmented Reasoning Chains Improve Error Identification in Math Reasoning with

    Mukherjee, Sagnik and Chinta, Abhinav and Kim, Takyoung and Sharma, Tarun Anoop and Hakkani-T. Premise-Augmented Reasoning Chains Improve Error Identification in Math Reasoning with. Proceedings of the 42nd International Conference on Machine Learning (ICML) , series =. 2025 , publisher =

  3. [5]

    Graph of Verification: Structured Verification of

    Fang, Jiwei and Zhang, Bin and Wang, Changwei and Wan, Jin and Xu, Zhiwei , journal =. Graph of Verification: Structured Verification of. 2025 , url =

  4. [6]

    2025 , url =

    Feng, Yu and Weir, Nathaniel and Bostrom, Kaj and Bayless, Sam and Cassel, Darion and Chaudhary, Sapana and Kiesl-Reiter, Benjamin and Rangwala, Huzefa , journal =. 2025 , url =

  5. [7]

    2024 , url =

    Asai, Akari and Wu, Zeqiu and Wang, Yizhong and Sil, Avirup and Hajishirzi, Hannaneh , booktitle =. 2024 , url =

  6. [8]

    Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , year =

    A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains , author =. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) , year =

  7. [9]

    2024 , url =

    Zheng, Chujie and others , journal =. 2024 , url =

  8. [10]

    2024 , url =

    Han, Simeng and others , booktitle =. 2024 , url =

  9. [11]

    2026 , url =

    Pham, Hoang and Le, Dong and Luu, Anh Tuan , journal =. 2026 , url =

  10. [12]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  11. [14]

    Transactions on Machine Learning Research , year =

    Chain-of-Thought Unfaithfulness as Disguised Accuracy , author =. Transactions on Machine Learning Research , year =

  12. [15]

    Computational Linguistics , volume =

    Probing Classifiers: Promises, Shortcomings, and Advances , author =. Computational Linguistics , volume =

  13. [16]

    Human Brain Mapping , volume =

    Nonparametric Permutation Tests for Functional Neuroimaging: A Primer with Examples , author =. Human Brain Mapping , volume =

  14. [17]

    Transactions of the Association for Computational Linguistics , volume =

    Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals , author =. Transactions of the Association for Computational Linguistics , volume =

  15. [18]

    Advances in Neural Information Processing Systems (NeurIPS) , year =

    Analysing the Generalisation and Reliability of Steering Vectors , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =

  16. [25]

    ICLR 2026 Workshop , year =

    Probing and Steering Chain-of-Thought Unfaithfulness in Language Models , author =. ICLR 2026 Workshop , year =

  17. [26]

    Iv \'a n Arcuschin, Jett Janiak, Robert Krzyzanowski, Senthooran Rajamanoharan, Neel Nanda, and Arthur Conmy. 2025. Chain-of-thought reasoning in the wild is not always faithful. arXiv preprint arXiv:2503.08679

  18. [27]

    Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207--219

  19. [28]

    Oliver Bentham, Nathan Stringham, and Ana Marasovi \'c . 2024. Chain-of-thought unfaithfulness as disguised accuracy. Transactions on Machine Learning Research. ArXiv:2402.14897

  20. [29]

    Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, et al. 2025. Reasoning models don't always say what they think. arXiv preprint arXiv:2505.05410

  21. [30]

    Kyle Cox, Darius Kianersi, and Adri\`a Garriga-Alonso. 2026. https://arxiv.org/abs/2603.01437 Post-hoc reasoning in chain of thought: Decoding and steering pre-committed answers . arXiv preprint arXiv:2603.01437

  22. [31]

    Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg. 2021. Amnesic probing: Behavioral explanation with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 9:160--175

  23. [32]

    Jiwei Fang, Bin Zhang, Changwei Wang, Jin Wan, and Zhiwei Xu. 2025. https://arxiv.org/abs/2506.12509 Graph of verification: Structured verification of LLM reasoning with directed acyclic graphs . arXiv preprint arXiv:2506.12509

  24. [33]

    Yu Feng, Nathaniel Weir, Kaj Bostrom, Sam Bayless, Darion Cassel, Sapana Chaudhary, Benjamin Kiesl-Reiter, and Huzefa Rangwala. 2025. https://arxiv.org/abs/2511.04662 VeriCoT : Neuro-symbolic chain-of-thought validation via logical consistency checks . arXiv preprint arXiv:2511.04662

  25. [34]

    Yoav Gur-Arieh, Ana Marasovi \'c , and Mor Geva. 2026. Faithfulness metrics don't measure faithfulness: A meta-evaluation with ground truth. arXiv preprint arXiv:2605.25052

  26. [35]

    Alon Jacovi et al. 2024. https://aclanthology.org/2024.acl-long.254/ A chain-of-thought is as strong as its weakest link: A benchmark for verifiers of reasoning chains . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL)

  27. [36]

    Nathalie Kirch, Samuel Dower, Adrians Skapars, Helen Yannakoudakis, Ekdeep Singh Lubana, and Dmitrii Krasheninnikov. 2025. The impact of off-policy training data on probe generalisation. arXiv preprint arXiv:2511.17408

  28. [37]

    Tamera Lanham et al. 2023. https://arxiv.org/abs/2307.13702 Measuring faithfulness in chain-of-thought reasoning . arXiv preprint arXiv:2307.13702

  29. [38]

    Parsa Mirtaheri and Mikhail Belkin. 2026. Catching rationalization in the act: Detecting motivated reasoning before and after cot via activation probing. arXiv preprint arXiv:2603.17199

  30. [39]

    Sagnik Mukherjee, Abhinav Chinta, Takyoung Kim, Tarun Anoop Sharma, and Dilek Hakkani-T \"u r. 2025. https://proceedings.mlr.press/v267/mukherjee25a.html Premise-augmented reasoning chains improve error identification in math reasoning with LLMs . In Proceedings of the 42nd International Conference on Machine Learning (ICML), volume 267 of Proceedings of ...

  31. [40]

    Nichols and Andrew P

    Thomas E. Nichols and Andrew P. Holmes. 2002. Nonparametric permutation tests for functional neuroimaging: A primer with examples. Human Brain Mapping, 15(1):1--25

  32. [41]

    Giovanni Maria Occhipinti, Alessandro Abate, and Nandi Schoots. 2026. Probing and steering chain-of-thought unfaithfulness in language models. In ICLR 2026 Workshop. OpenReview LocRunEIxK

  33. [42]

    Hoang Pham, Dong Le, and Anh Tuan Luu. 2026. https://arxiv.org/abs/2606.16151 GRACE : Step-level benchmark for faithful reasoning over context . arXiv preprint arXiv:2606.16151

  34. [43]

    Xu Shen, Zhen Tan, Song Wang, Pingjun Hong, Rui Miao, Xin Wang, and Tianlong Chen. 2026 a . Detecting unfaithful chain-of-thought via circuit-guided internal-external discrepancy. arXiv preprint arXiv:2605.25603

  35. [44]

    Xu Shen, Song Wang, Zhen Tan, Laura Yao, Xinyu Zhao, Kaidi Xu, Xin Wang, and Tianlong Chen. 2026 b . https://openreview.net/forum?id=lN3yKqqzF1 FaithCoT-Bench : Benchmarking instance-level faithfulness of chain-of-thought reasoning . In The Fourteenth International Conference on Learning Representations (ICLR)

  36. [45]

    Daniel Tan, David Chanin, Aengus Lynch, Dimitrios Kanoulas, Brooks Paige, Adri \`a Garriga-Alonso, and Robert Kirk. 2024. Analysing the generalisation and reliability of steering vectors. In Advances in Neural Information Processing Systems (NeurIPS). ArXiv:2407.12404

  37. [46]

    Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman. 2023. Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (NeurIPS). ArXiv:2305.04388

  38. [47]

    Kerem Zaman and Shashank Srivastava. 2025. Is chain-of-thought really not explainability? chain-of-thought can be faithful without hint verbalization. arXiv preprint arXiv:2512.23032

  39. [48]

    Chujie Zheng et al. 2024. https://arxiv.org/abs/2412.06559 ProcessBench : Identifying process errors in mathematical reasoning . arXiv preprint arXiv:2412.06559