Pith. sign in

Paper Citation Record · LEDGER

Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2307.08678.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.08678 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:14:41.316054Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 98f75df4-e94b-4224-aa3e-00aa9fcf32a8 · inbound

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing cites this paper.

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-12T11:42:30.886515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:42:30.886515Z digest=sha256:e0a136f5fed7fdeddd7f67449a998e7b3b27e049dc6f0183e3adc414990d51ef

Observation 44df6470-a2a8-4fa4-80fa-1949b9dd7150 · inbound

Let your LLM generate a few tokens and you will reduce the need for retrieval cites this paper.

Let your LLM generate a few tokens and you will reduce the need for retrieval Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:54:56.985524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:54:56.985524Z digest=sha256:717db1674332c9dabe5343130d61b99f1b8a3323d540236d1a8f6777e667c8d8

Observation 29df79b0-d079-4839-bd57-ba240b292bfa · inbound

The Science of Evaluating Foundation Models cites this paper.

The Science of Evaluating Foundation Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.646158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.646158Z digest=sha256:3737877fa40603c6349f25386c3e3e2f4b57b2f1530d016f7db0a6528de61762

Observation d5c304b5-0ab3-438d-ad9b-0814d8cde3eb · inbound

Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models cites this paper.

Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:41.316054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:41.316054Z digest=sha256:f2576d839c8c37c6e7c62d4951c9ae52adea78a950b9c3e82f4bb134a40fd5d2

Observation 54d38c93-7ef8-4d85-8794-6adf0cf2c45a · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.721948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.721948Z digest=sha256:617ab984c374f203f8977d4cf412249c601d41cc5e4652214cbf01dcde9a5513

Observation 9cad1910-5b6a-4cc6-ab7e-72aca740de40 · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:43.957104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:43.957104Z digest=sha256:e5c0ee00f1a291d12bf4465192ad9acc19139c4dc28dbbb8fd4c201952fb155c

Observation e8eceb94-7d19-431f-ab49-a252e88c46d4 · inbound

From Features to Actions: Explainability in Traditional and Agentic AI Systems cites this paper.

From Features to Actions: Explainability in Traditional and Agentic AI Systems Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:14.955808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:14.955808Z digest=sha256:7c5fff24091b40b76fbf343aa95c41c402a0b166bb0d9e3f7af68bae31e12ddc

Observation a60abfb8-b0aa-4f5a-b5e0-a0bf1a849545 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:07.946422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:07.946422Z digest=sha256:7d82528c5dc7bddbc96b2e7bae0777cc3e5cd052220d030040c068007142c2c0

Observation c8f91e2e-1fb2-40cc-a98f-2f7011873347 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 110

Resolution
verified exact
local_arxiv, observed 2026-08-01T08:38:35.965866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-01T08:36:27.791922Z digest=sha256:017e155c575f13296303b72d26d30d8211ba10d0ab4404d3c13e0e2ed2d67004