Pith. sign in

Paper Citation Record · LEDGER

Measuring Faithfulness in Chain-of-Thought Reasoning

As of 23 July 2026, this Paper Citation Record lists 26 of 26 outbound references and 100 inbound Pith citation observations for arXiv:2307.13702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.13702 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T20:51:38.390234Z

measured 126 of 126 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-23T06:31:01.910684+00:00

measured 100 of 180 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-15T11:31:26.851193Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T00:37:43.053081Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact8
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch13

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c3e0558-329d-4c7c-b10d-de80518435de · outbound

This paper cites Language models as agent models.

Measuring Faithfulness in Chain-of-Thought Reasoning Language models as agent models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T20:51:40.534792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:966b43ecbfde9ea253920e675d5b9a60707d712173314c53c0fbccb898c38097

Observation ed0c44c0-072d-4a14-8996-c1787ed64d9f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Measuring Faithfulness in Chain-of-Thought Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:51:39.588347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:b6c96ef7b5a5326087e938cb530c6bf4a85f8b456fbed111c6dd52af649a837f

Observation 310a7225-23d5-403c-b16f-07a60dbc2a9a · outbound

This paper cites Measuring Progress on Scalable Oversight for Large Language Models.

Measuring Faithfulness in Chain-of-Thought Reasoning Measuring Progress on Scalable Oversight for Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:01:41.422297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:de3e1aa0ae0133ab60b5e8ca4572640059c0e8c35e3c57f94687a28e9f8fa850

Observation 42f3679c-a503-42ca-aa6b-0899b98e3555 · outbound

This paper cites Language Models are Few-Shot Learners.

Measuring Faithfulness in Chain-of-Thought Reasoning Language Models are Few-Shot Learners

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:51:39.665457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:c26e19e7d5c2b727222bbe95754cfb1a6bc2e7a8f0f5f960ec6a90a8e903829d

Observation 77362fa4-78fa-4bed-a2a3-f43687b01787 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Measuring Faithfulness in Chain-of-Thought Reasoning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:51:39.784832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:ce9485da73a9c38e5e2c40eef7c1d86c15e763be584cf8d274345653a62dc018

Observation 36cec669-0e30-408e-b77f-a4540d8d5eee · outbound

This paper cites Faithful Reasoning Using Large Language Models.

Measuring Faithfulness in Chain-of-Thought Reasoning Faithful Reasoning Using Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:39.889349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:bc7a4cea44b281172b772fe05534b855fb9dfeae11acf5e90fd69f36b2f81515

Observation fb721806-0646-4dc9-bd9e-d92265480f8a · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Measuring Faithfulness in Chain-of-Thought Reasoning Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:45.301651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:6d92c8da14cfe4dc8998c6658870a84b36df4ed842d180756094344f344f7446

Observation d4e529b7-9eed-48c8-a95e-99f098f8c2d3 · outbound

This paper cites Successive prompting for decomposing complex questions.

Measuring Faithfulness in Chain-of-Thought Reasoning Successive prompting for decomposing complex questions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T20:51:40.464859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:48631d3c148c95d9272331cced4df80b84f075e4124b3f6cc2ce3cd711e4309f

Observation af2f82d5-f17d-4fbb-a365-01ce01d58a69 · outbound

This paper cites URL https: //aclanthology.org/2022.emnlp-main.81.

Measuring Faithfulness in Chain-of-Thought Reasoning URL https: //aclanthology.org/2022.emnlp-main.81

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T20:51:40.354755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:3f82120ac7bba1aa241aa57a163371fbdeeaa749f7e45f1636fde8c6bb0f0d52

Observation e40fb9df-7487-4bcd-8245-9f109f257360 · outbound

This paper cites URL https: //www.science.org/doi/abs/10.1126/sc irobotics.aay7120.

Measuring Faithfulness in Chain-of-Thought Reasoning URL https: //www.science.org/doi/abs/10.1126/sc irobotics.aay7120

Reference 10

Resolution
verified exact
doi, observed 2026-05-11T20:51:39.016749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:bd4930600b15bcd55d9dc130c22f54109d1845bdb7b403b98dcac78b7c1687b2

Observation d2dd1ed0-0477-45e1-98b7-40dcec18eb05 · outbound

This paper cites What do we need to build explainable AI systems for the medical domain?.

Measuring Faithfulness in Chain-of-Thought Reasoning What do we need to build explainable AI systems for the medical domain?

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:39.992068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:db4a7ca52c2a6f537304402aac2473cd21db6fbb25d4f360c88a389d3469ed4c

Observation 3ffba062-af03-495a-837b-fa6ec45e4d69 · outbound

This paper cites Jacovi and Y.

Measuring Faithfulness in Chain-of-Thought Reasoning Jacovi and Y

Reference 12

Resolution
metadata mismatch
doi, observed 2026-05-11T20:51:38.564923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:d5a34c44b4338829f20641845d196413a1ade5920f7d00dc8f4d2acdefcd2d0c

Observation 0f986712-a651-420d-bafc-d8cc13a3dd6f · outbound

This paper cites Explanations from Large Language Models Make Small Reasoners Better.

Measuring Faithfulness in Chain-of-Thought Reasoning Explanations from Large Language Models Make Small Reasoners Better

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:51:39.295774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:518cd1cbe51f5e8dd94442ed20c2dfca86b10ad957db85905cfec5b4d23d60c6

Observation e1e93a3e-cbca-4efc-9096-4dfc2eaa1e27 · outbound

This paper cites URLhttps://doi.org/10.18653/v1/2022.acl-long.229.

Measuring Faithfulness in Chain-of-Thought Reasoning URLhttps://doi.org/10.18653/v1/2022.acl-long.229

Reference 14

Resolution
verified exact
doi, observed 2026-05-11T20:51:38.694637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:0a24a9d58409f4b7e3fc1a5db0fcf88d587f4f3ee5de90078ab531c098bf458e

Observation 5c609946-0b58-480c-b7a4-ba8f1a5fd35d · outbound

This paper cites Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems.

Measuring Faithfulness in Chain-of-Thought Reasoning Program Induction by Rationale Generation: Learning to Solve and Explain Algebraic Word Problems

Reference 15

Resolution
metadata mismatch
doi, observed 2026-05-11T20:51:38.797352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:993f7ff0a440f53f1257741e1765894caaf05418995a44a0cfbca1ca120de05d

Observation ea21a291-bede-4ecb-ace3-6b8f8d6e5a0c · outbound

This paper cites Logiqa: A challenge dataset for machine reading comprehension with logical reasoning.

Measuring Faithfulness in Chain-of-Thought Reasoning Logiqa: A challenge dataset for machine reading comprehension with logical reasoning

Reference 16

Resolution
metadata mismatch
doi, observed 2026-05-11T20:51:38.887808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:68dd7fb5e15dab5717984e347cc711226589fa6b96a4846b0838b0059c658d09

Observation fcb165cf-84cd-46db-ab64-6d9d8b811db1 · outbound

This paper cites Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango.

Measuring Faithfulness in Chain-of-Thought Reasoning Text and Patterns: For Effective Chain of Thought, It Takes Two to Tango

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:39.389361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:b41d212dd06c1d9843d57f86e1ad0a1c62b39d506564d40e5b3067694bb164c1

Observation 03ed4b30-b1fd-4684-b6b9-f46839a06416 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Measuring Faithfulness in Chain-of-Thought Reasoning Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T20:51:40.274730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:8774b948ad442e56dc4730809427503182b2f28ea97e17430de29a0c3b1a3505

Observation 3789a7bc-0c9b-47ec-9d70-e7d5ee5f5d9e · outbound

This paper cites doi: 10.18653/v1/D18-.

Measuring Faithfulness in Chain-of-Thought Reasoning doi: 10.18653/v1/D18-

Reference 19

Resolution
malformed identifier
doi_truncated, observed 2026-05-11T20:51:39.147604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:624ed61672d91bef55f3eca78992fa9b242b345209d3ea468fb630c59bce0dd0

Observation e96e1bb5-a7f1-42c2-a4c1-a9ebea700ed0 · outbound

This paper cites Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead.

Measuring Faithfulness in Chain-of-Thought Reasoning Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead

Reference 20

Resolution
metadata mismatch
doi, observed 2026-05-11T20:51:38.627737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:0cc41e0fe8e0c7b4cdfc7eef38c0ba5628453dc593ad66cce77c85ca2b1b3490

Observation 216a096d-5215-4c86-9e01-548008079792 · outbound

This paper cites Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

Measuring Faithfulness in Chain-of-Thought Reasoning Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:01:19.882845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:a9a5e08a551157573bfe16ca3f7a3ad526f2ff9d0073790cb3a6e1b37a2a4b29

Observation caab02dd-a9cf-48b4-b704-53104a30a3b5 · outbound

This paper cites Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models.

Measuring Faithfulness in Chain-of-Thought Reasoning Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:43:46.984626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:ff547d8db6233a14ba528a6de090997fa3c433a15a09d3bc0608e1c786fb273c

Observation 0aa03b69-680b-47c3-85d0-94f7496fbdc1 · outbound

This paper cites Rationale-Augmented Ensembles in Language Models.

Measuring Faithfulness in Chain-of-Thought Reasoning Rationale-Augmented Ensembles in Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:40.105665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:d251cefe5fdafdbcd0be0998efb82b37f8db2fb50940aa2080ed45958e01bb83

Observation 9b16ba66-0ca5-4dc3-b6c0-7001266737d9 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Measuring Faithfulness in Chain-of-Thought Reasoning Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:51:40.196870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:06f86f84868dd7c114c2f27fb2f896c120a127a81cb6564c09d014ebb2a75c10

Observation ca50e230-0241-45d5-b244-545a6f76d50d · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

Measuring Faithfulness in Chain-of-Thought Reasoning URL https:// doi.org/10.18653/v1/p19-1472

Reference 25

Resolution
verified exact
doi, observed 2026-05-11T20:51:38.471882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:22eb687b98b9fe9ef6fed3f4dae3b79eb20eb8b44a6ba687a1f41946430707f2

Observation e9e57f8b-ce17-4b67-8cb6-26ac124ece5f · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Measuring Faithfulness in Chain-of-Thought Reasoning Fine-Tuning Language Models from Human Preferences

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:51:39.464894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T20:51:38.390234Z digest=sha256:52636bbee4c88e31cac56413a372d63d785f3e8c5250e106443ee8679a86a536

Pith citing papers

Observation aba4b3ea-1a09-4a71-8543-8dea535c0a13 · inbound

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions cites this paper.

A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 163

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:46:27.325938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T02:46:26.957539Z digest=sha256:4f6d8e5ec839a726d6f7a5b1e9abe7b4beccd0dd4005aea9f81d5be20ca4dbab

Observation 18f2c11e-d450-48c0-a9b0-48e6926c7bba · inbound

Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs cites this paper.

Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:52:02.536551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-18T00:52:02.421389Z digest=sha256:4d5c5f036c28bc3e339d377a13af3689a322bd494f359e35f904c07dbea8f7f2

Observation 5851f7f6-f5c2-4c93-ab7c-1ed8b0d0c431 · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:22:01.663013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:78dde5d42db8ba745fd94a0cc962884b660517af2c3491943f4085c93faadc7e

Observation cd134cf4-e6f8-4304-ac66-49af03c7cc62 · inbound

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations cites this paper.

Compressed Chain of Thought: Efficient Reasoning Through Dense Representations Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T04:47:40.416837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-17T04:47:40.335477Z digest=sha256:8af2759e0bc6c11875ee0c9c266f3e57044c0c5ce139062105435a9b9b8fbcec

Observation fd8ae6bc-7013-42c1-b2d9-fddb5c5d0ab7 · inbound

OpenAI o1 System Card cites this paper.

OpenAI o1 System Card Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:42:39.679765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-23T06:39:44.542350Z digest=sha256:2562414027dccc8101d7daa8c97d67e847bc32e09986259633578bd4da8eb542

Observation d14e9e89-2584-4b06-a015-94d3fd653eea · inbound

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation cites this paper.

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T07:24:12.943511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-21T07:24:12.845841Z digest=sha256:16d0ca6a3428f8d059fc39ccac4e4a609f8252cb0d212443360393c0b2303426

Observation b9d98a70-2587-4a2e-beee-1d2fd6814b44 · inbound

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens cites this paper.

Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T01:21:58.115138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-19T01:18:31.661827Z digest=sha256:f31e35edf0cb440857f613ba8382287aca5d866afbcd71b8f4c39a374db630b2

Observation 1f1ed01d-e616-4c8c-95f3-1d2634c8e406 · inbound

Do Activation Verbalization Methods Convey Privileged Information? cites this paper.

Do Activation Verbalization Methods Convey Privileged Information? Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T15:42:42.176623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-18T15:41:46.771905Z digest=sha256:3fdbc676e521e685a62f8e92aff63dd68a623f5810865764e23ca660c81886e5

Observation bc544b31-60eb-4c70-87c9-7faa7fd02fbf · inbound

On the Reasoning Abilities of Masked Diffusion Language Models cites this paper.

On the Reasoning Abilities of Masked Diffusion Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:56:01.524577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-18T06:54:37.021287Z digest=sha256:4b152b494b64aaf05d39ce204d9c5a0d0c6597c5c1eb0c3fd39cc30c8567df08

Observation eb3391ef-9335-48b4-92a9-0631e664be8a · inbound

AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning cites this paper.

AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:25:58.670854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e3f19f2481442fabb8791a143affd1cabfd9fcad585d7be173c7de10819d3a98

Observation 87df96fe-1941-4bcd-ac8a-1aebc393bd8a · inbound

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought cites this paper.

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-18T02:45:46.166458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-18T02:44:48.729794Z digest=sha256:ba657a78cebc3b8f62aa1431f2c99582dff4b35e593df9de82838a27e359d869

Observation e44d9eae-e7ca-46e6-8140-7cf1e8565be7 · inbound

Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction cites this paper.

Implicature in Interaction: Understanding Implicature Improves Alignment in Human-LLM Interaction Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:00:48.265638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-18T02:58:26.908673Z digest=sha256:87a31ebdbd9b5a793fa29d29b6ae0216a1c0fbe32bd4df45567da3cda94dc993

Observation 6b75f5c5-8fc9-41bb-ac1f-6b6ab0ec9660 · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T20:50:15.230417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-17T20:45:15.034196Z digest=sha256:b0712b1119f1279c61638d02ba41aa6986d1c0ce4a6beed593e8af428f16b0a1

Observation afde311a-65c5-44f6-b0b6-2a5c23048158 · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T07:20:28.800415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-25T07:17:26.381232Z digest=sha256:d1ffcc4bb9e5a87a9dcee40c6e8a3d1b15ee9dee6ed7ca89fabbf22c42e36dd7

Observation 0e404c5f-e040-452d-aa46-07da2fcb06bd · inbound

Training Language Models to Use Prolog as a Tool cites this paper.

Training Language Models to Use Prolog as a Tool Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:18:48.415038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-17T01:15:43.096577Z digest=sha256:656fc550dab1ecb094a270838f8cd60d1fe3b5bacb4cc453b808d4f3aabd1572

Observation d67d0cc8-e759-4601-a4d4-1a63e4b64206 · inbound

Do LLMs Encode Functional Importance of Reasoning Tokens? cites this paper.

Do LLMs Encode Functional Importance of Reasoning Tokens? Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T17:08:08.199173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-16T17:07:33.113459Z digest=sha256:26440cc86571626d5c9b17b062e976bf1e4a703e2c7c7e59c8724af81e281888

Observation cb918adb-2ab5-4c03-9386-9e01e190fcaf · inbound

Emergent Manifold Separability during Reasoning in Large Language Models cites this paper.

Emergent Manifold Separability during Reasoning in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:16:35.023235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T20:12:05.809065Z digest=sha256:f98c4817f63352bcdc6200012e24469ab4f4cd2439b0d52ea063ae8e1a48f7db

Observation 564d2891-2522-4d02-ba55-c7e7775478a7 · inbound

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness cites this paper.

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:19:37.294342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T01:18:48.798851Z digest=sha256:078a688321b9c03b5723c03500b6d6e9af32b1c4ef92cb2a6e8717997ed45960

Observation 4af4b4f8-9f97-456a-97e2-dde7aa2bba73 · inbound

LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset cites this paper.

LongTail Driving Scenarios with Reasoning Traces: The KITScenes LongTail Dataset Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:33:23.329178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T00:30:56.156113Z digest=sha256:0493ee3b1faea88791f67f71f5ce79e263f4a5bb9cd74b5ca9cd26a1663ba8e7

Observation 7c760d15-b8e7-4b34-b174-059e21028a23 · inbound

WMF-AM: Probing LLM Working Memory via Depth-Parameterized Cumulative State Tracking cites this paper.

WMF-AM: Probing LLM Working Memory via Depth-Parameterized Cumulative State Tracking Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T22:03:03.088799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-14T22:01:15.475633Z digest=sha256:324c5cd66236d67f6da8ee2cc327c84243758857934eef597f155f5aecc5c275

Observation b170f153-aaaa-436f-8ba5-c16c61734b54 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 99

Resolution
malformed identifier
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:038d616973f9cd2d4fdb181cfd30dbda45a243d5ab7ac4e938df48514f238e9e

Observation 81d51213-3ad6-4240-b321-3f13ff510d62 · inbound

From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment cites this paper.

From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T19:20:45.106436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T19:17:17.979115Z digest=sha256:71319715cdecf0d3ccdf9b8261bc3576569a73d0955da7d9527c34c69f2caf9f

Observation b5b10c18-a722-4b5b-9cba-80b7b8db5fa6 · inbound

Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space cites this paper.

Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:35:31.803014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-15T11:31:38.301306Z digest=sha256:7184146da90fe3c9d53e101733bb2f923d6859d402f56aa85b36e7ef5b2e33ba

Observation f8eaaee1-2398-4186-9353-b85434c0fc43 · inbound

CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation cites this paper.

CUE-R: Beyond the Final Answer in Retrieval-Augmented Generation Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:05:51.388070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T19:20:12.655900Z digest=sha256:f24be6f372c55da769fb16635dab196445fc6188837b1f78df2961ddbc941cfb

Observation 26c8ad3b-0bdb-421d-b79a-5aabcf73c63a · inbound

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models cites this paper.

Act or Escalate? Evaluating Escalation Behavior in Automation with Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:28:26.275369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T23:25:56.299371Z digest=sha256:11c8ad8dbd3e66a9f4b28d3f450f0a5a38d3773bd418a3d8d7de1220cceccb0c

Observation ab7b35be-9be7-4c64-9d79-c8c0ce0e5325 · inbound

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning cites this paper.

FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:04.746109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T15:04:23.688239Z digest=sha256:9c3bc49996bd64c04172f4e490957b31572bb11993cc1d8a55e61e39aadf9016

Observation e0b9da85-83c3-4a12-b658-d432b0898801 · inbound

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning cites this paper.

From Answers to Arguments: Toward Trustworthy Clinical Diagnostic Reasoning with Toulmin-Guided Curriculum Goal-Conditioned Learning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:45:59.470167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T16:30:34.434984Z digest=sha256:927cea3e78d41713fcb540b070c74d4120706e91c204b75f15a678146fbc112a

Observation 60de5e31-f21b-4c03-b2bc-2f4b72015464 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.737021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:0ad3601bc5aa764543d9fe9b8d7415585e09693cc87fdedcbd51ab9ac02e3485

Observation 87f611cd-8187-47f7-8302-50fc0ef08e66 · inbound

Mamba-SSM with LLM Reasoning for Feature Selection: Faithfulness-Aware Biomarker Discovery cites this paper.

Mamba-SSM with LLM Reasoning for Feature Selection: Faithfulness-Aware Biomarker Discovery Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:30:18.694587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T11:28:04.724476Z digest=sha256:bb5bfdb05327b5cda1c8d2ccb8ea1d4677e4ae9b46d5bb57de5562ac5549f110

Observation 8d8efdad-6c19-4a8a-9af2-7d283590dde6 · inbound

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models cites this paper.

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:00:04.054341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-10T10:58:30.370708Z digest=sha256:ae5b963254741fdd45c159f6b4a696e5497daace3d33192fea51e894ce3a87be

Observation 0721e63e-7504-4777-99f6-e70dcb41ba58 · inbound

LLM Reasoning Is Latent, Not the Chain of Thought cites this paper.

LLM Reasoning Is Latent, Not the Chain of Thought Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:04.684782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T08:49:05.178087Z digest=sha256:67f714adeccf5f105c42ee8eac1d5e72b82b1966763374e2618a856916c48775

Observation f6e18d1a-161a-4b4f-97ee-4eac71a51f4f · inbound

Disambiguating electrical detection of magnetization dynamics in magnetic insulators cites this paper.

Disambiguating electrical detection of magnetization dynamics in magnetic insulators Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-15T11:31:26.851193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:31:26.851193Z digest=sha256:e7e05103cf01fc0be8cc8eb753478b746540174d0b71e47d9443093fe1cc3824

Observation a16104a0-3e80-414c-9fd4-24c3f9f3f839 · inbound

MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition cites this paper.

MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:53:03.529009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T08:49:05.452010Z digest=sha256:2baca29a7ad152e2abc49b2a9717fa8e3fa8f94c6c165c0a832a98357d82fd49

Observation 25b07929-7bb7-4673-a367-d581912f7c84 · inbound

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency cites this paper.

AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.704156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T08:08:50.330857Z digest=sha256:9caca7aa85f03cb0bd1be68a8b37525603eaf701326565c5cb40cedce807b6df

Observation 7c46f2c2-848b-468f-8532-498c66be7c42 · inbound

Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs cites this paper.

Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:48:25.079532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-15T00:43:38.513913Z digest=sha256:149deb706cb19337577c845874a1f05ff5cb7e726670f80a70df9dbdf6d8fa5a

Observation a17569e0-1859-421d-999f-6808001f2f36 · inbound

Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI cites this paper.

Escaping the Agreement Trap: Defensibility Signals for Evaluating Rule-Governed AI Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.535953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T00:19:27.905893Z digest=sha256:c3381ec3d299207f8e47a8587c4c99260147b4b5f4baf97d634f7e3bc3cc7798

Observation f5343d46-4208-47d7-afc4-0a4356dee58c · inbound

Large Language Models Decide Early and Explain Later cites this paper.

Large Language Models Decide Early and Explain Later Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:09.351795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T12:06:33.200365Z digest=sha256:0afc0225bb1c9ba8f6cc20ce6b7ced63714118ca0f8b5f95c3421f8556e1700b

Observation 1befc306-73ae-4f3a-9a28-719e3a170ca3 · inbound

Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought cites this paper.

Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:07.340462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-08T11:52:00.969949Z digest=sha256:9d4857ad7de7311a5c06b91f4638d00bdfad886ce2eff7160e595a9c7e8c94f5

Observation a96fec87-3cbd-49ae-8271-f38e3686b98d · inbound

VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs cites this paper.

VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:46:13.361153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T08:04:28.925895Z digest=sha256:4dc2fa661707b0ccae6b53c6674f4d4dc92c0ebb70efdf234af0ed94c25affa0

Observation e2580677-2470-4f3c-a050-56e761fe6bb3 · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:56:24.795220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:6aeba9d6470020f7e9fef3de5a30dbdd4ae1462186b25d284fd4e3142323644d

Observation fe1ae87b-2552-4231-886c-777dd0c47f69 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:16:16.870590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:5494a04f644a69a536228519e7922103923350b62258868f2e8f23921d01f1d3

Observation 56fd84f4-db2a-4094-949d-139391522620 · inbound

Analyzing LLM Reasoning to Uncover Mental Health Stigma cites this paper.

Analyzing LLM Reasoning to Uncover Mental Health Stigma Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:13.266803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T03:19:25.314797Z digest=sha256:1fd8898fcdb826a3e59d7b70df4bc9f22996b531cb90a0a1c69e7cfd194ed60b

Observation 79f754a7-f661-41f5-af97-8b77fa20f2df · inbound

Knowledge Distillation Must Account for What It Loses cites this paper.

Knowledge Distillation Must Account for What It Loses Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T23:26:18.035472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-07T16:58:41.449172Z digest=sha256:cc9ffc8f67c14dc4712309818b1c677df825575a80e88dd6e47184146afce8e8

Observation bba5edc7-ddca-4570-920c-dac398fb4e68 · inbound

Knowledge Distillation Must Account for What It Loses cites this paper.

Knowledge Distillation Must Account for What It Loses Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:01:13.619368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T03:31:56.787201Z digest=sha256:9f425226d023fe3b35b274c5e649abf0748b4da4344664220d5a7f7bc1c1aebc

Observation c24b04f8-2163-4b3d-81b7-26cb7f0dd357 · inbound

Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models cites this paper.

Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:23:27.048663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T23:18:34.388467Z digest=sha256:8b4272d311a69279156394791f0dc118b578c1c870ce5cd3c099030d2bfb4853

Observation 4cd7c7a9-f377-4f1e-a55e-49cb265cbdeb · inbound

TRUST: A Framework for Decentralized AI Service v.0.1 cites this paper.

TRUST: A Framework for Decentralized AI Service v.0.1 Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:01:28.779069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-07T08:22:14.239443Z digest=sha256:7d57ffbc085d35f3ce2b9b6c066d73d2b5e60ad8e81efc26e52bfdb81759980d

Observation 3be2bbe0-763a-45cf-a3a6-0350b6c1a13f · inbound

Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models cites this paper.

Compliance versus Sensibility: On the Reasoning Controllability in Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-12T09:51:29.001424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-07T08:57:05.520335Z digest=sha256:6ca127d744ddaeb3bd79e4b259e0541a79a74c0e2a9d4b41db878bba716936f9

Observation ba615150-76da-4775-8a6d-10904e26dcef · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T19:05:10.564047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:4629b57c9b777ce63219fc6d36fc6f2a83c24d4b5286b8a1e6f39d092f5b2df1

Observation 6c010835-f72c-4cc2-b8f9-e622b38551a1 · inbound

LLMs Should Not Yet Be Credited with Decision Explanation cites this paper.

LLMs Should Not Yet Be Credited with Decision Explanation Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:33.415201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-09T18:41:24.552939Z digest=sha256:e52df12ed1a832682a9d4c53ad27cee7ea24359949b06dda738fc6eca5ff72f8

Observation 1230fe8e-5e98-4835-983f-56ba238a1b80 · inbound

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles cites this paper.

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:04.443272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-10T14:54:08.156511Z digest=sha256:9aac24333a8401dec900639af5fe8df506e1c82508fb664d7c19a0d4f7012a51

Observation 7eb559aa-70e6-40cc-a7b8-2cad07c9e15d · inbound

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles cites this paper.

NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:59:49.269139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T06:56:44.316029Z digest=sha256:dfb477ed9c8126b9d0e893f356df44f57397bd558a98dce2b02f41d7581a551f

Observation df4f4a7f-91f0-4203-a3c2-3c2c0bc91048 · inbound

Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective cites this paper.

Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:29.475339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T19:25:42.362245Z digest=sha256:52be294f8d9f350f3894cbc623d5248c49012733f1962799276967b52ed36159

Observation ceda3898-0d96-4113-8b3b-7a986f22698b · inbound

Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective cites this paper.

Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T00:15:09.066564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-01T00:12:54.321254Z digest=sha256:d917b97eb918cb2ef55ec1be5258439927e9db185158c6c896679dd6ce41379c

Observation 75257ba1-242f-4b8f-834d-db16f45fe96b · inbound

AgenticPosesRanker: An Agentic AI Framework for Physically Grounded Ranking of Protein-Ligand Docking Poses cites this paper.

AgenticPosesRanker: An Agentic AI Framework for Physically Grounded Ranking of Protein-Ligand Docking Poses Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:07.757075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-09T15:34:59.380508Z digest=sha256:a103c662b82ae190324103f7bb6f28c4a0b47133b4688c9dc25d5a593f8cec80

Observation 2cae353b-024d-4169-8776-f7ff64da7568 · inbound

Understanding Annotator Safety Policy with Interpretability cites this paper.

Understanding Annotator Safety Policy with Interpretability Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:21:08.586519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-08T17:41:00.933594Z digest=sha256:a22364a7635163cbc6e401679f19f38aab6396de75c4d1ff3cc7e4c82a278329

Observation bca19c7b-a35e-4574-8d32-de6d0a80e85b · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:10.434347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:4ed3924986356bf83d79777799184ffff81aaa9e6ee5f1149293a16e07434ed4

Observation 5c31ac8f-fc7f-4688-a675-2438f828919b · inbound

Evaluation Awareness in Language Models Has Limited Effect on Behaviour cites this paper.

Evaluation Awareness in Language Models Has Limited Effect on Behaviour Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.237243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T11:00:54.568772Z digest=sha256:ba4c05be4113d2b1ff10c8ba158cf8be7448c8a7009a05f5c83cb96511f05418

Observation 660311e6-73b4-4a10-bdc8-32f3674c4244 · inbound

Measuring Black-Box Confidence via Reasoning Trajectories: Geometry, Coverage, and Verbalization cites this paper.

Measuring Black-Box Confidence via Reasoning Trajectories: Geometry, Coverage, and Verbalization Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:10.002120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-08T10:08:02.008581Z digest=sha256:b46aa4d894fb0730f653c8192cca6a7842ffacae7736584149c269b1fbe704ab

Observation abbc8eea-1d17-469d-ac9f-b482e4cfd3bb · inbound

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning cites this paper.

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:05:56.696347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-11T00:52:59.406190Z digest=sha256:1460ad1714d37783d6d1905a338e1d9a1cc0110b6dee8332eef91329ca08f71e

Observation bf8d002a-6941-40af-abd8-092da4a505d8 · inbound

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning cites this paper.

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:41:28.746209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T02:23:55.323250Z digest=sha256:6345ea74287fa4401bee0e79aa9b7b2c839f5055eb2793d2360cb88c85e0deb9

Observation 69ddab04-b3bb-428d-bce5-40ca4c877d0b · inbound

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning cites this paper.

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:22:23.211873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T06:21:00.353334Z digest=sha256:22a3a0e970edb31c25f8a87ac09e0fa396b9adee4d87be85f23667c533306a09

Observation e1e4e93a-2949-470e-9566-da6a793c6b84 · inbound

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning cites this paper.

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:59:27.813479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-14T20:55:31.770238Z digest=sha256:7669cebcb19e7546eaa2bc49c3fa84e44835d318d29913b716be89ba434e3677

Observation 3395586b-ba5f-4a76-bbdc-3627d3746ed6 · inbound

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning cites this paper.

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:00:23.408488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-25T05:57:58.487109Z digest=sha256:ece0233ef7b4bbfd4b5fd593e2c8225979f7bf3464aac6e9b4d4cba7b7ddb9c7

Observation 820abb87-4d33-411a-94d7-9a31c1c7b3d0 · inbound

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts cites this paper.

Rethinking Dense Sequential Chains: Reasoning Language Models Can Extract Answers from Sparse, Order-Shuffling Chain-of-Thoughts Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:50:57.488474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-11T02:11:19.295354Z digest=sha256:53976b8c936549d3b6651801e8c2743e915a684690934d8e84172ac0b820a61f

Observation c1d2ed23-1fe4-4172-bb6f-76a977e339a0 · inbound

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups cites this paper.

Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:01:28.555788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T01:24:44.369468Z digest=sha256:bb886aebe2d50f8423579d8a33489bde7b14282b401f5810bdccaa6d6661091a

Observation 9bc355f9-bca0-4ccc-87b2-9269ae3143ed · inbound

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning cites this paper.

Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:42:07.056737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T02:12:22.228631Z digest=sha256:0d2f0320f2e456a356d0e1a68a4e8ddbae2e9ed1a57cee4fdbf05419df32c3fc

Observation d1ea7ad7-0c96-4c78-8f4c-db1281957eb6 · inbound

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence cites this paper.

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:26.803499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T03:26:18.375974Z digest=sha256:b122145dc3575cd89d09b86c15f61eeb0a38f49964dd71eeec43bdce92a62851

Observation 4a77f969-43fb-44d0-8915-50e476c8dae2 · inbound

Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal cites this paper.

Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T05:31:25.418102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-12T05:11:51.012343Z digest=sha256:c3e3a04e97c5e99b7db31a2d0459eb8719b4340bda4359b4b1d9f81df2244bd3

Observation 4ec5485b-b928-4e74-be88-620b12d6f7f5 · inbound

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies cites this paper.

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:06:26.211209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T03:46:03.117347Z digest=sha256:29d72918df4de05d8ebe6bc1b37dd04d4d65804b0024156d9f745063ce8c8ce8

Observation 05a6b4da-db17-48d0-83db-c1ec0a987751 · inbound

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies cites this paper.

The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T17:27:41.365207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-19T17:24:53.084270Z digest=sha256:a49995fa542bf26bed9a68273c7c6b5c5219da31c2316ac1e44a63a341d666a7

Observation 74a87801-8f24-411d-b438-93cf68baf7e7 · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:21:25.621709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-12T03:27:29.025166Z digest=sha256:0d42b0e98b08262d69e5922fa1bb0c7be7e02276a17f4baa343e0490e282027e

Observation 637f22e8-657e-466b-8e93-f8dcc1f13f09 · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:53:47.161520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T21:53:37.211145Z digest=sha256:92f5698d949546203df9bd1c179fd1a418c6a3766fd96b88e85f8acc4af905d7

Observation 7f81a893-d4d0-4aec-8b69-69eb139dcd62 · inbound

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition cites this paper.

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 85

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:52:09.152169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-13T02:47:10.398564Z digest=sha256:d01ac2f52bd9e4a9f818ee93e8eb40dffac7af2315261daa1f60ed089f3b2e0f

Observation 1d143636-e5a9-4252-852d-841c5f22aa9b · inbound

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning cites this paper.

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:17:07.213602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T02:07:59.111403Z digest=sha256:8993f19e84bac5dd23d117ac8c143c20983554469e72c66bd9fcff15ae7589eb

Observation 4e91a1cd-5c6b-4b23-90ae-2488bea25c19 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:32:24.382672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:1533b7980371ee4e947760a651b305172c8f1f4f6537c7a2aba5bf9df08233c9

Observation 621b7727-b8a9-4d84-8ab1-5cf57b84cb06 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 251

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:17:54.883508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:385f51f956f5a504eb423af5a9a85440a01bafe23471c535e267dc939b180cce

Observation ef1241cc-4c20-49f6-8de7-26dd850af9d7 · inbound

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict cites this paper.

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:03:28.750587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T02:03:23.662430Z digest=sha256:cac9545fcac3297cb1908e7c5cfb444ea24c26218d6ffcc188fcb45765ae0ac1

Observation e16f1b66-512c-4c1c-94ae-16e280db0662 · inbound

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict cites this paper.

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:42:39.903716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-19T16:38:32.946287Z digest=sha256:2c082ddca79825c4a9c2e109ede2a0d5885932305895778a9618ca2e7e4bf78f

Observation 65e71b72-1da9-4f6e-9dbc-5fc7625af95f · inbound

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict cites this paper.

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:25:46.754523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T21:21:56.443981Z digest=sha256:c5d4bf0cfee1444adecd878337b5e77d67c7d578e519c29da25677b7d55938aa

Observation 81a1ceaf-ac29-4d60-b74e-3cc60c43b0ca · inbound

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning cites this paper.

SpeakerLLM: A Speaker-Specialized Audio-LLM for Speaker Understanding and Verification Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T03:24:55.834111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-15T03:20:37.488507Z digest=sha256:1059271b8a2cb1721d23ca618853e2fca9a19747f41bf48c574326ed25ece67d

Observation d4cfa9e1-dfdb-4afd-955b-db651014cec4 · inbound

Proof-Carrying Certificates for LLM Pipelines: A Trust-Boundary Architecture cites this paper.

Proof-Carrying Certificates for LLM Pipelines: A Trust-Boundary Architecture Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:33:46.375288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T21:30:55.435207Z digest=sha256:82f284571a779198ffe6bf51763494b5990e7a6bd702bce87bfd31da361d3253

Observation 790ba829-19bb-4bea-be47-7f05ef043516 · inbound

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models cites this paper.

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.276808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T13:30:19.265822Z digest=sha256:aabea21c6722320b08e9ee93a1e6a6347459af872198c950de2c3bde6ce5afd4

Observation 98ce7102-bef0-430a-b114-039ec4d62884 · inbound

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models cites this paper.

Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:45:00.769683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T19:42:04.998498Z digest=sha256:5fb207e1147fb70cf4ec6c4142422c1e53ea1f3a67acf561ae5ac1943f09abf0

Observation fc7bda6c-3943-4938-b6ff-69fc3389eff3 · inbound

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning cites this paper.

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:08:13.518077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T11:06:29.022994Z digest=sha256:ca62a652f24b0abd3372db9e9281cf057302d21294cc6c997bcf2b15156d8e1a

Observation 14649c9a-402e-4a02-9409-50aabb5f7529 · inbound

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics cites this paper.

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:03:13.896633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T10:54:20.285841Z digest=sha256:353b57d46b253309b2ec435eb078b61a9e158f4e3279d814175dd7433f72a9f1

Observation a8ab7e8f-fece-418d-ba31-6dd08cb11304 · inbound

Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries cites this paper.

Auditing Reasoning-Trace Memorization Claims after Unlearning with Head-Conditioned Canaries Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:13:21.046295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e745652fd91ff3b22908868b9195ed3985f4163f0caf1b516fb7a7e4c8c381ef

Observation 301e8aa8-31aa-4f62-bae1-7f31ced0b1a8 · inbound

Counterfactual Likelihood Tests for Indirect Influence in Private Reasoning Channels cites this paper.

Counterfactual Likelihood Tests for Indirect Influence in Private Reasoning Channels Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:08:17.850961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T13:05:41.969600Z digest=sha256:340274cf2eef820150f28c4d66d4af76d75c867c9e319249ac5ab586b49f34f4

Observation a38f8d6f-7206-4c3b-9e7a-33a480131a0a · inbound

Probabilistic Tiny Recursive Model cites this paper.

Probabilistic Tiny Recursive Model Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.900954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-20T05:39:38.909349Z digest=sha256:283aeab8fae8d386b15275ec49a216003a7439b89c3a98dbd4eb769ee21a3b61

Observation d2e5c683-3e22-4d3a-8381-1b7d4ad53476 · inbound

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation cites this paper.

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T08:06:15.339510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-22T08:05:42.212459Z digest=sha256:b5503c221162b9dba4e40516f195b610bc211c2bb045f62889af9c8154a0485f

Observation 95ebcf97-bc83-4d24-aa57-1c515d2f7188 · inbound

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models cites this paper.

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:30:25.614868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-05-25T06:27:42.024094Z digest=sha256:a0c89a949b63755c2bb98ed9ba2ab26b4a1ee9b0f6e4e309129d581aca92531b

Observation e865cffd-f1c8-4933-9ffc-812a06915a5d · inbound

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems cites this paper.

The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:30:22.721236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-05-25T05:29:39.640753Z digest=sha256:9be2d2934cc6755b522ad9f0330d84a526e748e88ca8c29ee7c50efad379e368

Observation 6b6f0827-30fd-488b-b616-92e9d6ef061d · inbound

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges cites this paper.

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:55:05.917805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T21:50:51.351735Z digest=sha256:09556e6cf91e99d3daa2fd4e29437d1318d0f8ddb47b18e7b8bd2337445cb030

Observation 4e8771ab-b875-412f-b970-bca2cd12e1ed · inbound

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning cites this paper.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:34:47.976533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T15:29:34.096277Z digest=sha256:f8c786951f346d4d0fb6f9074c19cbe135ca56e5526b1331d6e541dce07c78d2

Observation 59d826d6-53ac-42d2-931a-5b508ebd55ef · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:44.431053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:31e5f4893462a76df229372a21312923fb0ff2cb74da7ec18092c7a6c98b0ab0

Observation 1a1bb1f5-6f6a-4757-ba02-c70694a3ae0c · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:34:40.585534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:c155ae205f6d4eda5172d50cd8d3cb51a77919ebdaadd0db8705866c880e9785

Observation fa334a4f-1415-47fc-8b5b-c10b56ceeee9 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T08:05:31.918695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:e2b72fa10ea913d00097ccb2e359c59ba1ac6ad64b337a068b925fc01ba9a60a

Observation 0ffcc3db-fbc7-4d8e-bf47-537cc5b7b153 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T23:27:26.729039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:b1839ae6d8309271e17042fb3edb0d5f528f14145eaee30f03a646a40a0ad12f

Observation 1d12e536-83c5-4cd3-9e61-49cdbdecd62a · inbound

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization cites this paper.

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:24:39.906418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T12:17:12.602012Z digest=sha256:6f650003fe09581ead9dc4260150ee1433bd7aba8823b2136f67a13a3d3dbe47

Observation 5663c35f-c492-430a-9f31-f76f80bc92b8 · inbound

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth cites this paper.

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:04:39.065815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-30T11:56:53.355299Z digest=sha256:158d0dbd95752dfdfc11ed3e8663d1bf28eaf4d71918d69040d316583c8eebf9

Observation 45ca3f4c-0862-4d92-a20f-2f31f702343d · inbound

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs cites this paper.

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:23:59.835106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-23T06:31:01.910684+00:00.

source=pdf_text observed=2026-06-29T22:23:38.178195Z digest=sha256:c2788bec97e713ccac92c760693755c3f6332ac5632bc088a52d8fa102a734ce