Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2503.08679.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.08679 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:20:56.946510Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9194b23c-8816-44fa-a0e8-38533ac977b9 · inbound

Phi-4-reasoning Technical Report cites this paper.

Phi-4-reasoning Technical Report Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T03:40:25.706499Z digest=sha256:332a18b9014953d834d67c934960bb09ce114530c565e3dcdcd6c21c53db5795

Observation 65660cc0-fbd5-49e6-a3ff-f2d982edfcc9 · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:44.340730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:44.340730Z digest=sha256:1971d69b1e19c07682e842074584a95295b582fb76c242f9c2d0ca771c76d30c

Observation 7a17aa33-fcff-43ff-a872-6e30272278cc · inbound

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs cites this paper.

Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T12:44:27.769465Z digest=sha256:5c74515d6e5aff64095868ae6a2ca455135c62ea402ad44bfc0a46b37db77672

Observation 615326e2-1ec1-4e27-996c-c9fcaa96e469 · inbound

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs cites this paper.

Answer-Centric or Reasoning-Driven? Uncovering the Latent Memory Anchor in LLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:33:52.801644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:33:52.801644Z digest=sha256:da8b705c20ecca00b0fd4d4e7f0ea134f2ec4dbf653a120f8295a2bbf34022f7

Observation e4047ab3-1937-40f7-8b9c-f670e5675eb9 · inbound

Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs cites this paper.

Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:24.498510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:24.498510Z digest=sha256:dce5db697ac198c23e27f8b0587182941fd3d9ebffcc2dd55720f59a04134dd0

Observation 311a0f42-6e0d-4e61-8cb5-2c1e3812b90f · inbound

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality cites this paper.

KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:33:08.719028Z digest=sha256:a667512015986ccb148ea78b49057598cfeed7ab98d81b5cf684e12a02d13598

Observation 2ad3bb70-9718-4362-9b7b-5b5f34ad5512 · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:50.831268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:50.831268Z digest=sha256:84c09a6f91f2d1faef2cb811c0e80fe8b3b87483085e7e4ada0f2f1428dccfd4

Observation dddb8680-2d31-4975-acb9-d35b4364e831 · inbound

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation cites this paper.

Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:12.997242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:57:12.997242Z digest=sha256:7871d46ff5c197d8a6beb26649a62aa720b701f501d70a52c23d16abc9672ec5

Observation 1cbfa818-0ff0-449f-8e6e-8defeacc35b9 · inbound

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit cites this paper.

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T04:54:38.908451Z digest=sha256:22629a5e5a09e76e3b45a95718718762f52a96e750b5ec491fa57f72aa43b15e

Observation 53e75451-60cd-4f23-857a-77f6bd6da8f1 · inbound

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit cites this paper.

TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:07.947618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:51:07.947618Z digest=sha256:d0dd784a4385bd5758e16e39c0ad437be441c3f892339ea0bda16ce9639d0278

Observation ecf95395-fedf-4c66-8987-a59e164ad1b1 · inbound

Position: Intelligent Coding Systems Should Write Programs with Justifications cites this paper.

Position: Intelligent Coding Systems Should Write Programs with Justifications Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:58.755678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:58.755678Z digest=sha256:3c8c58994ca93065380f03e211f1f887e742dd1be646baab899fbeb0c2c2a45b

Observation 87e6b6c7-e38f-4fa2-b7cf-efbaa83f6522 · inbound

Do Cognitively Interpretable Reasoning Traces Improve LLM Performance? cites this paper.

Do Cognitively Interpretable Reasoning Traces Improve LLM Performance? Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:40:52.629513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:40:52.629513Z digest=sha256:04084b12c33e971856f93d3222ba9b8b05ebb4d924fb3172f1998d3def0f3ed9

Observation 1331e4bc-038e-4a13-9748-b01e2c52454f · inbound

Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? cites this paper.

Analysing Chain of Thought Dynamics: Active Guidance or Unfaithful Post-hoc Rationalisation? Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:30:30.158836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:30:30.158836Z digest=sha256:3ab10285cdd56879006d8052791d292cc5a5bf405030b61e642a95017b92dc0f

Observation 19a9428a-3b77-4944-a35d-8b37187c5f34 · inbound

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control cites this paper.

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:50.382652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:50.382652Z digest=sha256:d02eb9c64325a2804f519cd30f22400c010b812473540646caa3fa969fd81487

Observation cdb19087-9a0a-42eb-939b-20c32f6f35da · inbound

CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs cites this paper.

CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T13:05:31.857966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:05:31.857966Z digest=sha256:44a4fc71425b199943fb4b3ac7f65a8895212c66b1572c980065d4ffc474984e

Observation 2dd18d34-1be3-434d-8793-e115c674b165 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:04.414376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:04.414376Z digest=sha256:0be9c92eea788ca3f4901421f1b361b571df178d07fa246d68563766ced96c6f

Observation f6edb149-d740-45bc-9538-5a21a51714b3 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:13762d4de0ec122924d74f66065f8be0201c161f9942d00f4df1d9a48aa0339a

Observation 77f1045f-9409-4096-9730-639a51ae7717 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.437398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.437398Z digest=sha256:fa8cecbbcb0b7ea0605c453c95c4ee72ab4912b13d3cc5669b5960b9062a6cab

Observation f1acdbd2-4445-433e-8b26-90251c95d133 · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:41.886859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:41.886859Z digest=sha256:5e244ec207f217bb2c31af6fe7a68e54da926d6d5c649963de2d317e2e05186a

Observation f709fc83-dfe9-4e63-8e2c-1254a68959fb · inbound

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations cites this paper.

Explainability of Large Language Models: Opportunities and Challenges toward Generating Trustworthy Explanations Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:08:53.755726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:08:53.755726Z digest=sha256:ebe290416eeea816ce2a72ea9d04b8c6d76a11052c38f0e6a51b69241442d65e

Observation e7d357cd-022e-450a-8303-13b307615b75 · inbound

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought cites this paper.

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T02:44:48.729794Z digest=sha256:0f7e513c92b94ea3de7617717be91c535186ae7ae1e6a2a654f208d42023b1a2

Observation 31e59c10-a00b-4c74-8dbc-1e53c2171a5e · inbound

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought cites this paper.

Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T07:43:09.181699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:43:09.181699Z digest=sha256:bd51bbd75c5d0df5c9886aefbc478b70cc1817910dd2a2e2de57473d1f12d8f0

Observation c6361836-f748-47cc-b3c4-a4e614fd20ac · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:5d6267d7a97964aedff96a74c2987dab321ca5a2c36b262e8b40206426429ea0

Observation 406e97bc-d75b-4dc4-bcde-e8963e4f5691 · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:49.221545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:49.221545Z digest=sha256:271bb96afbd3c8181df57161999fc60016cb50ed4d8e4f3ae2592b74ca505668

Observation d723168a-93a1-4914-9770-99db91369060 · inbound

Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments cites this paper.

Interpretability from the Ground Up: Stakeholder-Centric Design of Automated Scoring in Educational Assessments Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T21:02:32.788289Z digest=sha256:a7a8bf315ea38874f397674eb51242fae68e84f5c4733a2bbfa88afc162b529e

Observation 784a0bf0-189b-46ae-a61e-aee5a720143b · inbound

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making cites this paper.

When AI Persuades: Adversarial Explanation Attacks on Human Trust in AI-Assisted Decision Making Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T13:21:45.525765Z digest=sha256:384e80d3ef6c9ac8e8e9321f9afa8938ac1795b1f0c90df8b6f41966737377d1

Observation acdeba73-909b-4153-8070-569ed35b01c1 · inbound

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold cites this paper.

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T13:57:36.981142Z digest=sha256:e2cbfad5f714db56176e6b3198d9be355bc5bc21597cde8ba13edbe60b7bd47a

Observation ede4e4e2-5534-4676-af24-15b61800fd2b · inbound

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold cites this paper.

ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T01:00:20.188501Z digest=sha256:37d3684f3f491129af80aee48a38d25ede91d9955e995c078c136a9ec591eafa

Observation e25d97f8-ffe5-45e2-a9f8-e2f888310d97 · inbound

LLM Reasoning Is Latent, Not the Chain of Thought cites this paper.

LLM Reasoning Is Latent, Not the Chain of Thought Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:49:05.178087Z digest=sha256:63b16c1fc4348685c5dab93ac491174554aa6188ef6ee6523aa3cebb932ed1f2

Observation 3e39f9fd-9175-45fd-8256-cf26c7a51918 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:6bde92d66864e7caefc43c3c741d9743009b08833e867bdfab0d5af26772ecc0

Observation 12b37a5b-f300-456e-a9ec-00c3ebcb2ccd · inbound

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable cites this paper.

Confidence-Aware Alignment Makes Reasoning LLMs More Reliable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:10:40.020460Z digest=sha256:b5093993f335df082de1dd588166e8f0b288dd4f6f471d15cc61e08252a1a7b7

Observation 87db8dcf-3462-467c-a315-c31a58610d26 · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:27:29.025166Z digest=sha256:7ca6b3b0b7b3ea6c60f9cded4de80fdf36584498f23091a19400a45188f448d2

Observation 0de77f20-3079-4254-8484-40b09a7ed8fb · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T21:53:37.211145Z digest=sha256:e796175efc3b762ed7e47d96063a4b350c4930bb883c4fa8876e140e0bb8fce4

Observation 5fa3958e-b342-4153-adce-24cfe02f2460 · inbound

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition cites this paper.

Deep Reasoning in General Purpose Agents via Structured Meta-Cognition Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T02:47:10.398564Z digest=sha256:cf278aaf44be12fc3e1edc9e914221adce6c7a2c9ac1ae86a1bc7ae940a8b026

Observation 817312d5-2378-4b74-b334-5e8c9f5aaa83 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:f4b4f1359126b80560749d903ca38d65cb25e0f4919b94f7d00bca46e680749e

Observation bd88a687-d925-4417-9db9-72c31c0bdc41 · inbound

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution cites this paper.

Diagnosing Multi-step Reasoning Failures in Black-box LLMs via Stepwise Confidence Attribution Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T06:43:23.870127Z digest=sha256:fa37fd91ebece3e5496cf938587c5c4733ea26d46a01a34c569b28f682c3808b

Observation d8320dad-b0be-4e3b-918b-d58655a82e11 · inbound

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models cites this paper.

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-01T03:02:49.618587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-25T06:27:42.024094Z digest=sha256:57b384368032f47e5eb17b831326e4823812e69fe0d05f362dc3fd89d704f41b

Observation 891d14cb-99e1-4ffe-a516-c78d6ebd3e47 · inbound

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning cites this paper.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T15:34:48.035904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:29:34.096277Z digest=sha256:26c110806dd1afea213ef9e7f0e92cd126ae4ea99836165fd64cbcbe867e76ae

Observation 3fe2150c-f828-4c41-87a6-2444b1bc2959 · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:44.414531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:4fd093400d613506fcf4b1ad7ac8746911abbe0f319a1479477c6a0def93fd72

Observation 015fee4b-cd2e-466c-91cf-5c19db65ad7a · inbound

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy cites this paper.

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:53:59.509298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T21:47:17.894881Z digest=sha256:f212677511da8596113c84891a65f0333e02a3c7a593475f10edc86833b0341e

Observation e335a703-8459-43b5-bf75-eb80c96cf4cd · inbound

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation cites this paper.

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T17:33:45.512215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T17:24:32.401230Z digest=sha256:27bdc93a4efb0247b134bba6662ec5b5f90d41b2c7d8dd0600e8e7e5975eaaeb

Observation 919221a4-713c-4b7c-a11c-cac4fe10542a · inbound

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation cites this paper.

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T13:08:40.058805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:08:40.058805Z digest=sha256:206bf88a6ceb1e108f6009e67988642ae3817aa74ac5221b2aaf44f6085c33d1

Observation f65a3bd3-0475-4cd6-ba98-10aa0fe87439 · inbound

Forecasting Future Behavior as a Learning Task cites this paper.

Forecasting Future Behavior as a Learning Task Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:07:41.352894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:55:39.494339Z digest=sha256:dc2cde624c3c075035d724816d091010d8fcf5a0d73f0df15e18a8b69e6208cc

Observation ce64ce93-e475-4f2d-9290-61b6ba32b49e · inbound

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models cites this paper.

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:18:03.828935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:36:05.700067Z digest=sha256:5f8fb6d44096e4b51e40e61ef4b9dbc7e588ca0b34e90a6186560a350d1c4759

Observation 447b601f-3e95-4e48-9367-8fbd2feb2c6d · inbound

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation cites this paper.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:58:43.406916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:5a4359138fd38cfee2d00ccc7a78fa2fd15ffa96617cc6ba8a056d8fe6492484

Observation d2805322-f0dd-4723-9ca3-bf133385f47b · inbound

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models cites this paper.

Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:49:37.545786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:13:13.678245Z digest=sha256:ae7c93774282ae85a66109f794721293432e3593fb92a073014444643f757d0d

Observation 95de15de-01a6-4a55-b4fe-52017bee5406 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:30:00.655029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:18112f45d341066021cf4ebd1030c9d76c4ac0db15b10af4a5ade2eb8a4392ee

Observation 7bcea873-d6aa-4b6f-87f9-b44b317af510 · inbound

Where Do CoT Training Gains Land in LLM based Agents? cites this paper.

Where Do CoT Training Gains Land in LLM based Agents? Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:49:51.539339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T04:55:14.293452Z digest=sha256:fc4cab0ff7772a54a050d52fc8fe9b38a29f370f88b65a7c09892bf06c06916d

Observation c6c31034-9380-40e9-84d2-dae4f9cd07e9 · inbound

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates cites this paper.

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:46:56.219034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T12:46:40.700728Z digest=sha256:b578cf7b71f39bb2f705d1016c6fd9547569bde01bfc6c532282513104873101

Observation a3dd1913-a81d-4e3c-a99e-1394705840ac · inbound

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens cites this paper.

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T02:02:49.492856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:02:49.492856Z digest=sha256:3f0c0045131af444359c647dfe69a2e15e40cfe5433147e35077d0be97327ceb

Observation b1756608-8180-4040-92fc-3538414a25fe · inbound

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors cites this paper.

Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T17:05:29.695180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:05:29.695180Z digest=sha256:7b46c5d6b22cd9754603c1e3c88f00beba67c6e11927d3ec369cf64432a3f31b

Observation b2add7b3-a570-4564-b51a-b01d972bd8da · inbound

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning cites this paper.

Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T15:16:38.426751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:16:38.426751Z digest=sha256:278da586ac7360970fc52e723f5b3328c37425a4a2c900fee7cec9b8facedba6

Observation 0fd9aef5-5d36-495a-842a-128c892cf410 · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:56:41.060161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:2718656516a477a2f6fd25f3d273c36cfb28f23c54a25981e485a28764d359a4

Observation ec75fd25-568c-4c41-b2aa-3a13edd5f3dc · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:6ecb555e1230a44d4021b7a5302f7bcc03e8994034ecca4337d2cced3e2915f3

Observation a4cd360a-c10f-419a-b72d-18e4acb12420 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.759079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.759079Z digest=sha256:6924deee836ce8852e854ed3ffbf4ff38071a4354c4639fced439f1fc5e7e26c

Observation 82ffa36e-827e-434b-853f-66e4ade443c8 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:30:26.776436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:30:26.776436Z digest=sha256:33b6d50410d9fb6a20f649be92eb2ec8cb176c7b236635c5d73ec3b9cc8125d2

Observation 7e9ee9b0-a0b5-4d20-9c68-2ca6b429b1b9 · inbound

Controlled Reformulation Testing for Logical Consistency in Large Language Models cites this paper.

Controlled Reformulation Testing for Logical Consistency in Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T01:55:04.455545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:55:04.455545Z digest=sha256:0f74f3706f05d262488ff7cca742e8b715350a72599eae58bef928fba23c10d1

Observation 05f8c993-b201-41c7-b42c-5b575121ab38 · inbound

Mechanistic Attention Guidance for Agent Memory Refinement cites this paper.

Mechanistic Attention Guidance for Agent Memory Refinement Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T17:30:38.509216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:30:38.509216Z digest=sha256:c17ede7fbf32f647fb6658cc73771e9ff4e2b363832ec1e6776123e4d18a93a8

Observation 7de66025-d265-4f54-b74a-364463398050 · inbound

One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context cites this paper.

One Student, Many Teachers: Multi-Task On-Policy Distillation via Soft-Prompt Privileged Context Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T09:22:00.540306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T09:22:00.540306Z digest=sha256:6c31cdeea20fc536fd7ce8b56e82424f3441a802d9b0f780c890adbe9abe17b5

Observation 0106a6b7-a635-43f4-bb0a-93dca770eb44 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:41.224620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:41.224620Z digest=sha256:adfc63a5e08ea3b0f06c73e6eefedd84913c482a9503c4447c2492cc04a40ced

Observation 656cdbbf-656c-4531-827b-b2f40dc799bf · inbound

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong cites this paper.

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-30T21:44:25.171383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T21:44:25.171383Z digest=sha256:e9c046e41162ffecc1984dfde9cee0507bf0eee25427444b8f3ca95ed523d8ed

Observation 087608a2-96ff-469a-9bdd-0e8c391224a5 · inbound

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing cites this paper.

Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:52:56.666840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:52:56.666840Z digest=sha256:52f5543834cb62266ff8e2efb1cc6a9ba0624690f63c44b6a3e36cc90391da4d

Observation e7d0c5bc-7b1a-4d33-8a3a-a2e66e218315 · inbound

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models cites this paper.

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:41:53.515878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:41:53.515878Z digest=sha256:65317a77112ac4622dc972b646eb1eb7679a6e3322277c91576de9c2dbacc04a

Observation dfb39c0b-f025-446f-8f35-12f215f3da91 · inbound

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness cites this paper.

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:26:27.978746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:26:27.978746Z digest=sha256:de233de5ca8f3dde58c287cf55b19c5324d1d881494a97b117c73504ba1ef489

Observation a5540de9-5ae7-4408-a903-f8736ba54a1b · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:47.033588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:47.033588Z digest=sha256:168fb9e3b313dfd29111c4da0c55d7a67adfa8338b05400f9e36051404d0adf1

Observation 96f401f5-f452-466e-8c89-fc3e111287e3 · inbound

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings cites this paper.

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:51.013543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:51.013543Z digest=sha256:c10af7df8a4ab5f487234221ea0480077ea069b02692b0af3a4a5f9e77210846

Observation f6abacef-696b-4345-84ad-a1a670263898 · inbound

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking cites this paper.

Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:20:56.946510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:20:56.946510Z digest=sha256:70f8ba93fb5f7d109c778028d6d1f22de09b2d1936a9b1bf31eb8bded6ce6c01