Pith. sign in

Paper Citation Record · LEDGER

Reasoning Models Don't Always Say What They Think

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2505.05410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05410 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 124 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:06:55.819737Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

12
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b0d84af8-53d5-4908-8fbe-29a7d337e092 · inbound

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring cites this paper.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.819737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.819737Z digest=sha256:866a9f19042caf0eba3c3f593cbb1dfd0e118e60f4c7021a329432782a6431f6

Observation e584c35e-928a-4b1c-b8b7-db77b142f3fb · inbound

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? cites this paper.

Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models? Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:40.853592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:40.853592Z digest=sha256:df69f53a43ad86957e5933c4148cc776638327b45a7113665e66387e5d36258d

Observation eab382ac-af4b-45d1-a118-f797cd15a0b5 · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Reasoning Models Don't Always Say What They Think

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:48.472560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:48.472560Z digest=sha256:6cdd9f84998b9b1fabe304121a81813c3b5f7756dc8aab603c2d25e508b13f65

Observation e7f3d973-79fb-47d3-8a48-616c1f3674ff · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems Reasoning Models Don't Always Say What They Think

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:21.225815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:21.225815Z digest=sha256:861b8c567ccae0475a9fd0683dddab70b865068f543d210a47cd2ac3adced4ba

Observation 8fea7934-502b-4000-97ef-23decd091a22 · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Reasoning Models Don't Always Say What They Think

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:40.820858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:40.820858Z digest=sha256:bd2d63887aa13cdc77115f3569b7400863221f2b5b7286bb09e6a6c0d46158bc

Observation d658fca0-f320-4b32-9929-f9116c6da8c0 · inbound

Why do AI agents communicate in human language? cites this paper.

Why do AI agents communicate in human language? Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:49.677629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:49.677629Z digest=sha256:32a42ddaf53606bfcc518ff4ecbbd772fcde1447fc2b61251fab8b1f79d2d75b

Observation 3897f119-b35d-4837-b2a0-23f98bbac994 · inbound

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity cites this paper.

The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity Reasoning Models Don't Always Say What They Think

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T16:10:31.664485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T16:10:31.440921Z digest=sha256:5409816e1163054e07194a2b52192fd83d70a1240f5dc8bdc39198d936804272

Observation cf7e1040-bede-4c42-b203-413bc87f8b2f · inbound

A Survey on Large Language Models for Mathematical Reasoning cites this paper.

A Survey on Large Language Models for Mathematical Reasoning Reasoning Models Don't Always Say What They Think

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:45.045847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:14:45.045847Z digest=sha256:a2d4bb7a7ca23b82510beed0034756fc129c7d302a16bedce8e570e0d8acf779

Observation 0df62072-9738-438c-969c-a79048aecda3 · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions Reasoning Models Don't Always Say What They Think

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.327546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.327546Z digest=sha256:6ab3e9ad1b77acaf3735cb077fc73483ce0de6a77699ff5ae7e607fd51bbf4c7

Observation 2170fee5-fc74-4209-a415-4a639f0c0b89 · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:51.958119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:51.958119Z digest=sha256:3752a23e05e6a69f7300c5d6164ce3a3790bf945e7f925076e01983d0eac37bf

Observation 536ecd38-a839-4c65-af8c-9b27a028b125 · inbound

Listener-Rewarded Thinking in VLMs for Image Preferences cites this paper.

Listener-Rewarded Thinking in VLMs for Image Preferences Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T07:42:09.240618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:38:52.273903Z digest=sha256:a52512c552f740b3a45a3cd237c89659d72ab592b104430ea98e08635def79a1

Observation 465cd223-356b-406d-9dee-178694ad797e · inbound

Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models cites this paper.

Interactive Reasoning: Visualizing and Controlling Chain-of-Thought Reasoning in Large Language Models Reasoning Models Don't Always Say What They Think

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:52.644431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:52.644431Z digest=sha256:53d19e60dc52488bd8886d79426c7dfe44195035e8d6352641b70894036ca0d0

Observation 40f0cd15-8704-4053-a0ff-37067452c6a2 · inbound

Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models cites this paper.

Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models Reasoning Models Don't Always Say What They Think

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:53.363109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:53.363109Z digest=sha256:3b634bc0f6ebb96b0114772ed4a83b01c27200e373f40346143ffd9c4ae6c675

Observation 88018fc3-e8f8-4346-9d24-ad3dee5f9bf8 · inbound

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language cites this paper.

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Reasoning Models Don't Always Say What They Think

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:15.499693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:15.499693Z digest=sha256:efc77939c05816cca40365e0a7a4f7f1d125bddc65df1f4b3a8994c99f9b4be3

Observation caa0589b-8752-4308-b471-8fef84c0721a · inbound

WebGuard: Building a Generalizable Guardrail for Web Agents cites this paper.

WebGuard: Building a Generalizable Guardrail for Web Agents Reasoning Models Don't Always Say What They Think

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:12:47.648456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:12:47.648456Z digest=sha256:47093ebc23e9e0f6538d8d005df43eb1cc626623f72036e685addeae49f49a73

Observation 5795bd41-e36f-4944-a6bc-19e2b558e834 · inbound

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny cites this paper.

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny Reasoning Models Don't Always Say What They Think

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:10.865469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:20:10.865469Z digest=sha256:cc31935c0018ba7e27384b594d9211ab61c243d9c5756a185fa88335022accca

Observation ef73b372-82cc-494e-bfc6-1fdb4ff1aad1 · inbound

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks cites this paper.

Lightweight Language Models are Prone to Reasoning Errors for Complex Computational Phenotyping Tasks Reasoning Models Don't Always Say What They Think

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:06:07.060768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:06:07.060768Z digest=sha256:c3699117285416d399546211ac58ce70a821a5d790b4f122c21e5b531a2004cb

Observation 640802df-da7a-4574-9bb4-b6d7ea69a610 · inbound

Position: Intelligent Coding Systems Should Write Programs with Justifications cites this paper.

Position: Intelligent Coding Systems Should Write Programs with Justifications Reasoning Models Don't Always Say What They Think

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T23:01:59.326685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:01:59.326685Z digest=sha256:62b3ffe29de85f2e9ba4d051a0bb880625fb81adb011da884b209d08e06cfab0

Observation acfa9999-7359-4316-860a-4b79b5086acf · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents Reasoning Models Don't Always Say What They Think

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:48.111145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:48.111145Z digest=sha256:c4b8e668232d3936a8b51db7e42997628cb49ca7a303286a2a6bcfdfae3e1332

Observation e706203d-4698-45ef-9caf-4a22f187b0d5 · inbound

LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations cites this paper.

LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations Reasoning Models Don't Always Say What They Think

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:12:05.814286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:12:05.814286Z digest=sha256:8637740c85555706a3cdcf8a97acc8bccaa0d2a07623030ad4e40d9558a84a46

Observation 3653af64-4892-4a1b-a832-592267de95d2 · inbound

Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight cites this paper.

Dark Patterns Meet GUI Agents: LLM Agent Susceptibility to Manipulative Interfaces and the Role of Human Oversight Reasoning Models Don't Always Say What They Think

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:40:08.174781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:40:08.174781Z digest=sha256:ceab037471546cde341c0d5be664645893c16c842b3dcd6ebf2d2391518f75b9

Observation 33013157-cb8a-442f-9429-a3e535258762 · inbound

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs cites this paper.

Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:52:41.098958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T14:51:42.879926Z digest=sha256:8a5b7abd3e52ed67ecc4d7f633db3b108bbec266fc48ec24b9e7ef1184472dd0

Observation 399715de-4411-4b6e-a0b8-7346bfde06b6 · inbound

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts cites this paper.

CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:40:42.937652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:40:42.937652Z digest=sha256:95a9e4c85044a3395c7cf78e78d3148d04542bb79e7d0ff3e278881adce77f34

Observation 16b83c19-f98d-4cdb-8de1-2f7b2e26fe91 · inbound

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes cites this paper.

MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes Reasoning Models Don't Always Say What They Think

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T09:16:39.944551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:16:39.944551Z digest=sha256:aaa3f22645a5ff9d4a3b0d8494fa024a3434af178cd0ef380faa489e680317d4

Observation b0d2b242-9c7e-42cc-8bd1-5d10b8716736 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:12:23.661588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:11:29.205366Z digest=sha256:40e0eedf15f255760f6323fd0cc072b73b5acb22ff671f49a8a40be32667f2a0

Observation b41a6e13-c940-409a-8066-8b6103ea7860 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T20:10:34.943343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T20:06:16.172916Z digest=sha256:f45a0f4a98e82c695bfb0594449f01d006dd64a1cbc06d80b164f30a0b134fe0

Observation b4076495-9561-4ac2-a970-59b3cd42d405 · inbound

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning cites this paper.

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:53:05.547904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:53:05.547904Z digest=sha256:c754148d846d30c9779d0bd0263509167557b99bc230fed84dfe03f2f1e4ec85

Observation 962c22c5-bc25-4353-8cf6-52aa84e33783 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models Reasoning Models Don't Always Say What They Think

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:49.310593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:49.310593Z digest=sha256:8805696f99a7566c68cf15b39d1719a427adb51adcfd3012b80dd7d2214c5f19

Observation 15d4e5a5-a2dc-40f9-b01b-c4b79c422fd2 · inbound

Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment cites this paper.

Does the Model Say What the Data Says? A Simple Heuristic for Model Data Alignment Reasoning Models Don't Always Say What They Think

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T04:08:59.918559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:07:08.752350Z digest=sha256:2450d2a2c2586223e6a0c1eb480e292891d3fe4e0992c7ee7d6f70fc6f2412b9

Observation e2f94947-bc2e-4e8a-9c83-19f530296e93 · inbound

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning cites this paper.

TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning Reasoning Models Don't Always Say What They Think

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T15:51:05.408221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T15:49:26.750021Z digest=sha256:bbfa16ffb3db87da9f4e1c7219a216da8b4225cf730cf3dea2214b8dffb7abb7

Observation 5641d735-1424-4009-adf0-ba2249c06512 · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:06.392506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:06.392506Z digest=sha256:ae6826d39f51cfc5a7a2127486fbd8ad0d60c4982d178dd345cdc841857150ed

Observation ee42911f-be32-4e48-ad5d-dd22f84a1634 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:29.156561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:29.156561Z digest=sha256:e94f748fa61b400a7b6199c704253d410691f5a0866f2326d7b13bb1e037c3ec

Observation d052d0bd-f976-4c9f-a265-938c17305c5e · inbound

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring cites this paper.

A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:01:29.941429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:01:08.076475Z digest=sha256:194362a0a0b61faea51a7d235fcc635c2a5297d95401c4e411706b36f2c09f75

Observation 0208175d-d09c-48dc-a403-16ac3708cd52 · inbound

Decoding the Critique Mechanism in Large Reasoning Models cites this paper.

Decoding the Critique Mechanism in Large Reasoning Models Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:50:27.509774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:366bd0fe4b3288faba5244edf72cde966c000f150041c6df9de231881ea90ae7

Observation 2d8925cd-3703-41b9-b029-08f8412cb9f6 · inbound

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness cites this paper.

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:19:37.314423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:18:48.798851Z digest=sha256:3d99efb6f4b75a415285bfff372ad44c16c8738156a3e2ed9c251be37b4db36c

Observation d55b2e55-70c2-4e9f-9061-35dc68a0ca37 · inbound

Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor cites this paper.

Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor Reasoning Models Don't Always Say What They Think

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:31:37.034954Z digest=sha256:e031ab17c50adb5476966aac2e2d5fe6a3f7d5cfe7ec27aab37accda316f9a8f

Observation de4e5c57-26a4-4f22-9482-9b20151a6582 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Reasoning Models Don't Always Say What They Think

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:e8949a7daba0d6dc50089654c5a76a16c9a513bebb56e7b8e1eed1c653a18b35

Observation 1ea296e3-d673-435b-a44f-d0104d2989fa · inbound

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models cites this paper.

Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T10:58:30.370708Z digest=sha256:7b1520ba42939a6be55df4dcab172ea90717a2d0208b25aa8f4cb83ce9b98616

Observation 15a8e342-8d97-447d-9c15-1b6d2ef2d903 · inbound

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography cites this paper.

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:54:35.000783Z digest=sha256:90da713567db3e579a8ee9a2c2b52d2990b28043e9a3d0db7065f1bb3fdb507d

Observation d8fc774a-cf78-4284-9d2d-b5f5eb752d97 · inbound

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography cites this paper.

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T19:46:06.596310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:46:06.596310Z digest=sha256:734bb4295da07432c7b48a279763896790e2f36055bbf0a1be97fb37ca0d8cf3

Observation 0e9ff09e-a53d-40a6-90b6-00cb2d7f6d00 · inbound

LLM Reasoning Is Latent, Not the Chain of Thought cites this paper.

LLM Reasoning Is Latent, Not the Chain of Thought Reasoning Models Don't Always Say What They Think

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:49:05.178087Z digest=sha256:18a79a43fb3f9f29a8080a0e9e174af5e7686d80f90fab12650648b98b0f96d6

Observation 0ab57bc7-0888-4fce-bbdd-9816e3682432 · inbound

Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment cites this paper.

Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment Reasoning Models Don't Always Say What They Think

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:29:54.013922Z digest=sha256:efd285bfc3bcac417ec0f78401462ac416312a1efcf8c68b46d532aeb2890928

Observation f33251ec-8c90-4958-b480-5098bb9e89d9 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use Reasoning Models Don't Always Say What They Think

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:3a5bbd88d22692ecfa4d8913d05f543edb14cc2b3d92bc9afb5645e16e9c8763

Observation f01d4a45-c755-4998-a99a-18d72c0bc82e · inbound

Knowledge Distillation Must Account for What It Loses cites this paper.

Knowledge Distillation Must Account for What It Loses Reasoning Models Don't Always Say What They Think

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:58:41.449172Z digest=sha256:fa7441a82d6b14c0c135b402ffc33b0d898e3da1e82bb28c9df58bafaaf538c5

Observation 8f6e77be-77d3-46cd-8720-222d36b097d2 · inbound

Knowledge Distillation Must Account for What It Loses cites this paper.

Knowledge Distillation Must Account for What It Loses Reasoning Models Don't Always Say What They Think

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T03:31:56.787201Z digest=sha256:79d217171c4d3d52c732646f7acf40dede88185edfe0dfb334f9e3d6cd686af8

Observation 00e2d83c-05f3-4c81-8a42-fd6a2a80bea4 · inbound

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning cites this paper.

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T05:51:45.499422Z digest=sha256:522cea79e47f22e791e729f1f41cc9e30752c450dd8dc49d85f6ef0246074822

Observation 4990d3c3-1389-4523-8f49-7d3fbc3c8800 · inbound

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning cites this paper.

LLMs as ASP Programmers: Self-Correction Enables Task-Agnostic Nonmonotonic Reasoning Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T08:15:31.889929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:07:20.134045Z digest=sha256:641a794656769a0e5a77f34d578d7f231f6824665e27dddcfbdc30b188b30d92

Observation 27acda19-81ed-44f8-aa1c-25997ffef768 · inbound

Compared to What? Baselines and Metrics for Counterfactual Prompting cites this paper.

Compared to What? Baselines and Metrics for Counterfactual Prompting Reasoning Models Don't Always Say What They Think

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T19:02:46.991897Z digest=sha256:ae8c549bd5d488a6ab4efa845f02ae71068d14a164b980089f01fb0551f19398

Observation 4d178c74-9cdb-4d7f-b9d5-76aa514d3719 · inbound

LLMs Should Not Yet Be Credited with Decision Explanation cites this paper.

LLMs Should Not Yet Be Credited with Decision Explanation Reasoning Models Don't Always Say What They Think

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T18:41:24.552939Z digest=sha256:e81649e91bd637e93a70c94a99f0d055c9dcc6581700748a76332e6aa16f5b8d

Observation 443bff77-af49-4d53-abb8-cd4e56939818 · inbound

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem cites this paper.

How Well Do LLMs Perform on the Simplest Long-Chain Reasoning Tasks: An Empirical Study on the Equivalence Class Problem Reasoning Models Don't Always Say What They Think

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:29:33.453354Z digest=sha256:c29aa4e44f7252fa8b8853b614a1643e4f52459f70a492fcbe9a2dee03eb728d

Observation 905a0f71-d091-47e3-95df-9c3ae5bb58dc · inbound

Weighted Rules under the Stable Model Semantics cites this paper.

Weighted Rules under the Stable Model Semantics Reasoning Models Don't Always Say What They Think

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:38:04.965252Z digest=sha256:d61b288e90b126c344caa8d786d8ca1b92612921f18582aaf66f1707bd23b10b

Observation 0890147d-6344-46d2-b322-4049a2b5ae88 · inbound

Medical Model Synthesis Architectures: A Case Study cites this paper.

Medical Model Synthesis Architectures: A Case Study Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:27:59.466519Z digest=sha256:54d78ca01c28a866deabb03a888186eb8eb609c9a72e76aa107d94211f396447

Observation 1379bc7a-10d2-4f7c-bb5c-e48a32918246 · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:13ab50901eb026ae63d7e1e9823fbf7a633fa993a7cccb4943248b9bb07a85b0

Observation 97761f00-d0d9-4985-8d59-97100c0a2408 · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Reasoning Models Don't Always Say What They Think

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:27:29.025166Z digest=sha256:58e3a37d14eb38400ecd8d2fe45b680bb06141bb22d5b07b44426cae9e40749c

Observation 18fac225-07d6-4ccb-9680-e9ea6f8dc4c4 · inbound

Evaluating the False Trust Engendered by LLM Explanations cites this paper.

Evaluating the False Trust Engendered by LLM Explanations Reasoning Models Don't Always Say What They Think

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T21:53:47.150254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T21:53:37.211145Z digest=sha256:cb35351cd069fc16a7db32fe55fb456035aa6c766128884ff8278272feb30d81

Observation a6c9611b-04e9-46f6-a459-6342b5d71e30 · inbound

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning cites this paper.

Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T02:07:59.111403Z digest=sha256:040d461cdaf4b0b090f8200d068ca60b9b2b5e9914762a8a96b23926b9cb8141

Observation 298e183b-3d79-4f02-82bb-6a5e067eaa22 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Reasoning Models Don't Always Say What They Think

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:1d5160a3e8c6cf353fc539a23d02c9cc506eea7597ab838e27e3e7dac51867e9

Observation cd56e294-43cc-4ae3-a435-b5a8a2df0670 · inbound

Trace ideals and uniserial modules cites this paper.

Trace ideals and uniserial modules Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-15T11:13:18.391956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T11:13:18.391956Z digest=sha256:6457de5a9f43b2b55f478e58c03912de9e68c74f6fc8769de268f0cc8ba7f624

Observation 2d5bbaf8-b8ad-4afa-8f43-8315ffe1a154 · inbound

Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs cites this paper.

Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:56:11.373242Z digest=sha256:f591a0d37819f1b755507ed2b1dab5b5a4f04234c76de2dbc809c0516390c6cc

Observation 24b201f8-68b0-46b3-aafc-03a13b094649 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:32:56.763151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:9e3ab4b4b8991777ec067ef96cace840572fc6858157601605943ca76ebc2241

Observation 7df7c8d5-f52b-4b61-afeb-25c3fea8e00a · inbound

CoT-Guard: Small Models for Strong Monitoring cites this paper.

CoT-Guard: Small Models for Strong Monitoring Reasoning Models Don't Always Say What They Think

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:39:16.264272Z digest=sha256:eddcbea9b58234419a400e3b2cd60d39f75bda8ba2efcae28ff548d772e177a5

Observation 14716553-4a52-4687-a8cb-4b99e8b0d481 · inbound

What properties of reasoning supervision are associated with improved downstream model quality? cites this paper.

What properties of reasoning supervision are associated with improved downstream model quality? Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:18:01.027228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:36:30.699633Z digest=sha256:b9427beb7237d3c4a50e175ca37a2e6ac2bbeab180cfa12f0b7afbbb3012ecaf

Observation 72267283-16c0-42eb-a0c2-f9d345092281 · inbound

Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning cites this paper.

Bridging Legal Interpretation and Formal Logic: Faithfulness, Assumption, and the Future of AI Legal Reasoning Reasoning Models Don't Always Say What They Think

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:25:04.091839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:23:11.698521Z digest=sha256:bc22addfe9894e172f65e7b7241441beef7a5d0530af01c5037d6c6305c0e4d8

Observation 0c687c55-bc96-45e7-bec1-1d32ecf88a00 · inbound

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades cites this paper.

SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades Reasoning Models Don't Always Say What They Think

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:33:32.331296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T02:31:18.183715Z digest=sha256:bcc3db4fc35247108b0277b24f4e4d808dca3cbebe72459cd395d4e0bde323d1

Observation 2bdf5cbd-10a0-41f5-aab4-9f883173e6a3 · inbound

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations cites this paper.

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations Reasoning Models Don't Always Say What They Think

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T16:03:33.155685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:59:56.231210Z digest=sha256:6fc5aa8faaa4ba65841dc20578cd3e940766026bc70d1deab22f706e1a16b222

Observation 8b6033d8-dc92-4d3b-a1ed-04c886432846 · inbound

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics cites this paper.

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics Reasoning Models Don't Always Say What They Think

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:03:13.933854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:54:20.285841Z digest=sha256:507c7b15f43c6fa004efee4e0f183dbc61b0b5fef8cad59993b3ea5de33f2a55

Observation 5de14b6c-8e1f-4937-b5c3-9ee0e0899c22 · inbound

Probabilistic Tiny Recursive Model cites this paper.

Probabilistic Tiny Recursive Model Reasoning Models Don't Always Say What They Think

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.907422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:39:38.909349Z digest=sha256:5d369cc54130bc66ed302c45cdc7cd98cfa068a11dd542fb024a76da6e6eb69b

Observation b0f81a3c-ea24-4265-ba64-297fdec4f1e5 · inbound

Neurosymbolic Learning for Inference-Time Argumentation cites this paper.

Neurosymbolic Learning for Inference-Time Argumentation Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:13:03.466611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:11:44.118384Z digest=sha256:80b861ce1b9cec6d79235c8ce6b7ee6ae94f2079f5b511bec8cbb9497d87ca54

Observation 3338abe3-aff3-4a91-9965-6d1cae5ca940 · inbound

Neurosymbolic Learning for Inference-Time Argumentation cites this paper.

Neurosymbolic Learning for Inference-Time Argumentation Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:04:57.778119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:04:21.013167Z digest=sha256:b8a46d90e632ebefdde06f1c52b00f88b1677f5e38c059393142fde6f36975e5

Observation 72b1b6c2-5744-47eb-b542-7568fe0aa45f · inbound

On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective cites this paper.

On the Cost and Benefit of Chain of Thought: A Learning-Theoretic Perspective Reasoning Models Don't Always Say What They Think

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:13:58.676304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T05:09:37.588841Z digest=sha256:b413c305194896e75d7a93034c1182d52ae9daa95773a907af7cc8d3eb235ff5

Observation 036aa5de-343c-4c35-a9cb-dba17059b446 · inbound

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models cites this paper.

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T06:30:25.628182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-25T06:27:42.024094Z digest=sha256:2744ff215bff1ca15a161a08c8b8a56aa6e642f33ba113a60cfd16ac3a46896d

Observation 93f60722-4667-4c28-bae8-603703863689 · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:44.433322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:dfcf833e936016bd0e16e787f880a697126db3eaae77fa4ba9e58d324d08e380

Observation 36615e4a-62d5-42f8-a29c-05719f7a5fe3 · inbound

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization cites this paper.

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization Reasoning Models Don't Always Say What They Think

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:24:39.911749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T12:17:12.602012Z digest=sha256:fe0e5cfbc041fd1c3aa743b1e7abe0592e344f3cf39d037bd6d8965674f13efe

Observation 1267243d-b84a-4a28-a165-467128bd3df1 · inbound

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth cites this paper.

Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth Reasoning Models Don't Always Say What They Think

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:04:39.043070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:56:53.355299Z digest=sha256:5b8d40f201a9057c7fd8d587bd703f145baf5029687a57dd86dddb8beb1bfe97

Observation cc117f66-49fd-45d4-9e01-d012482f881e · inbound

The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure cites this paper.

The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure Reasoning Models Don't Always Say What They Think

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T12:03:24.170793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T11:56:54.670979Z digest=sha256:2599c85a0479795576ffcc5d79e93a242e4fd5637b4048af89c22e48d60a325b

Observation ea7361b8-2675-4d0b-b8d1-99997e495135 · inbound

"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents cites this paper.

"I Strongly Suspect This Website Is a Scam": Benchmarking PII Leakage and Detection without Defense in Autonomous Web Agents Reasoning Models Don't Always Say What They Think

Reference 118

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:42:36.153240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:53:41.420255Z digest=sha256:bf1f47399dace3d75bca6acf67016ea0d66fb8bebc37f516e255dda79979e971

Observation 5a3e24ab-bec8-4d1b-b8b5-2164b985e945 · inbound

Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs cites this paper.

Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs Reasoning Models Don't Always Say What They Think

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:52:35.744075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T18:48:36.147561Z digest=sha256:cb20d315d7e71c4c328cafe23ae3ba24a70d37ebff0af82aa9636750a10addc9

Observation 14c8dcfc-e9df-4fe1-b08c-47d49a33643b · inbound

Quantifying Faithful Confidence Expression in Large Reasoning Models cites this paper.

Quantifying Faithful Confidence Expression in Large Reasoning Models Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.237014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:24:38.335417Z digest=sha256:0cccb08df5be5c395cd8410dc1a64e4245a9fc737c2a18b0cb02e1447ee13e1a

Observation 0edb1f89-ee21-4a21-811b-fe05be39555c · inbound

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense cites this paper.

ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense Reasoning Models Don't Always Say What They Think

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:16:58.775410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:26:22.434043Z digest=sha256:dddfc4c92316c3f42acb7ce53a8448fd2f68ea1047de2a472b5f8063dd7a8272

Observation abf2824c-88dc-4bea-8a86-df21dcbead16 · inbound

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models cites this paper.

The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T01:31:29.339523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T01:25:07.890796Z digest=sha256:5c716afbd9870d89743962d78ebcc4fe87a7f95d29e78648ef6710aa4ea3e555

Observation b22e2154-4f51-47c1-a388-6dbf500856e1 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:36:59.388654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T01:07:17.014002Z digest=sha256:875828336d807675a86dcb8a514200af48202be2196ff12015e75e7f74e64a34

Observation 9efaf398-c88f-4b7f-aaff-b12045c5eb38 · inbound

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents cites this paper.

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents Reasoning Models Don't Always Say What They Think

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:19:46.889450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:19:46.889450Z digest=sha256:8d4f6f7627d85c2c7a2bb8177b4e896d89c7b8339e35ca3c336ef7c1deef0982

Observation ced19d64-699c-42a9-9dfb-ba8fc5a9e584 · inbound

Sycophancy Towards Researchers Drives Performative Misalignment cites this paper.

Sycophancy Towards Researchers Drives Performative Misalignment Reasoning Models Don't Always Say What They Think

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T22:27:26.205760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T18:52:29.375027Z digest=sha256:d045fd18139db9c538fecff2e1e87336f616643abbc6d8c30fbb871873e275d4

Observation a6ef89e3-3b6c-49e2-9372-19cded203eeb · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reasoning Models Don't Always Say What They Think

Reference 239

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.562997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:5d28a5da6ea21d314a227a24e79ad72a1da5290468146c232a4357a2db739723

Observation ed797390-a567-404b-af25-f504bd1c365f · inbound

The Distributed Detectability Band Against Marginal-Preserving Attacks cites this paper.

The Distributed Detectability Band Against Marginal-Preserving Attacks Reasoning Models Don't Always Say What They Think

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:57:41.849792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T12:58:22.056355Z digest=sha256:75ebea4b0dcdd2d920ae9c262ea206fe8596d356aa92341c0c87b286ef3bba70

Observation f4fc79be-0cac-4a79-9e02-5734febcfbbf · inbound

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models cites this paper.

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models Reasoning Models Don't Always Say What They Think

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T11:18:03.796570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:36:05.700067Z digest=sha256:205a00195a3315ecc73dd99540c61d07de839006db0d8e8020da6d7a396a8adb

Observation 60250f88-3172-43ba-bf75-dd6341c2c6f1 · inbound

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation cites this paper.

Beyond Accuracy: Measuring Bias Acknowledgment in Chain-of-Thought Reasoning for Responsible AI Evaluation Reasoning Models Don't Always Say What They Think

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T16:58:43.404342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T04:43:04.740925Z digest=sha256:b8540ffd3935e9cb64880186ddf7d0b33a321b70308c74310ee4b9bd2eb723fa

Observation 8e9fc536-173e-4202-ad20-be7a769556ad · inbound

Know Your Limits : On the Faithfulness of LLMs as Solvers and Autoformalizers in Legal Reasoning cites this paper.

Know Your Limits : On the Faithfulness of LLMs as Solvers and Autoformalizers in Legal Reasoning Reasoning Models Don't Always Say What They Think

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T17:18:43.460462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T04:24:11.949884Z digest=sha256:b22315371925e29a7cef86a20183b4c324566012b020d13486c52cce84310924

Observation 166f1bd7-68e5-493a-8294-b0190d2ace62 · inbound

Analyzing the Narration Gap in LLM-Solver Loops cites this paper.

Analyzing the Narration Gap in LLM-Solver Loops Reasoning Models Don't Always Say What They Think

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T20:49:57.677402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T20:41:54.964459Z digest=sha256:09fc32ad4f3a935152bdfb1733057cbd857cebff1089bc5030f786c48d42fdb8

Observation e3156763-063a-496d-a562-cc7412994ac9 · inbound

Local Causal Attribution of Chain-of-Thought Reasoning cites this paper.

Local Causal Attribution of Chain-of-Thought Reasoning Reasoning Models Don't Always Say What They Think

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-26T12:49:29.099013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:6acf8edf16eaee92eadec375a39e18673fab9cce6c96bfb1ac1886ba3047fa49

Observation 1cc15de3-a976-4434-bcec-2e1b23359500 · inbound

Against Proxy Optimization cites this paper.

Against Proxy Optimization Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:49:46.634634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T08:26:50.595370Z digest=sha256:4115aacc5a7d0c73e25fe3f50ff65f9db5d4d0a1d9a8ba6c5701d2ef12220a92

Observation 1023fca4-0c2c-4ba5-bf14-f47e8ba7b344 · inbound

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation cites this paper.

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation Reasoning Models Don't Always Say What They Think

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:14:36.929953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:12:02.227976Z digest=sha256:e56f68d8eec1dcc82a44b6eb8c06750efb18caaddade958bd96d82ba2c5f7032

Observation 19ccb1b6-dec5-480c-a109-8a5ce6b01850 · inbound

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs cites this paper.

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs Reasoning Models Don't Always Say What They Think

Reference 133

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T16:25:50.038326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T00:38:21.949283Z digest=sha256:7b8e4d03b1243b3192795f50bcfea168c29db5c00f001df2e7e1113e2f5995fb

Observation bea6fe0e-0506-4155-ba9b-1b29eb4c9774 · inbound

Defeat Devices in AI Systems cites this paper.

Defeat Devices in AI Systems Reasoning Models Don't Always Say What They Think

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T08:44:27.899211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T08:34:58.879346Z digest=sha256:bd31173209a8b6f4d65488e631e965ec52a0c5a33058c388f62aa4287798c478

Observation cf4e6467-6a82-45aa-a571-d41bfbbb3270 · inbound

Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision cites this paper.

Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T10:35:42.816633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T05:18:30.190239Z digest=sha256:3f0e1b14f8c3ad25eb5c49f7d254996925b45556a881ba89293d91714a9de597

Observation 3c96079a-738b-489f-9f4e-b7300b8b2eb0 · inbound

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates cites this paper.

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates Reasoning Models Don't Always Say What They Think

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:46:56.222938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T12:46:40.700728Z digest=sha256:998681ed8a505fe883516b15533ba05fcee51c239655e88c93d1f7ae6695658d

Observation 632a689a-da3f-445c-bf59-005be81801ac · inbound

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens cites this paper.

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens Reasoning Models Don't Always Say What They Think

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T02:02:49.492856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:02:49.492856Z digest=sha256:68824598229e9468335a58aef9b688bc84f5ddc392effbee08a5f41095df826e

Observation 80f741cc-f051-4447-8d5a-6a40f24b9ea1 · inbound

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets cites this paper.

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets Reasoning Models Don't Always Say What They Think

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T11:57:03.259382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T11:54:32.051780Z digest=sha256:ce4c5b904fa6b6912ebcbc49d081ad100acf5cbb3274e09a19e0d73320b91e91

Observation f7628fda-e636-4363-ae46-64d25abe2a80 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Reasoning Models Don't Always Say What They Think

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T15:45:54.532529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:45:54.532529Z digest=sha256:569746ed83ff44500b9b368716b10e64358943aed659680fa972ee395c13d6e9

Observation f25cadd4-831e-4196-aafc-d85e71457bf1 · inbound

Length Penalties Make Chain-of-Thought Less Monitorable cites this paper.

Length Penalties Make Chain-of-Thought Less Monitorable Reasoning Models Don't Always Say What They Think

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T08:06:10.773565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:06:10.773565Z digest=sha256:b75fd9daacd7f23f7733af12012d5bf502b327bf523b7b29c5fe1b6364312476