Pith. sign in

Paper Citation Record · LEDGER

Preventing Language Models From Hiding Their Reasoning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.18512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18512 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:41:22.141140Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 06a3266a-2cee-4b34-8223-3dcf28092312 · inbound

Alignment faking in large language models cites this paper.

Alignment faking in large language models Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T22:50:11.946185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T22:50:11.846863Z digest=sha256:300bf422892a015a84b7058aeae8b9b0f95a484eff21ea1d6dc7511467f10234

Observation 42aceb86-a0bb-481f-b238-dfaa9b454102 · inbound

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking cites this paper.

MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking Preventing Language Models From Hiding Their Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T16:41:22.141140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:41:22.141140Z digest=sha256:1d6653f858c82d5e3ff2a6f4af7b4850416f94c7356b112b3dece43217eb12ea

Observation 67be5f1f-89f9-49de-b9de-97969602ee1a · inbound

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning cites this paper.

Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:54.383501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:03:54.383501Z digest=sha256:ccdb291e57845b08c86e9bc86249233475e37979181f3741d8b05088b10af527

Observation 31c74176-d014-4c06-b0de-dc539a7e1980 · inbound

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors cites this paper.

When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors Preventing Language Models From Hiding Their Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:55.843136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:55.843136Z digest=sha256:d9a8219f6ee84328077cba4bc162008c8c8a0e56c7a0601699455ff75c1443c0

Observation 1bf448f2-6dcd-4830-905a-704b7d8ece0f · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Preventing Language Models From Hiding Their Reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:19:44.761298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:f453548b1d0f6edf0ca3581ef6fca93a8e856e2e49ac7a2b7f4ebfd1fd98be47

Observation 1f21787f-b906-4a30-bddd-cd0a03fb9630 · inbound

Diagnosing Pathological Chain-of-Thought in Reasoning Models cites this paper.

Diagnosing Pathological Chain-of-Thought in Reasoning Models Preventing Language Models From Hiding Their Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:26:30.628642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:26:30.628642Z digest=sha256:06e0f52f5fa1b8620884f4925469d62d9007d03ebde933f9e510504e0cecc83b

Observation 3d71b3f4-33e8-4484-be57-83039ffb90f6 · inbound

Detecting and Suppressing Reward Hacking with Gradient Fingerprints cites this paper.

Detecting and Suppressing Reward Hacking with Gradient Fingerprints Preventing Language Models From Hiding Their Reasoning

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:22:37.572076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T08:18:37.665350Z digest=sha256:d1044380ded8bfc84a92441eea1f32ebe988a87562820b58b118192444eefd9c

Observation bc293fa8-4e7f-4c12-b12d-7002267f0c87 · inbound

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel cites this paper.

When Reasoning Traces Become Performative: Step-Level Evidence that Chain-of-Thought Is an Imperfect Oversight Channel Preventing Language Models From Hiding Their Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.393171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:30:12.558660Z digest=sha256:ec8f6e5ce1d39ed600e1b6b8c1ff544949b3086188d100248931258c8f0e82f8

Observation 03f5e166-a971-4774-ad45-9f8e85f5c4a0 · inbound

Understanding and Mitigating Premature Confidence for Better LLM Reasoning cites this paper.

Understanding and Mitigating Premature Confidence for Better LLM Reasoning Preventing Language Models From Hiding Their Reasoning

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T14:04:44.385114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T14:03:25.913615Z digest=sha256:cb7610a0f5fa3292515e4c77e61a880cc571055886c6ae82ef211652c6b79320

Observation f752b928-caa1-4e30-a480-7f9f7a324065 · inbound

Conceptual Steganography cites this paper.

Conceptual Steganography Preventing Language Models From Hiding Their Reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.621461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T18:46:39.761975Z digest=sha256:b82bdbfc644049b48e8de5b38d40ebae82967a8ce6ce124fe1417e79688e560f

Observation 4ad341bd-0be5-4de4-80fe-cc441e544fac · inbound

A Note on the Strategic Confinement Problem cites this paper.

A Note on the Strategic Confinement Problem Preventing Language Models From Hiding Their Reasoning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:31:06.842944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T17:30:12.165074Z digest=sha256:d3409d10680a1ed76899b8ce41ab0e739fc7983b3cc282e8b328237721f21c60

Observation 122c54e7-21d8-478a-b498-880e09ae5697 · inbound

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs cites this paper.

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs Preventing Language Models From Hiding Their Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.423665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T13:23:28.061924Z digest=sha256:65798ad412efd02526890153f10cf3f595c86616a2ee8987cbdfe964bff2e9e8

Observation e8a0d71d-82ff-46f5-ac74-7566df5fb004 · inbound

Comparing Linear Probes with Mahalanobis Cosine Similarity cites this paper.

Comparing Linear Probes with Mahalanobis Cosine Similarity Preventing Language Models From Hiding Their Reasoning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:18.923298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T20:40:03.169825Z digest=sha256:077e2963d430fd83936a9d5c650863f720eeb1ba4de8d4a163915a055a11402d

Observation 0b3af246-dbce-44ae-b803-3e0a0600999d · inbound

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs cites this paper.

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs Preventing Language Models From Hiding Their Reasoning

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:50.102094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T00:38:21.949283Z digest=sha256:2310a4cc2c3028bf84bb35c8adc9bb99e1f3e7074072b931189ba4aa7c214bf0

Observation 3fd1c6e4-9836-428f-8115-7be3bbfda935 · inbound

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation cites this paper.

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation Preventing Language Models From Hiding Their Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:50.791477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:29:50.791477Z digest=sha256:2beb1f3049dc9353609ef67e9d00b7780d78faedfae1ed14aec641633662be8d

Observation 9df42a77-34b4-467c-8d69-28d57c845de5 · inbound

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations cites this paper.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations Preventing Language Models From Hiding Their Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:59.970123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:02:59.970123Z digest=sha256:4ae1abe9d76e8c84493593c90159aed613540f043cd7a8f7b9936413b6484a76

Observation c7dbb06f-cfdd-481b-a3dc-6ebe56957df8 · inbound

Not All LLM Reasoning is Visible in the Chain-of-Thought cites this paper.

Not All LLM Reasoning is Visible in the Chain-of-Thought Preventing Language Models From Hiding Their Reasoning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T04:14:45.406064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:14:45.406064Z digest=sha256:7fb2967d33701ece7561457ffe7a99804e13d5f3f48d1975d131761f84c797b1