Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:22:55.728237Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2608.02089.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:22:55.728237Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a373dd73-b609-4f09-97fd-80d1fbab6cd3 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8199497-cb67-4eed-a4ba-2fb9fa508bdd · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba035bc8-4993-4dff-ac0f-280eacbe58e1 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Monitoring monitorability
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d9de79-eb3e-4151-8588-b0e7e667031c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Learning to reason with LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00dd6d45-59b4-434d-bdba-bdad35b6d4d9 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Reasoning models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fc9aca7-83dd-4247-8d85-cc1216172457 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Gemini thinking
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2f15e35-9ecf-4171-b629-6090ef04fc9e · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Extended thinking
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee19f08-34c6-4dcc-902d-64bca9bf7e06 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Detecting misbehavior in frontier reasoning models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a84790b-1427-421b-8fef-1e05d624f26f · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Evaluating chain-of-thought monitorability
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d31e4bc-575d-4e23-ae8b-65a98f9702fe · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Weinberger
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3830b88-790d-46b6-a444-51592cf03256 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Language Models (Mostly) Know What They Know
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 670ec74b-dca9-47c8-a921-a4b1b4dc5314 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Discovering latent knowledge in language models without supervision
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a60e5fd-6ff8-4cf1-82f8-481d7bc715e7 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models The internal state of an LLM knows when it’s lying
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8de3072-45d2-49d5-bdb2-b18bc8598df4 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Qwen3 model collection
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbf82d27-7fd1-43e1-a317-6855fc7398c7 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Qwen3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c905e8-16e3-4985-928c-ab192a798ba0 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Introducing gpt-oss
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fca9c4e1-5f2f-4d9f-b351-103c57a10434 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models gpt-oss-120b & gpt-oss-20b Model Card
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39e291d4-6ebd-4712-9ce7-047c3036de32 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models GPQA: A graduate-level Google-proof Q&A benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d9e42c-34f4-49ef-a96f-d119cd9e6672 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models MMLU-Pro: A more robust and challenging multi-task language understand- ing benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26bead1-62ce-4b4b-9afc-07b0f2c4c903 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Benchmarks saturate when the model gets smarter than the judge.arXiv preprint arXiv:2601.19532, 2026
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69bd3f39-8981-4b20-963f-77b5d753461c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Le, and Denny Zhou
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9285d41e-9d1e-4e76-9931-60e05bacd976 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Large language models are zero-shot reasoners
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd54062b-bbea-4263-b1a7-d638b3788517 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Training Verifiers to Solve Math Word Problems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4590fb20-cd6a-4419-9b69-45f05f9bd1b8 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Let’s verify step by step
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b3ff758-276c-4af8-addb-b5a0e8efe634 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models The relationship between reasoning and performance in large language models–o3 (mini) thinks harder, not longer.Scientific Reports, 16, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe8b4e2-2a48-4e3a-b026-7ade65457f15 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Measuring Faithfulness in Chain-of-Thought Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c34a3c-1df0-497c-82f2-e95c5ac0b4f9 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d986b2-c8e4-4cf2-8906-e615e3486f3c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Faithful chain-of-thought reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 891febc3-c56d-4222-8264-e34320446d0f · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Reasoning Models Don't Always Say What They Think
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd74589-00ab-48b7-abdc-77fd4de9211d · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models MonitorBench: A comprehensive benchmark for chain-of-thought monitorability in large language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20b9d831-1796-4841-996c-999c335b3812 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Recent frontier models are reward hacking
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c240b39-bb8d-4940-955f-43cd49f44316 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Zimmermann, David K
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0faacaec-9e86-4592-84f2-9481b9a78dd7 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models ReasonOps: Operator Segmentation for LLM Reasoning Traces
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c431f89-a143-491d-b303-ff92ff1eb2f6 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Length Penalties Make Chain-of-Thought Less Monitorable
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ced02226-d5e1-4daf-9070-b8fb59ddb0a8 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models How to Steal Reasoning Without Reasoning Traces
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27b77f84-96c7-4573-a65a-68ff4c96956a · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Measuring weak-to-strong legibility of reasoning models,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e51d651-eda2-4133-ac24-a4090bbc8243 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Selective classification for deep neural networks
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bbd4f15-5e8e-4dde-b0cd-b9aee3327416 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Selective question answering under domain shift
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9137fcb8-3be6-43bf-8ec2-a59b6810b193 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ecc528-7dfe-4465-a29e-dcb142d04643 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Post-abstention: Towards reliably re-attempting the abstained instances in QA
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c21d660-8403-4a1e-9477-6584cfbadd2a · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Understanding intermediate layers using linear classifier probes
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5b3e427-3d92-40ff-857b-e043099550bd · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c7eac5-c1a8-4b8c-866e-0237616497d1 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Probing classifiers: Promises, shortcomings, and advances.Computational Linguistics, 48(1): 207–219, 2022
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51cf9d20-99a4-4bfa-856b-4a1a76cf83e3 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models The geometry of truth: Emergent linear structure in large language model representations of true/false datasets
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dbfacce-095a-4e25-a322-d302c24fd408 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Representation Engineering: A Top-Down Approach to AI Transparency
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fc31e30-c431-4330-9f14-5a07614571ab · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 128f845d-abca-4eb0-8cff-572e2d44cd27 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Catching rationalization in the act: Detecting motivated reasoning before and after CoT via activation probing
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 553a984d-8060-496d-b3fb-d9e8d817bdaf · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Can LLMs predict their own failures? self-awareness via internal circuits,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb32995-7a31-4b13-bee8-d60196feb79b · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Unresolved cited work
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa0ea57-4e87-4d6d-ab0a-7fa9a9957c7a · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Lexical Hints of Accuracy in LLM Reasoning Chains
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f50feb57-7cfa-40dc-b940-c4f52dfef437 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d4a926-5ff9-411e-8ea0-cd4757ef639c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Cohere’s embed models: details and application
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e09c6ee7-226e-4fb6-9f28-dd3a366f1009 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Bag of tricks for efficient text classification
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b31ef1b5-be68-4428-9359-8161787a2b71 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Efficient Few-Shot Learning Without Prompts
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 023b3a8f-f40a-4cfa-b264-9767a7fd109c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Ma- tryoshka representation learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f456af3-86a9-40a0-a709-5e0595d9c956 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Accessed 2026-05-12
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a3d63aa-70c5-43bc-afc0-142966bf912a · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Task prompt:
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc37a3c5-5168-4efe-9585-d90f0ea45894 · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Cohere Embed v4 model parameters for Amazon Bedrock.https://docs.aws.amazon
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73985422-df85-4de7-a118-8b29d80e8f3c · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Unresolved cited work
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b310046e-0b9c-4079-8694-d3b6bee995cd · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a502dba-f086-44ab-a5fb-40a1dea5224f · outbound
How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models Measuring Weak-to-Strong Legibility of Reasoning Models
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.