Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:07.861843Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 9 inbound Pith citation observations for arXiv:2505.14352.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:40:07.861843Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T22:20:51.989267Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 1f0d0b3f-737f-484a-85b1-811e08771f24 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511f238a-e3aa-4478-8dce-1dd08fed1d9a · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Tell me about yourself: LLMs are aware of their learned behaviors
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b876a4f-14aa-499c-9a3a-99deede786ab · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceccbf25-6bf5-4eaf-b38f-82f0527d3524 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability E., Hume, T., Carter, S., Henighan, T., and Olah, C
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce6d09c6-23cc-41fb-a75f-90b362471e25 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Eliciting latent knowledge: How to tell if your eyes deceive you, 2021
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ab3fae9-90a7-4b18-935f-485cfe3fc634 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad1cdce-3136-4de4-8d8a-f201ef897d63 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9020be08-2811-4fb1-9c81-85823fbfd8f9 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Safe RLHF : Safe reinforcement learning from human feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f1fc3d7-4b82-437e-9b8f-da650e9b9df9 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Qlora: Efficient finetuning of quantized llms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8419faf-5ad4-4c82-9089-e10388bd00af · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Pal: Program-aided language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f78732ec-7c97-41db-acb6-5b6f3c15ed75 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 flash, 2025 a
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5f075c8-433c-4927-8594-ea3ab04015d6 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemini 2.5 pro preview, 2025 b
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5635376-a01a-460b-8ad0-e949039a0f6c · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Alignment faking in large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc2b0c15-274a-432d-a95d-c421c32a8667 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dc4628e-ae13-40c4-a7fa-3fa88ab731d5 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability and Lee, S
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 16de2527-ffb6-46a0-801c-457ab90bc719 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability u chemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., G \
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 480279d4-bbad-4752-9cec-e4b7594543c9 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability M., Bommarito, M
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13023214-b0bf-4312-be60-747afb101da2 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4523999c-11ea-4675-8629-f91213f4ebe6 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Auditing language models for hidden objectives
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac39182b-c4bf-4d4d-9319-78b89bf67f87 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Frontier Models are Capable of In-context Scheming
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68671958-3532-4304-b1c0-16ed2186ead8 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability interpreting gpt: the logit lens
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation efdf8ba3-4011-479b-b770-f283ec91767b · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Learning to reason with llms, 2024b
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a02bb6a-e6a2-418c-8618-b60bfa998db6 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6bfe05-6f84-46d8-91cd-3a1ca2a3bbbe · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability D., Ermon, S., and Finn, C
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 49d393c5-a973-41d8-ae91-34bfd73cfeda · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a6f0fd4-9efd-4e5a-82cd-e68e69c925ce · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Top 1000 english nouns, 2019
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87906699-feee-48b8-9fe9-90d16caa5973 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Large language models can strategically deceive their users when put under pressure
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12d7963d-ac67-4556-8d4c-d04a4424e592 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e8a314b-57e8-4e7d-923c-ea400b0faca4 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6ee307d4-e2d9-4355-ad8b-597a447b6ec0 · outbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability N., Kaiser, ., and Polosukhin, I
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4bd3ba3-9c28-48ce-82c3-e71a7d58af84 · inbound
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e6064b-8eca-414d-913f-f9b0a427f6ea · inbound
DECOR: Auditing LLM Deception via Information Manipulation Theory Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71882747-c6a5-4999-bdb2-b6a74edf1618 · inbound
Building Better Activation Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf2dbefc-e764-4a38-9b8d-6660d1c8a799 · inbound
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a44a5147-ccd4-4129-a339-7880f3ca094a · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a8b2a06-b7ed-42ab-b5fc-883db645c7ee · inbound
"Don't Say It!": Constraints, Compliance, and Communication when Language Models Play Taboo Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9cd5712-4558-422c-9c58-77fc742c2630 · inbound
The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76e1f9be-c3ec-4f0e-9a94-c8040b4a08fe · inbound
MUX: Continuous Reasoning via Multiplexed Tokens Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4e3907-474c-4909-9ba9-029b28afc797 · inbound
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.