Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:26.477946Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2506.05451.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:26.477946Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:42.313410Z
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dbf922d8-bb63-4726-ace1-fb395551418c · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refusal in Language Models Is Mediated by a Single Direction
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 062a7351-50e3-44b0-a0e1-ff7da1d591b2 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58ac760-0795-4097-8005-5aaff2c6130e · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Latent Knowledge in Language Models Without Supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a510d602-2812-4b6e-ad08-998edc02fac3 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e97702f2-a65c-42fc-8bb2-dd2a8b388c03 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Nitay Calderon and Roi Reichart
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8ea653f-a2ff-40a9-9f0f-2877412ba779 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Improving Steering Vectors by Targeting Sparse Autoencoder Features
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9e87b34-a8e6-4d4d-a985-b455905c92b1 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Finetuning Language Models to Emit Linguistic Expressions of Uncertainty
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08dd8f6-c46d-439f-8826-c1fd52b48c8c · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Faithful Reasoning Using Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f977c197-fd2d-41d8-aa68-130cebbea1dc · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Eleventh International Conference on Learning Rep- resentations
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07d74fd0-9172-4674-a7fe-6404471bb4ba · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Variable Binding Circuitry with Desiderata
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c72c0a66-3748-4fb3-b5e9-3db3bf864d17 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Studying Large Language Model Generalization with Influence Functions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2711a5e4-750b-4a0e-abe1-ae7a0bf9a485 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55a3048b-7a0b-43f1-926f-4d7fd1a3dd6c · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d010205d-0dd2-4b85-9ba4-9d27f86083e7 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3488b3e3-dca5-4d5e-8340-3e9f76278872 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Alon Jacovi and Yoav Goldberg
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86483a8-da5a-482b-b580-4a6b03accbec · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271a24eb-f94f-46b3-b5ae-aaff77699d9a · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 237d5644-9cb1-4701-bc63-1347536b7e4c · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9c0bf9b-88c6-4aae-b144-8b6267082701 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Twelfth International Conference on Learning Repre- sentations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6532def3-7402-47d1-89d6-aeee8c336770 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Rethinking Explainability as a Dialogue: A Practitioner's Perspective
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0a5476d-6ef2-4063-9cb9-57abc06527db · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d704216-f085-43b4-ad3d-bf434267d1ac · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ef54eb-60ab-41e0-90dd-9dd62454179e · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f2c7cf-d0a3-41ba-bb9f-906cba85e495 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2e084d-6a75-4237-9fca-33dbbf571323 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SaRO: Enhancing LLM Safety through Reasoning-based Alignment
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 686d0779-98ad-4775-9f85-46cf60b7591e · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbeb1e1e-31e7-498f-af5f-ca4aa00efc3b · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Model Refusal with Sparse Autoencoders
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08cddfeb-c7d8-4643-b109-c0b071d77632 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f08816-1d81-41f5-87dd-077d24031298 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25099a16-f626-442c-8027-75aa87f62c6f · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InICLR 2025 Workshop on Building Trust in Language Models and Applications
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18f8d42c-9149-455b-a29f-653f71b8df80 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Models With Activation Engineering
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7093f4e-d86c-4039-9e92-9ea347a812a3 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6077e1dc-c411-4eb3-8de1-7e23de6a6069 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d632adf3-891d-4c43-9da3-7e8ddee5ac65 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2c1e6d-62a1-48f1-99ca-61df1f20c819 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety why should i trust you?
Reference 483
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99d8ab01-3987-47fe-a284-cd5793d76918 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55ba416a-3575-4f8b-982a-9f11b3f56127 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety What do you learn from context? Probing for sentence structure in contextualized word representations
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9acd6c1-29b8-4742-8f4c-bda55b6e9ef2 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2848e556-47e8-4bc3-8182-b773c9fa1e82 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1215c80c-f397-4304-89cd-3a2e21d794a6 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety AtP*: An efficient and scalable method for localizing LLM behaviour to components
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1c746e-ede6-4330-bb77-3f4bf6832ede · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebd8f26f-5dc9-4d52-9ac6-bef8b3335c81 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Get my drift? Catching LLM Task Drift with Activation Deltas
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 560521ff-a5b7-43da-a648-05155e0ef7b0 · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e69ad70-b6de-42ab-8d6e-6831fa30682f · outbound
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aaquib Syed, Can Rager, and Arthur Conmy
Reference 3328
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0ba13cc-7577-4cd3-9ca2-2dfa84908394 · inbound
Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
Reference 2000
Source-reported events for the cited work
Unavailable: canonical work link unavailable.