Pith. sign in

Paper Citation Record · LEDGER

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2506.05451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05451 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:24:26.477946Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:45:42.313410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact2
  • verified fuzzy10
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dbf922d8-bb63-4726-ace1-fb395551418c · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.305781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.305781Z digest=sha256:e9783c30c01d23b750a2d0ee5b57713bc7cdc7eb5268ba4edccd644c2bc40eba

Observation 062a7351-50e3-44b0-a0e1-ff7da1d591b2 · outbound

This paper cites Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.309978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.309978Z digest=sha256:f5c444d7aea0d1371ee0bc19e8a8bdcbf676db3120ea954085cf6ecac1a559fa

Observation b58ac760-0795-4097-8005-5aaff2c6130e · outbound

This paper cites Discovering Latent Knowledge in Language Models Without Supervision.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Latent Knowledge in Language Models Without Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.314508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.314508Z digest=sha256:6dc83eaa0cefc56ef04941c6873f205ed2bf98c762b315af622debd78dd333f7

Observation a510d602-2812-4b6e-ad08-998edc02fac3 · outbound

This paper cites V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety V ojtech Cahlik, Rodrigo Alves, and Pavel Kordik

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.266765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.318627Z digest=sha256:d6e4c07b2278dd35fbdb3b6fa9f5a325a66eba47056350682391931a4f8760ae

Observation e97702f2-a65c-42fc-8bb2-dd2a8b388c03 · outbound

This paper cites Nitay Calderon and Roi Reichart.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Nitay Calderon and Roi Reichart

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:27.077942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.323532Z digest=sha256:c688f2ba7ffbb13bebe8566fa5a63279cc0ba43105c0577ed86c306d085f6933

Observation d8ea653f-a2ff-40a9-9f0f-2877412ba779 · outbound

This paper cites Improving Steering Vectors by Targeting Sparse Autoencoder Features.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Improving Steering Vectors by Targeting Sparse Autoencoder Features

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.327706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.327706Z digest=sha256:5a86fe174db3af134d3260c76699683e666d6faa126c34916fdb781e67640e2a

Observation c9e87b34-a8e6-4d4d-a985-b455905c92b1 · outbound

This paper cites Finetuning Language Models to Emit Linguistic Expressions of Uncertainty.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Finetuning Language Models to Emit Linguistic Expressions of Uncertainty

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.332694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.332694Z digest=sha256:9c6afa7795778fb01e06e076c73500b0b63020ab4eef0090704d9d3a8cbeef07

Observation a08dd8f6-c46d-439f-8826-c1fd52b48c8c · outbound

This paper cites Faithful Reasoning Using Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Faithful Reasoning Using Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.340928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.340928Z digest=sha256:8648a0c9810bd8e45280bf7e6b7186e830d394ec4469aa37268c437e62c5c454

Observation f977c197-fd2d-41d8-aa68-130cebbea1dc · outbound

This paper cites InThe Eleventh International Conference on Learning Rep- resentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Eleventh International Conference on Learning Rep- resentations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.345165Z digest=sha256:1650e19dec0e9a48b283215218b490d349b124a080db5587a965d565f2fd813c

Observation 07d74fd0-9172-4674-a7fe-6404471bb4ba · outbound

This paper cites Discovering Variable Binding Circuitry with Desiderata.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Discovering Variable Binding Circuitry with Desiderata

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.349316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.349316Z digest=sha256:8407b7a8fb595f80b75ad7f7bd6b4c113a71ddc888657f7d03fc3c59da6ecb17

Observation c72c0a66-3748-4fb3-b5e9-3db3bf864d17 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Studying Large Language Model Generalization with Influence Functions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.357474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.357474Z digest=sha256:c06aa66fd7b4ac78e585491dc15612c739178fe70fa02971d285c1aa5a04a420

Observation 2711a5e4-750b-4a0e-abe1-ae7a0bf9a485 · outbound

This paper cites InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th Interna- tional Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.220441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.366075Z digest=sha256:18fcd4b48004058c02fd2223a2f948a3cf1eacf0c26b814e522d273ad62562d4

Observation 55a3048b-7a0b-43f1-926f-4d7fd1a3dd6c · outbound

This paper cites Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.370726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.370726Z digest=sha256:d3164b9896ce8961ee6abd0750e541feadec2e6558e0381362d8e61f91d2dfaa

Observation d010205d-0dd2-4b85-9ba4-9d27f86083e7 · outbound

This paper cites Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Zhengfu He, Xuyang Ge, Qiong Tang, Tianxiang Sun, Qinyuan Cheng, and Xipeng Qiu

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.375257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.375257Z digest=sha256:c9a1e1c71d82be0a5d6d85faecd1973a4cd5f61b29135cb0698e806fdcb795e0

Observation 3488b3e3-dca5-4d5e-8340-3e9f76278872 · outbound

This paper cites Alon Jacovi and Yoav Goldberg.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Alon Jacovi and Yoav Goldberg

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.380068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.380068Z digest=sha256:5a5d210e98397b0719b608eb4fc6e0bae3d805f1f86c388a332d0a378ef44a0e

Observation d86483a8-da5a-482b-b580-4a6b03accbec · outbound

This paper cites How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.384700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.384700Z digest=sha256:81d66f43219ba11672c56690a978ac7984f846e3d2d29c48b1927bb70a735f1e

Observation 271a24eb-f94f-46b3-b5ae-aaff77699d9a · outbound

This paper cites InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY , USA

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.209641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.392317Z digest=sha256:9842c2ec552d4abf7dd861152a91c1424b91931b382b7f9c02ef8bd3c6a9d1ef

Observation 237d5644-9cb1-4701-bc63-1347536b7e4c · outbound

This paper cites InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InProceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 3: System Demonstra- tions), pages 42–50, Toronto, Canada

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.196551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.396707Z digest=sha256:9871405b9a9ab11ecf295038fc552f8b42f3b5fe8dbcadf2f1fdc15bf075c749

Observation d9c0bf9b-88c6-4aae-b144-8b6267082701 · outbound

This paper cites InThe Twelfth International Conference on Learning Repre- sentations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InThe Twelfth International Conference on Learning Repre- sentations

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.184327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.400384Z digest=sha256:19a6374d53242ffbb128d402712082bb9be8293cde0ae11a096d5449e4ee9bcb

Observation 6532def3-7402-47d1-89d6-aeee8c336770 · outbound

This paper cites Rethinking Explainability as a Dialogue: A Practitioner's Perspective.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Rethinking Explainability as a Dialogue: A Practitioner's Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.403709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.403709Z digest=sha256:085b36cb971d8204f52e3b5edf06f56ca0ab89f45fba46cb4bdf3be8712da0b5

Observation f0a5476d-6ef2-4063-9cb9-57abc06527db · outbound

This paper cites HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety HuDEx: Integrating Hallucination Detection and Explainability for Enhancing the Reliability of LLM responses

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.766381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.408324Z digest=sha256:cec37e3e4d75abb3001a934f4dc3d3d0dc31977ffb5e40e90e4c8e6a92d3abc2

Observation 9d704216-f085-43b4-ad3d-bf434267d1ac · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.412268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.412268Z digest=sha256:661e7d3681c404b4639d145354b3fa34ec7e48b988024c0fe76e3c0eac40baff

Observation d9ef54eb-60ab-41e0-90dd-9dd62454179e · outbound

This paper cites Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.415954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.415954Z digest=sha256:e1039ba3354435322036de3d91e1ca290cf2750c6c44cb5ef9aae707310dd74a

Observation b7f2c7cf-d0a3-41ba-bb9f-906cba85e495 · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.419920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.419920Z digest=sha256:583c80a55938b95fdfc217f8eef1b3937406902bdbb1e31ad7f3e8fd6aa5e394

Observation de2e084d-6a75-4237-9fca-33dbbf571323 · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.427788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.427788Z digest=sha256:7c4dd4e2efffa1956b6f56a6dd2d209b9af1a99ed7298a239da2abce65beb525

Observation 686d0779-98ad-4775-9f85-46cf60b7591e · outbound

This paper cites Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Gradient-Based Automated Iterative Recovery for Parameter-Efficient Tuning

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:26.697247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.431523Z digest=sha256:57be370dfda613dc88bb6096c9e20635e05af85ad2a23292abe98362725cdeeb

Observation cbeb1e1e-31e7-498f-af5f-ca4aa00efc3b · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Model Refusal with Sparse Autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.435845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.435845Z digest=sha256:2f0d38bf36abddb3419586809597df1991a25e91e4af38745bdd1ae868191071

Observation 08cddfeb-c7d8-4643-b109-c0b071d77632 · outbound

This paper cites Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Refining Input Guardrails: Enhancing LLM-as-a-Judge Efficiency Through Chain-of-Thought Fine-Tuning and Alignment

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.440712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.440712Z digest=sha256:f9b66dfe65012ea13aa9c40e9f2f6e86864c4a924d0909b535e96cc7f365e9c9

Observation d2f08816-1d81-41f5-87dd-077d24031298 · outbound

This paper cites RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.452708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.452708Z digest=sha256:70ec56598ceb8c07a6e2a5e90e8740894f3f8f0596522c32b086fde13500953e

Observation 25099a16-f626-442c-8027-75aa87f62c6f · outbound

This paper cites InICLR 2025 Workshop on Building Trust in Language Models and Applications.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InICLR 2025 Workshop on Building Trust in Language Models and Applications

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.460867Z digest=sha256:f6ade45e8d5ce774e0448ac18b2bac31b20756e93e92896c2f959aab90b5397d

Observation 18f8d42c-9149-455b-a29f-653f71b8df80 · outbound

This paper cites Steering Language Models With Activation Engineering.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Models With Activation Engineering

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.465201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.465201Z digest=sha256:149fd816cd6716977e0e4b92a56631449bef054f7485f4a8d089fe54ac359bf3

Observation b7093f4e-d86c-4039-9e92-9ea347a812a3 · outbound

This paper cites an unresolved cited work.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:24:27.148900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.469980Z digest=sha256:dc0101afd17e2bac3839234017a77c49d1988baf6167c207783af8f20d3d1425

Observation 6077e1dc-c411-4eb3-8de1-7e23de6a6069 · outbound

This paper cites Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.474285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.474285Z digest=sha256:2294f125d3bb996e93a365f9a12b4af467c57b64ce382e702a35e4e33d3796a4

Observation d632adf3-891d-4c43-9da3-7e8ddee5ac65 · outbound

This paper cites Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.477946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.477946Z digest=sha256:f2ad52b3073c4f841a68c54d73df9296b2186e02ba258d53f546f938030d1247

Observation fb2c1e6d-62a1-48f1-99ca-61df1f20c819 · outbound

This paper cites why should i trust you?.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety why should i trust you?

Reference 483

Resolution
verified exact
raw_fallback, observed 2026-08-07T10:24:26.660635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.444527Z digest=sha256:7befd3d925c7edb8517e2efa1d86d8a4a5d52d6240468c7ba8089c7bcd948e11

Observation 99d8ab01-3987-47fe-a284-cd5793d76918 · outbound

This paper cites Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.423582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.423582Z digest=sha256:729e3b76dcd78afbc55c99a7d2ef17f8f7f9ef72f6bef2da703f2f392a2714b0

Observation 55ba416a-3575-4f8b-982a-9f11b3f56127 · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.456815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.456815Z digest=sha256:47e9bf00a014451b6b3db56be24f7422f0a389ad0a7876e2629dfce72472e226

Observation d9acd6c1-29b8-4742-8f4c-bda55b6e9ef2 · outbound

This paper cites In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 5553– 5563, Online

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.232109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.362341Z digest=sha256:f061ea747028bb2afed6ebdc4f0e6f6d0ea0ad08ddafc9703c361dfad2d3d50d

Observation 2848e556-47e8-4bc3-8182-b773c9fa1e82 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.353163Z digest=sha256:efe58b3065f2969545d66427937fea5949789a5605f1dd5a7eb4fcaab04c515e

Observation 1215c80c-f397-4304-89cd-3a2e21d794a6 · outbound

This paper cites AtP*: An efficient and scalable method for localizing LLM behaviour to components.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety AtP*: An efficient and scalable method for localizing LLM behaviour to components

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.388763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.388763Z digest=sha256:6d4533b41e20d9433bd622c272e7441a53294af4ed9323e2f435a578fe9a7a42

Observation 2f1c746e-ede6-4330-bb77-3f4bf6832ede · outbound

This paper cites InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety InAdvances in Neural Infor- mation Processing Systems, volume 36, pages 16318– 16352

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.255379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.336515Z digest=sha256:9514451a444437651507ede9753acd3a9049cc9e5c81d449aca9dfb76aab0f29

Observation ebd8f26f-5dc9-4d52-9ac6-bef8b3335c81 · outbound

This paper cites Get my drift? Catching LLM Task Drift with Activation Deltas.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Get my drift? Catching LLM Task Drift with Activation Deltas

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.300328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.300328Z digest=sha256:faba3a03223640aba04a0a7feb6c5eed0aa55fc6cb3eaa3cd017e0eeb13047ea

Observation 560521ff-a5b7-43da-a648-05155e0ef7b0 · outbound

This paper cites SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety SAFE: A Sparse Autoencoder-Based Framework for Robust Query Enrichment and Hallucination Mitigation in LLMs

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:24:27.136657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.294848Z digest=sha256:446a5c42ea195591600eff7e1da4b3a8c456930444f8d97f57726b17c94755e3

Observation 9e69ad70-b6de-42ab-8d6e-6831fa30682f · outbound

This paper cites Aaquib Syed, Can Rager, and Arthur Conmy.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Aaquib Syed, Can Rager, and Arthur Conmy

Reference 3328

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:24:27.172301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:24:26.448608Z digest=sha256:4fe9527e2631f717d840e14e2948f4a247b654cd0b1a5a0536803c98d3a93c33

Pith citing papers

Observation c0ba13cc-7577-4cd3-9ca2-2dfa84908394 · inbound

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models cites this paper.

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-03T05:45:42.313410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:45:42.313410Z digest=sha256:9fc6ad5af9819c55879a1d408df479e0d927a2d7b3f367bd19c765dc518d9715