Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:17:07.402209Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:2505.07610.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:17:07.402209Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:51.818562Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T09:05:36.555847Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bac0d01e-ce84-4d53-b949-dfc76ceef72b · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Hello gpt-4o
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 45fcea3c-3621-4852-b92e-68671c83e7eb · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses The dark side of generative artificial intelligence: A critical analysis of controversies and risks of chatgpt
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 86f892e6-3403-496a-b12a-3a745283dbc5 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Survey of hallucination in natural language generation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318b843c-b369-4074-b837-e9db22e0b3a1 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eba732d-fa67-4ab9-a1b8-b3e9acf0c758 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Spear Phishing With Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41396e7-6318-438e-9f90-ab13c566e92a · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Training language models to follow instructions with human feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf86fb3-7fa5-4abb-9ecb-0072a0c96075 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Constitutional AI: Harmlessness from AI Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2b17460-a338-4259-9cac-7056b4ccb6a6 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Pretraining language models with human preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dafeab6-0b5a-4734-b201-e5ca3d6c52c7 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e11a20c-13d9-497c-a289-773f62e206fa · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Can LLM-Generated Misinformation Be Detected?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3501c2d3-7799-4708-8c87-99a33c64db35 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Ai model gpt-3 (dis) informs us better than humans
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 761c1b7a-e4c5-473d-b0ea-ac4b882f353b · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses The operational risks of ai in large-scale biological attacks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 064faa6d-e629-497f-b57b-b5bd301cce07 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35081a60-27a7-4c0e-a018-57a1a893c672 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses LLM Agents can Autonomously Hack Websites
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd4b3541-5924-427a-9a8c-f7905a998f12 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96908015-f375-4ffd-8af3-81ab4d62194d · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18ba5bb-3b5e-4ae5-bb6e-335684b6a272 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Frontier Models are Capable of In-context Scheming
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b6a6ec-0b17-4d14-b23a-7911b5a49ecc · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa40994f-4b69-4860-a2c4-87a7da7808d0 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa737317-87a5-41a2-9b0a-2edc2231b49b · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses SyntaxShap: Syntax-aware Explainability Method for Text Generation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f4d51d47-8053-46e0-a1dc-9d533c3f2ac1 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Investigating the impact of linguistic errors of prompts on llm accuracy
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 908a955f-fc69-4801-ad0c-99db80f5b14c · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e0e68d1-fc97-44d9-a6f9-3a8cb601876f · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Conceptnet 5.5: An open multilingual graph of general knowledge
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e10233de-dd21-4122-b1f6-fb870fa5464a · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Hashimoto
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6397b1b-1d9b-471a-8ad0-ec57d9acdc38 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af50549-af41-4b6f-9a53-8f2a4e562276 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Definitions, methods, and applications in interpretable machine learning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a628282d-7156-4ad7-b020-88ce0aa63fba · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Techniques for interpretable machine learning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0a6ef210-56d7-4e90-840b-5a2182d8bbaa · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses A survey of the state of explainable AI for natural language processing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a0cd394b-65d2-415c-ad01-510b20f8c705 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses On the explainability of natural language processing deep models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13c4ece8-690a-4a0f-ace5-89c65da0bf65 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Interpretability in activation space analysis of transformers: A focused survey
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 788db6e4-600c-4c9f-90b9-725ba6436a76 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Neuron-level Interpretation of Deep NLP Models: A Survey
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0f800d80-f707-4fa9-9093-9e90f89f4ffb · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses A value for n-person games
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 90792b78-b420-45d0-8099-b4a7e118abf5 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses why should i trust you?
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b6712469-0d7d-4b62-8a1c-c31bcfbb1190 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses BERT meets shapley: Extending SHAP explanations to transformer-based classifiers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c0e73b65-4983-4c97-b7d1-3db6797eef80 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Generating hierarchical explanations on text classification via feature interaction detection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a1f66367-8446-4169-bada-fe4da118d417 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee143436-0090-4d70-b195-7e983c05ef79 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Explainability for large language models: A survey
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d37153b2-4a86-4c66-a0bc-3534713d77c6 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Ten levels of ai alignment difficulty
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 409b987b-19a3-42f5-a77a-1d7e2f3d0b64 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses A Survey on Fairness in Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b16f809f-dfdd-4179-a469-a33503eccbb8 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Aligned probing: Relating toxic behavior and model internals
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61508227-3c81-450d-a549-ea966a2f268a · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Toy Models of Superposition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2db2240e-1800-45a3-aae5-9eca2783b067 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Formalizing convergent instrumental goals
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 854b13ba-f2d2-49d3-90a5-3120df4fe382 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Circumventing interpretability: How to defeat mind-readers
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c856c7-6096-4266-afe9-271843c6fa6b · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ef9ad3-5011-4a51-81b3-be73ea613aee · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Parafuzz: An interpretability-driven technique for detecting poisoned samples in nlp
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50b6c9e1-5818-47ac-ab9c-55025625e684 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea027d8-1962-4ffb-91e0-58686bb29653 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Protecting your llms with information bottleneck
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f4d68c51-1582-43be-ae90-20658e656865 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Defending LLMs against Jailbreaking Attacks via Backtranslation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7369042-7366-4822-93c2-1b76075e8ef7 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses IMBERT: Making BERT Immune to Insertion-based Backdoor Attacks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f9ab14bf-c74d-4e11-b9f9-c94e9643bec2 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Defending against Insertion-based Textual Backdoor Attacks via Attribution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c2490a-07fd-459f-8340-96c717217984 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27ea07c-a248-4f61-b078-a5a85af29c3e · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Multilingual Jailbreak Challenges in Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78efed91-2622-40ac-8f34-ebea2440af36 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Defending chatgpt against jailbreak attack via self-reminders
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82cf7bc5-488a-405d-b066-6d3dacc68af0 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Scaling and evaluating sparse autoencoders
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0420b3f-9101-4fcf-a6f5-742953987e98 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Inference- time intervention: Eliciting truthful answers from a language model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46704d75-9f44-4dba-94d8-1eec70010c58 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Towards monosemanticity: Decomposing language models with dictionary learning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cedf41c0-b192-402b-a851-4a00ae89419d · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650c80e1-3e86-4142-b986-ce49f4974040 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses spacy: Industrial- strength natural language processing in python
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3e971645-5a0c-496d-ab41-69d6d18c3a65 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967d5286-2ecc-49fa-a25c-ff77cf19efbd · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ea726d-0cda-4914-adc9-30f5637d5d20 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2592c22d-c44c-4271-a395-5defbe1026f9 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Gpt-4o mini: advancing cost-efficient intelligence, 2024
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ce51501-ff3e-4b52-8415-992c55a7502f · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Manning, Andrew Ng, and Christopher Potts
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e9c8d542-69a1-4c08-bcc2-37d198805965 · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Analyzing sentiment polarity reduction in news presentation through contextual perturbation and large language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b9f7cb44-8f69-4289-a243-29596037956f · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e76e2382-4f89-44d7-bde0-4974df77e60f · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a15849-097a-4bd7-9424-8167c6e7fb6e · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 467cd83a-bdb6-480a-97bd-f5e86ec6082c · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c25210f-1a26-44eb-b958-4200a95032de · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81a50a37-21e6-4d3d-bd7d-35e67e9fedbc · outbound
Concept-Level Explainability for Auditing & Steering LLM Responses The Llama 3 Herd of Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb72a985-5ae2-4caf-bdc6-5473b80e50bf · inbound
Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP Concept-Level Explainability for Auditing & Steering LLM Responses
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9769ed81-c868-4115-a94f-e940e0998d1b · inbound
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion Concept-Level Explainability for Auditing & Steering LLM Responses
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9168fb-7653-4a6e-847c-0fd7ce1785c6 · inbound
Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures Concept-Level Explainability for Auditing & Steering LLM Responses
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.