Pith. sign in

Paper Citation Record · LEDGER

Concept-Level Explainability for Auditing & Steering LLM Responses

As of 20 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:2505.07610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07610 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:17:07.402209Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:51.818562Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T09:05:36.555847Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bac0d01e-ce84-4d53-b949-dfc76ceef72b · outbound

This paper cites Hello gpt-4o.

Concept-Level Explainability for Auditing & Steering LLM Responses Hello gpt-4o

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.497141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.094844Z digest=sha256:bb721482266b1017d0fbd3ee1d7b78ce5621c41694a1ff6e9515f74d3834d8a6

Observation 45fcea3c-3621-4852-b92e-68671c83e7eb · outbound

This paper cites The dark side of generative artificial intelligence: A critical analysis of controversies and risks of chatgpt.

Concept-Level Explainability for Auditing & Steering LLM Responses The dark side of generative artificial intelligence: A critical analysis of controversies and risks of chatgpt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.483311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.099884Z digest=sha256:164ab978e89493f6797b630235fab684fbabd054c290bea104d20de3b84a4e14

Observation 86f892e6-3403-496a-b12a-3a745283dbc5 · outbound

This paper cites Survey of hallucination in natural language generation.

Concept-Level Explainability for Auditing & Steering LLM Responses Survey of hallucination in natural language generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.104508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.104508Z digest=sha256:03bf1fa879195ad114e448fbff97cea1e232aa87b87ace9ce049a966049a973e

Observation 318b843c-b369-4074-b837-e9db22e0b3a1 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023.

Concept-Level Explainability for Auditing & Steering LLM Responses Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.108993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.108993Z digest=sha256:cd5db7ef3234d83488759e7fa039b4b12056157bec3e96fe57dc79805fe9d0be

Observation 8eba732d-fa67-4ab9-a1b8-b3e9acf0c758 · outbound

This paper cites Spear Phishing With Large Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses Spear Phishing With Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.113423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.113423Z digest=sha256:be8e3a07dd543da6fb10a8015468854a03ffd3f84b5a8df6ec2fe55f91086221

Observation b41396e7-6318-438e-9f90-ab13c566e92a · outbound

This paper cites Training language models to follow instructions with human feedback.

Concept-Level Explainability for Auditing & Steering LLM Responses Training language models to follow instructions with human feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.118226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.118226Z digest=sha256:cc07a089829d58dc8df7ec0125f7fc30a2dc74404034dded7745f09d08238220

Observation ddf86fb3-7fa5-4abb-9ecb-0072a0c96075 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Concept-Level Explainability for Auditing & Steering LLM Responses Constitutional AI: Harmlessness from AI Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.123143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.123143Z digest=sha256:13278938754ff136577f257165f4acdc68c3695b20e55c634f513f1b20901264

Observation d2b17460-a338-4259-9cac-7056b4ccb6a6 · outbound

This paper cites Pretraining language models with human preferences.

Concept-Level Explainability for Auditing & Steering LLM Responses Pretraining language models with human preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.127814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.127814Z digest=sha256:dd2899894fcd881516b3d5eb5aa2b38ed7649bcbe89e22eee0ea52a3fb21e57e

Observation 8dafeab6-0b5a-4734-b201-e5ca3d6c52c7 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Concept-Level Explainability for Auditing & Steering LLM Responses Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.131947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.131947Z digest=sha256:b282a81edbd4fbb41ec2edddebc2a122e86defd69a6d198843c474fea219ee0a

Observation 4e11a20c-13d9-497c-a289-773f62e206fa · outbound

This paper cites Can LLM-Generated Misinformation Be Detected?.

Concept-Level Explainability for Auditing & Steering LLM Responses Can LLM-Generated Misinformation Be Detected?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.136572Z digest=sha256:cd7a7ca71c45d334ade32f2f927583e97e5a229dcdbc942da3a0a798bd515b7d

Observation 3501c2d3-7799-4708-8c87-99a33c64db35 · outbound

This paper cites Ai model gpt-3 (dis) informs us better than humans.

Concept-Level Explainability for Auditing & Steering LLM Responses Ai model gpt-3 (dis) informs us better than humans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.434478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.141622Z digest=sha256:ea7af6b513e9cb4a3d8c76b4e7135a544a0d48318140a5455ad3fec09a2288e5

Observation 761c1b7a-e4c5-473d-b0ea-ac4b882f353b · outbound

This paper cites The operational risks of ai in large-scale biological attacks.

Concept-Level Explainability for Auditing & Steering LLM Responses The operational risks of ai in large-scale biological attacks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.420698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.145941Z digest=sha256:12a4bef902c3d059a91da238f9a3758fc82ddc3288cea7086acf2103d9fb5603

Observation 064faa6d-e629-497f-b57b-b5bd301cce07 · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.150065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.150065Z digest=sha256:99b95f8e4781bfb2dfb79c3f8cae35953e50a8a75732fe1289ee2499b9b391a1

Observation 35081a60-27a7-4c0e-a018-57a1a893c672 · outbound

This paper cites LLM Agents can Autonomously Hack Websites.

Concept-Level Explainability for Auditing & Steering LLM Responses LLM Agents can Autonomously Hack Websites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.154425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.154425Z digest=sha256:2bb7deea436e6c9818fff478626821f377bcc8c503462186a539ddae993c79ce

Observation dd4b3541-5924-427a-9a8c-f7905a998f12 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned llms.

Concept-Level Explainability for Auditing & Steering LLM Responses Emergent misalignment: Narrow finetuning can produce broadly misaligned llms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.159214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.159214Z digest=sha256:7f20964f883da89a9abae60def5506c3090fe2aeec842f9a8205c2a2b9c44f01

Observation 96908015-f375-4ffd-8af3-81ab4d62194d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.163342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.163342Z digest=sha256:3de4ec5e1c91f6a875967d3111abad5f32bc496900991167eaa3bba54d775664

Observation b18ba5bb-3b5e-4ae5-bb6e-335684b6a272 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Concept-Level Explainability for Auditing & Steering LLM Responses Frontier Models are Capable of In-context Scheming

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.168159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.168159Z digest=sha256:8925e965dfd1307286c5bf9b5b7d876e3d95344f0df263d035336ddb1f7b0580

Observation f4b6a6ec-0b17-4d14-b23a-7911b5a49ecc · outbound

This paper cites Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era.

Concept-Level Explainability for Auditing & Steering LLM Responses Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.172594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.172594Z digest=sha256:e1f36555486ec7b88721b0e4fee0648cd495460c0b5acc6b701e24a51f6a3971

Observation fa40994f-4b69-4860-a2c4-87a7da7808d0 · outbound

This paper cites TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation.

Concept-Level Explainability for Auditing & Steering LLM Responses TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.178029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.178029Z digest=sha256:58a86ea0d40e797b79efa3f6563e41ab231d0a9fbe1a75a38ff985ecb9d3bf60

Observation fa737317-87a5-41a2-9b0a-2edc2231b49b · outbound

This paper cites SyntaxShap: Syntax-aware Explainability Method for Text Generation.

Concept-Level Explainability for Auditing & Steering LLM Responses SyntaxShap: Syntax-aware Explainability Method for Text Generation

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:17:07.854631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.182475Z digest=sha256:ad9a802b685e16d88a64990d33214f99d14dc0608973d5e9395e918c610deb69

Observation f4d51d47-8053-46e0-a1dc-9d533c3f2ac1 · outbound

This paper cites Investigating the impact of linguistic errors of prompts on llm accuracy.

Concept-Level Explainability for Auditing & Steering LLM Responses Investigating the impact of linguistic errors of prompts on llm accuracy

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.406370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.186969Z digest=sha256:67fd9303127817523242af4836e6ddcb95cd2a5001f6f2a3de908e4855aaaa39

Observation 908a955f-fc69-4801-ad0c-99db80f5b14c · outbound

This paper cites Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection.

Concept-Level Explainability for Auditing & Steering LLM Responses Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.191242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.191242Z digest=sha256:e274ab1db647d474bd76d23bd58ed240942db7040aff410a42ba06fa41a17b9f

Observation 5e0e68d1-fc97-44d9-a6f9-3a8cb601876f · outbound

This paper cites Conceptnet 5.5: An open multilingual graph of general knowledge.

Concept-Level Explainability for Auditing & Steering LLM Responses Conceptnet 5.5: An open multilingual graph of general knowledge

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.195619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.195619Z digest=sha256:7ddcb5c3695cdfd187c89bebe955aa1975dc41b3ab3694bd3d01180293ee10a8

Observation e10233de-dd21-4122-b1f6-fb870fa5464a · outbound

This paper cites Hashimoto.

Concept-Level Explainability for Auditing & Steering LLM Responses Hashimoto

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.200529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.200529Z digest=sha256:6c89bf0dc690f15e5630ce7a91eb58a8145ad13d9c4f9ba48575b6b50e7d473e

Observation c6397b1b-1d9b-471a-8ad0-ec57d9acdc38 · outbound

This paper cites Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM.

Concept-Level Explainability for Auditing & Steering LLM Responses Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.204518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.204518Z digest=sha256:f8e1feb48c8450918dbf986b2c553f79baa5bd78502ebc1d04ba80f037b725f3

Observation 8af50549-af41-4b6f-9a53-8f2a4e562276 · outbound

This paper cites Definitions, methods, and applications in interpretable machine learning.

Concept-Level Explainability for Auditing & Steering LLM Responses Definitions, methods, and applications in interpretable machine learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.208956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.208956Z digest=sha256:e0a1e96d0456014bfc46970ec44fdaf710d5d69967434784f90f8c19374e919c

Observation a628282d-7156-4ad7-b020-88ce0aa63fba · outbound

This paper cites Techniques for interpretable machine learning.

Concept-Level Explainability for Auditing & Steering LLM Responses Techniques for interpretable machine learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.364846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.213197Z digest=sha256:b8451171c979366f4a4db05573c191462e6ed9c62dadad0f1cddb86c3caad8e8

Observation 0a6ef210-56d7-4e90-840b-5a2182d8bbaa · outbound

This paper cites A survey of the state of explainable AI for natural language processing.

Concept-Level Explainability for Auditing & Steering LLM Responses A survey of the state of explainable AI for natural language processing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.350187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.217300Z digest=sha256:155f45833e846324f5fb022e3a68aefb8bfbfd1d91d0e17b428a72e89d521c18

Observation a0cd394b-65d2-415c-ad01-510b20f8c705 · outbound

This paper cites On the explainability of natural language processing deep models.

Concept-Level Explainability for Auditing & Steering LLM Responses On the explainability of natural language processing deep models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.335689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.221493Z digest=sha256:a4f4ef3ccc5d5331dae8377c3ab3eea9e6fe40377e44fe676a17a2aa9614bf58

Observation 13c4ece8-690a-4a0f-ace5-89c65da0bf65 · outbound

This paper cites Interpretability in activation space analysis of transformers: A focused survey.

Concept-Level Explainability for Auditing & Steering LLM Responses Interpretability in activation space analysis of transformers: A focused survey

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.321415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.225547Z digest=sha256:d37cd6a5611755e0a5f109270405150c97d5a73ec504af75570c6a618f99db3c

Observation 788db6e4-600c-4c9f-90b9-725ba6436a76 · outbound

This paper cites Neuron-level Interpretation of Deep NLP Models: A Survey.

Concept-Level Explainability for Auditing & Steering LLM Responses Neuron-level Interpretation of Deep NLP Models: A Survey

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.306137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.229673Z digest=sha256:e5ee338783102035b0ef9ab5605325964398dac207b24c4f19e717923194594a

Observation 0f800d80-f707-4fa9-9093-9e90f89f4ffb · outbound

This paper cites A value for n-person games.

Concept-Level Explainability for Auditing & Steering LLM Responses A value for n-person games

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.292011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.233734Z digest=sha256:a00e18d276cf362c41760d8a1213ab046d560f1ded3cf3d6ab79fccd2c7e0a2d

Observation 90792b78-b420-45d0-8099-b4a7e118abf5 · outbound

This paper cites why should i trust you?.

Concept-Level Explainability for Auditing & Steering LLM Responses why should i trust you?

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.278054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.238591Z digest=sha256:2aac17405d26ce0082136b32c2ba8496df0af02ec5b701157bd88e0f156e0fb6

Observation b6712469-0d7d-4b62-8a1c-c31bcfbb1190 · outbound

This paper cites BERT meets shapley: Extending SHAP explanations to transformer-based classifiers.

Concept-Level Explainability for Auditing & Steering LLM Responses BERT meets shapley: Extending SHAP explanations to transformer-based classifiers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.264337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.242845Z digest=sha256:978b0ec4131c546d39d15220edd7a2defe5fd843e49c96e88fb6b0c8ec289670

Observation c0e73b65-4983-4c97-b7d1-3db6797eef80 · outbound

This paper cites Generating hierarchical explanations on text classification via feature interaction detection.

Concept-Level Explainability for Auditing & Steering LLM Responses Generating hierarchical explanations on text classification via feature interaction detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.249633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.247715Z digest=sha256:069ac6a3ef743dcb5f16c49ad56651cfc71b9bcf10dc0698b21b881153a4cf20

Observation a1f66367-8446-4169-bada-fe4da118d417 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Concept-Level Explainability for Auditing & Steering LLM Responses Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.251939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.251939Z digest=sha256:753f40532a629cfcb7935cc41f0513707a0c97b2408d50bc072788b19e2947e2

Observation ee143436-0090-4d70-b195-7e983c05ef79 · outbound

This paper cites Explainability for large language models: A survey.

Concept-Level Explainability for Auditing & Steering LLM Responses Explainability for large language models: A survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.256221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.256221Z digest=sha256:8c25f2dcbeb40fcfa91548df85e1c08e4db280e3b208417ca912e9dd8190a06c

Observation d37153b2-4a86-4c66-a0bc-3534713d77c6 · outbound

This paper cites Ten levels of ai alignment difficulty.

Concept-Level Explainability for Auditing & Steering LLM Responses Ten levels of ai alignment difficulty

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.227087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.260674Z digest=sha256:ad29c8a76663ab0b07e03448ec420c5de5fdbc6d116f0273f3c57cbf69f6c629

Observation 409b987b-19a3-42f5-a77a-1d7e2f3d0b64 · outbound

This paper cites A Survey on Fairness in Large Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses A Survey on Fairness in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.264825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.264825Z digest=sha256:70e16199c371e94bce75a0469b3132cacfb361df131814d1f009d70e16743616

Observation b16f809f-dfdd-4179-a469-a33503eccbb8 · outbound

This paper cites Aligned probing: Relating toxic behavior and model internals.

Concept-Level Explainability for Auditing & Steering LLM Responses Aligned probing: Relating toxic behavior and model internals

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.269168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.269168Z digest=sha256:3dd98c7d11f890df40ef631d4802717394c8b5a80abf8d3cc436abd6446d2597

Observation 61508227-3c81-450d-a549-ea966a2f268a · outbound

This paper cites Toy Models of Superposition.

Concept-Level Explainability for Auditing & Steering LLM Responses Toy Models of Superposition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.273064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.273064Z digest=sha256:9821d191fb6d058f9d67c5bb0d296988286a74696c45022c8efc0cb9ef8eabd1

Observation 2db2240e-1800-45a3-aae5-9eca2783b067 · outbound

This paper cites Formalizing convergent instrumental goals.

Concept-Level Explainability for Auditing & Steering LLM Responses Formalizing convergent instrumental goals

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.213540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.278372Z digest=sha256:3810fd6cde97931ec59c40a9a800ac49597b0143fa646101e10e3f788cc26841

Observation 854b13ba-f2d2-49d3-90a5-3120df4fe382 · outbound

This paper cites Circumventing interpretability: How to defeat mind-readers.

Concept-Level Explainability for Auditing & Steering LLM Responses Circumventing interpretability: How to defeat mind-readers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.282747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.282747Z digest=sha256:4c4bd4b53c90bdd99991a13cb799c6d7d8dfddabe152e846f3efcf5f987a70e6

Observation 08c856c7-6096-4266-afe9-271843c6fa6b · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

Concept-Level Explainability for Auditing & Steering LLM Responses SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.287109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.287109Z digest=sha256:3a80afbe7e8e96227df1120e6bc23ef95060a32067e226acad17a6509f18e09e

Observation 68ef9ad3-5011-4a51-81b3-be73ea613aee · outbound

This paper cites Parafuzz: An interpretability-driven technique for detecting poisoned samples in nlp.

Concept-Level Explainability for Auditing & Steering LLM Responses Parafuzz: An interpretability-driven technique for detecting poisoned samples in nlp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.200116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.291512Z digest=sha256:6e07a697df33309ab0d153db6dcc961e3eae77e0eabf2e571ba24308e1e56f0d

Observation 50b6c9e1-5818-47ac-ab9c-55025625e684 · outbound

This paper cites SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner.

Concept-Level Explainability for Auditing & Steering LLM Responses SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.295771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.295771Z digest=sha256:7a2521f6ac813549ddafe7d3a6e75d106a5e907bccb0137a8e931698f0d14784

Observation dea027d8-1962-4ffb-91e0-58686bb29653 · outbound

This paper cites Protecting your llms with information bottleneck.

Concept-Level Explainability for Auditing & Steering LLM Responses Protecting your llms with information bottleneck

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.186282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.300843Z digest=sha256:248d4f6024c3a8a6318f0a6aedd3c1dd7e5e0c39ea5ce7fe4a21c76c4547d5b9

Observation f4d68c51-1582-43be-ae90-20658e656865 · outbound

This paper cites Defending LLMs against Jailbreaking Attacks via Backtranslation.

Concept-Level Explainability for Auditing & Steering LLM Responses Defending LLMs against Jailbreaking Attacks via Backtranslation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.304943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.304943Z digest=sha256:0449d47c5a91b816d5c165979c4becc09477a10aeada9eef84e9c0a9a8232d76

Observation b7369042-7366-4822-93c2-1b76075e8ef7 · outbound

This paper cites IMBERT: Making BERT Immune to Insertion-based Backdoor Attacks.

Concept-Level Explainability for Auditing & Steering LLM Responses IMBERT: Making BERT Immune to Insertion-based Backdoor Attacks

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:17:07.573180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.309449Z digest=sha256:b7def399cece08d9cb9732c6397009ed714aae968090bd862615d716b98760bd

Observation f9ab14bf-c74d-4e11-b9f9-c94e9643bec2 · outbound

This paper cites Defending against Insertion-based Textual Backdoor Attacks via Attribution.

Concept-Level Explainability for Auditing & Steering LLM Responses Defending against Insertion-based Textual Backdoor Attacks via Attribution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.314042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.314042Z digest=sha256:4a6a33c496eac9a4a4109cd3bcb472c5a421dd58ea31f31dcd25d79c786cbb11

Observation b2c2490a-07fd-459f-8340-96c717217984 · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

Concept-Level Explainability for Auditing & Steering LLM Responses LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.318596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.318596Z digest=sha256:c951564a79ebeed92c21c722e074035ec751633d24e80537a47ab5d7364987fc

Observation e27ea07c-a248-4f61-b078-a5a85af29c3e · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses Multilingual Jailbreak Challenges in Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.323068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.323068Z digest=sha256:0ded29dad09d02b5192aa28aa25c464aa4e00823a7d748b97e8765843e63e236

Observation 78efed91-2622-40ac-8f34-ebea2440af36 · outbound

This paper cites Defending chatgpt against jailbreak attack via self-reminders.

Concept-Level Explainability for Auditing & Steering LLM Responses Defending chatgpt against jailbreak attack via self-reminders

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.327313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.327313Z digest=sha256:314b64407c5ea938c8cf2eafdb6d757254448c7b3037cd0b76eacb44aaf2c383

Observation 82cf7bc5-488a-405d-b066-6d3dacc68af0 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Concept-Level Explainability for Auditing & Steering LLM Responses Scaling and evaluating sparse autoencoders

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.331441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.331441Z digest=sha256:9980d80bfba9af150cb3d591eab371a7837eb12328c772c2472b7d2b689c7b71

Observation a0420b3f-9101-4fcf-a6f5-742953987e98 · outbound

This paper cites Inference- time intervention: Eliciting truthful answers from a language model.

Concept-Level Explainability for Auditing & Steering LLM Responses Inference- time intervention: Eliciting truthful answers from a language model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.335878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.335878Z digest=sha256:0f498263d0939fd0a56a90030dea4074bd8900dd756bb4471272c13519fb1c79

Observation 46704d75-9f44-4dba-94d8-1eec70010c58 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Concept-Level Explainability for Auditing & Steering LLM Responses Towards monosemanticity: Decomposing language models with dictionary learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.339995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.339995Z digest=sha256:a4c7e8dbc3165fef0bef053169b3c76aeb2d5bcf7d7d19e8cbc482b582bb1109

Observation cedf41c0-b192-402b-a851-4a00ae89419d · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.344144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.344144Z digest=sha256:844e10ca00dfd1a0f76327001d27d029f813a56eaf70a11aab535043794fa878

Observation 650c80e1-3e86-4142-b986-ce49f4974040 · outbound

This paper cites spacy: Industrial- strength natural language processing in python.

Concept-Level Explainability for Auditing & Steering LLM Responses spacy: Industrial- strength natural language processing in python

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.146288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.348717Z digest=sha256:ce5c860e1c61f2be5aebf1953ca8578726caf096752a4e806e364a8060694116

Observation 3e971645-5a0c-496d-ab41-69d6d18c3a65 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

Concept-Level Explainability for Auditing & Steering LLM Responses Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.352934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.352934Z digest=sha256:97081286cbf3ee2ea505d155421d1570ef2ae998fd9d8e015b743b795e64d31a

Observation 967d5286-2ecc-49fa-a25c-ff77cf19efbd · outbound

This paper cites an unresolved cited work.

Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.357912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.357912Z digest=sha256:3c9223cb703ce68587b84896719fe993bd73ef94d21d1fbe4e608c2bd6997162

Observation 15ea726d-0cda-4914-adc9-30f5637d5d20 · outbound

This paper cites an unresolved cited work.

Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.362031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.362031Z digest=sha256:eb51ce20ea7563c71f5ebfb682b6c9b864782a803f5003a7daf387a4b157bd39

Observation 2592c22d-c44c-4271-a395-5defbe1026f9 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 2024.

Concept-Level Explainability for Auditing & Steering LLM Responses Gpt-4o mini: advancing cost-efficient intelligence, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.103285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.366222Z digest=sha256:cc07c4640d230f805f8dcbc7cfe632a405edd0e4459e33815115d498f81594b1

Observation 5ce51501-ff3e-4b52-8415-992c55a7502f · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Concept-Level Explainability for Auditing & Steering LLM Responses Manning, Andrew Ng, and Christopher Potts

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.089392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.370196Z digest=sha256:484e09bbcb7aa7c04a4cdec87bca6f7a19778c489d676365b603f863cbd05447

Observation e9c8d542-69a1-4c08-bcc2-37d198805965 · outbound

This paper cites Analyzing sentiment polarity reduction in news presentation through contextual perturbation and large language models.

Concept-Level Explainability for Auditing & Steering LLM Responses Analyzing sentiment polarity reduction in news presentation through contextual perturbation and large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.075044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.375418Z digest=sha256:bbcd54b83590cc46522df9e08235cedcd0276ad6e2284ff9fa67ca12751e04da

Observation b9f7cb44-8f69-4289-a243-29596037956f · outbound

This paper cites Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders.

Concept-Level Explainability for Auditing & Steering LLM Responses Interpreting and Steering LLMs with Mutual Information-based Explanations on Sparse Autoencoders

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.379461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.379461Z digest=sha256:333d793cfe4abeb4b9565de0c9eec9716884e0dd04bf231c8114c38c8b6c8c03

Observation e76e2382-4f89-44d7-bde0-4974df77e60f · outbound

This paper cites SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models.

Concept-Level Explainability for Auditing & Steering LLM Responses SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.384006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.384006Z digest=sha256:6c58db7bc0865859964c40a2b66e4e3ede8fe8c273c7978bd6e87336eed9b441

Observation 02a15849-097a-4bd7-9424-8167c6e7fb6e · outbound

This paper cites an unresolved cited work.

Concept-Level Explainability for Auditing & Steering LLM Responses Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-15T22:17:08.060462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.388911Z digest=sha256:c9cdf3d248130ae0af28b2a0c532e66690eea1b048fe8bcb01f49ce88a42d024

Observation 467cd83a-bdb6-480a-97bd-f5e86ec6082c · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Concept-Level Explainability for Auditing & Steering LLM Responses Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.393168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.393168Z digest=sha256:42f6a2d6bd4a8bfabbfad5377ffc395849da129867b731e6bb77249c206690fa

Observation 2c25210f-1a26-44eb-b958-4200a95032de · outbound

This paper cites Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.

Concept-Level Explainability for Auditing & Steering LLM Responses Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:17:08.046077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T22:17:07.397708Z digest=sha256:b5288be26d529246dd218bf1e2a73195a28b78477576107a04a6394ca8b49806

Observation 81a50a37-21e6-4d3d-bd7d-35e67e9fedbc · outbound

This paper cites The Llama 3 Herd of Models.

Concept-Level Explainability for Auditing & Steering LLM Responses The Llama 3 Herd of Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:07.402209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:17:07.402209Z digest=sha256:4c99190876f531895f451d1e4a021a2c5ad01a05a47369c8b1652c3987b4a056

Pith citing papers

Observation fb72a985-5ae2-4caf-bdc6-5473b80e50bf · inbound

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP cites this paper.

Can Global XAI Methods Reveal Injected Behaviours in LLMs? SHAP vs Rule Extraction vs RuleSHAP Concept-Level Explainability for Auditing & Steering LLM Responses

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:51.818562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:51.818562Z digest=sha256:565523854b44035c222a2704eb238cc020268da12e8e6ee8a3dfbd29886b66eb

Observation 9769ed81-c868-4115-a94f-e940e0998d1b · inbound

MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion cites this paper.

MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion Concept-Level Explainability for Auditing & Steering LLM Responses

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:33:25.666132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:33:25.666132Z digest=sha256:5d3b449b40f8a3ba43126323f1a232f6ae25df3908397ea6391f894af12fede2

Observation 8e9168fb-7653-4a6e-847c-0fd7ce1785c6 · inbound

Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures cites this paper.

Investigating Linguistic Steering: An Analysis of Adjectival Effects Across Large Language Model Architectures Concept-Level Explainability for Auditing & Steering LLM Responses

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:05:36.558082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T09:01:34.038992Z digest=sha256:61f6d825a8cd3a02fdd462f2bc22d1c3b32aca6f5d4bd6a8f22ce8def8ddc3cd