Pith. sign in

Paper Citation Record · LEDGER

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2411.12701.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.12701 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:20:47.964154Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1eb6100-59c3-4b36-98b8-2cc0e7f6c2af · outbound

This paper cites Understanding intermediate layers using linear classifier probes.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Understanding intermediate layers using linear classifier probes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.789293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.789293Z digest=sha256:79f4fb16db86c1c1932c65812225db69f51f917236d7bc8467ae50942a7adfaf

Observation f178567c-0611-4702-a76d-4077d70ba11e · outbound

This paper cites Eliciting Latent Predictions from Transformers with the Tuned Lens.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Eliciting Latent Predictions from Transformers with the Tuned Lens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.793842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.793842Z digest=sha256:793b865427cb055f8734e670a5dd8f7bcc05f6870406876bf0ec32659bc05c76

Observation 27c905c8-6ad6-47fb-bc92-6c226359ecde · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.798094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.798094Z digest=sha256:a9765a06cb0530f1797c0c5d09d52b09de180199d55df3667ab3752d3d625a11

Observation 637fa55d-1d2e-45ba-8087-bb5981de2d11 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.802316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.802316Z digest=sha256:c4ee26a5c1359b83d03c13716c138197cddc6a3243aa1d674eb992077c77d676

Observation 53106c5f-ed7d-47c4-8f45-8f0b456594a3 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.431158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.805946Z digest=sha256:bad303fc0f0b637b40953fea4d74bba57aadf8a37ed4c6e140fec273fc7b9382

Observation 7e84de85-9ca3-4042-b9c6-24a429937849 · outbound

This paper cites Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.809937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.809937Z digest=sha256:cd3e86a3b77567dd182ab1195c2f884cfe7d866247867827ce17dbed1449c13f

Observation d3d0e54c-e0b5-4894-b46f-9b10b8d31218 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.419496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.814235Z digest=sha256:b87213c948386b10e15631af844618baa82b85ef2e488f18dc08a5fd96b74be1

Observation c36ca47f-55d3-468f-98c9-1f6a47776e4f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.817894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.817894Z digest=sha256:f3b48d5e2ba3096867b1b372c2dfaf00c935f56ec6e332980d0f58892c939f02

Observation aca82476-07ea-4f10-bc94-08849207c958 · outbound

This paper cites Triggerless Backdoor Attack for NLP Tasks with Clean Labels.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Triggerless Backdoor Attack for NLP Tasks with Clean Labels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.821833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.821833Z digest=sha256:815e2b48381eec70e21b8cc102ce4bf0bad59111cf21a16c2d6123d363265b36

Observation fadf73e0-49b4-453e-8f93-68d04a6cc602 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.826004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.826004Z digest=sha256:0adf43698bb4f7253b1e742cfd3937df389e5f24c3f41137a87f3e9a430211ad

Observation 3d89cbbe-6302-4101-b536-5a4952ea89be · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.400936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.829829Z digest=sha256:cced51dcaf07dceb8f9029a28e1814cc4590eaee10b288133349a2bd4629331c

Observation 186b9804-67c8-4b61-ac49-354c8b15a793 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.388665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.833786Z digest=sha256:def02a26218bfd68ee47c6030e470fb3f6b65c01fec8b137ae08932e70765b1d

Observation 325ab4dd-55eb-461d-9f05-2c73828076b5 · outbound

This paper cites Weight Poisoning Attacks on Pre-trained Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Weight Poisoning Attacks on Pre-trained Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.837338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.837338Z digest=sha256:72fca08467d849003e600a61dc044b067445d0ac63dd866cd8b826b5e53b7e4d

Observation 102dd777-73dd-40ee-abd7-3faa44deb3e8 · outbound

This paper cites BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.841557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.841557Z digest=sha256:c589d62be9a3d0fe9e1b00e84e8425891c56a8de674921e99e84a83c2c35195b

Observation 8a6dc7e3-981d-44e7-8413-5d21954d5860 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.377489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.845489Z digest=sha256:ddda78161edc29ef757a14fb855cad3d6135355c18e4a4fbbe0c4b8d212ef898

Observation 1d46b3f7-bf71-44a1-8d75-00b08198cff8 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.364962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.849280Z digest=sha256:2fbe6763e24b4192c7ee87120c768273a0d9417a67852ae78903af3aef402f97

Observation 581408bd-5971-44b6-9da1-418806bb18b6 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.853081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.853081Z digest=sha256:87c3ad66bcac90a73bd6ad4e4f2ab312afb18f19942bb5d135ab3c8e74f39c81

Observation 267bf8d8-1047-4a5f-89be-98078480cae2 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.345482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.856909Z digest=sha256:a81520af254b3ea43df0a7f8b26956fa21063f13ba06e9dad5f2f03d88b30350

Observation dd48b56c-81b7-4730-a59f-60255ab08680 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.334735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.861122Z digest=sha256:54077666b56af649c69aa6f3ec0943fab95d1524a9889b426c28ff3463d93111

Observation 32301e36-bc1c-41bd-89f0-00daf1912226 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.321842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.864881Z digest=sha256:167c7b280cd1bd0a63abb109f6924e702150468a823687fbaf2f1bef0619f206

Observation b347b450-3b4f-4704-a495-d75e18890df3 · outbound

This paper cites WT5?! Training Text-to-Text Models to Explain their Predictions.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations WT5?! Training Text-to-Text Models to Explain their Predictions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.868417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.868417Z digest=sha256:2db0a5a1fb98e8da26af1b706ed870a5e8b11e356801784c0199c3d1c11af897

Observation 48aaefce-1dec-4c5c-a9bc-06121786601c · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.309952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.872280Z digest=sha256:889615018f003e46f0939f5e920dd15fe45a306f091addd38c41f07a9df1a174

Observation 9ccc6ebd-8851-44d4-a026-ced087df5e60 · outbound

This paper cites GPT-4o System Card.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.876450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.876450Z digest=sha256:964e9ea4de013196f9dc3beb3f7518290b8e25e9c320d9ee7e6b0c4065b80e80

Observation d5ce55ae-c121-4750-91e1-6592de525067 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.298126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.880247Z digest=sha256:d4058ae8291d141fffacc830c0b204c23fdcb3d10964e89af468a4013ed9167e

Observation fc59da8b-0f4d-4237-95c6-d73502ddc931 · outbound

This paper cites Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.883792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.883792Z digest=sha256:004a82c14d025b39c70588fc0e4d81007543db442e6975ae1665ea66c09677ac

Observation 84750a0f-f502-48d2-995a-8129036479c5 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.286530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.887895Z digest=sha256:b046db332280449425dd09539116d227738dbd6b6640f05dae8d01df126b6154

Observation bfe9e137-fcb2-4fae-9a75-dd4f03e51e0d · outbound

This paper cites Explain Yourself! Leveraging Language Models for Commonsense Reasoning.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Explain Yourself! Leveraging Language Models for Commonsense Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.891761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.891761Z digest=sha256:a7a86810afaafa4e719c4ccebb3fed4c4b4d28d94a017bd736d626613be2cc16

Observation cabb687a-784b-4f96-b0a6-53419eabc550 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.274947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.896208Z digest=sha256:0de958658593e705bdfd7daec329d8f9d48cad51560106dfb0f650b056808bd4

Observation c6b94e1e-7282-434b-aa34-c34e4c7b6d8e · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Manning, Andrew Ng, and Christopher Potts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.900803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.900803Z digest=sha256:9250a39b224c6a8499256c7b01446d3277ae9665d2b156837a6a47d25104f34d

Observation 416d7128-de65-417d-8066-dffb55ce40ab · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.256275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.904878Z digest=sha256:5b33e508d627aee0da7df266be094c9f0a0bf88f97db01c8806330f50d9cc8b7

Observation 270d12a4-f5bf-4bad-9d1a-44d7d41ea3db · outbound

This paper cites Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.908765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.908765Z digest=sha256:4624895a3b0e8cbe9ffb2f67ec18464dca9f553d8c76ee25a3be42793258740b

Observation 5d758a62-24bc-4551-a0c3-1eb27c043fd2 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-12T17:20:48.243331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T17:20:47.913728Z digest=sha256:f04634913808d7898c1b1c6469a922051a1fa364cb604aec1816ca0e386ed931

Observation b37e0629-3f99-435c-9ab2-518a2d8f542a · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.917519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.917519Z digest=sha256:eff921b80f91b4e2960032ad7a3a30cd4c881bae68c79f9c152cae88c23af596

Observation 85933582-f49c-4345-a080-874ae0a43d74 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.921306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.921306Z digest=sha256:80e287fe2b9d3ce87a6ff4b1fb49688c3d7a6fe258f6f193f5ca0dc9c9df0b07

Observation 2179665d-6bd6-4ce5-8ef2-53b70a7e7ddc · outbound

This paper cites Concealed Data Poisoning Attacks on NLP Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Concealed Data Poisoning Attacks on NLP Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.925071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.925071Z digest=sha256:7f031ca2cc321cfe471d7a876e18378c8bc0c148d744aa4ae5b7f30f9946612b

Observation b0282666-d307-4c75-906f-0a5beda1308a · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.929218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.929218Z digest=sha256:08dc8551398cd8636c95970cb3716cdb84191294071ba0443242d1db63bb5751

Observation 63f11b28-45ac-4144-844b-1c0a5b001817 · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.932852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.932852Z digest=sha256:2e54a4cca5fec0c994c1d99bf8c53096fe7998cc2b31afb3cd5ad7c70a2114e8

Observation c4d7a69f-6c88-4de7-b71a-c5d8bd329d9a · outbound

This paper cites Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Usable XAI: 10 Strategies Towards Exploiting Explainability in the LLM Era

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.936979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.936979Z digest=sha256:6923aee7a6de1b9e7c491035ed2265c1e9de380a5785dc78edee76bc821fdb77

Observation 191b7ac9-7b76-4796-97dd-639d1dc7d929 · outbound

This paper cites Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.940866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.940866Z digest=sha256:d2236e3e850c23eb2e1d67d5783784096b3cc6c554d621ef268d26106b8d83c8

Observation 6eac4be4-6f81-44cd-a736-92449211f721 · outbound

This paper cites BITE: Textual Backdoor Attacks with Iterative Trigger Injection.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations BITE: Textual Backdoor Attacks with Iterative Trigger Injection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.944924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.944924Z digest=sha256:c4e88ea44c1221f9331fde290a858602b62c0dc37591ebef4cc0d0d896e3a877

Observation 1a5e70c6-d9b6-48fa-aceb-4e392b373cde · outbound

This paper cites The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations The Unreliability of Explanations in Few-shot Prompting for Textual Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.948928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.948928Z digest=sha256:052575169d2aef74c79aae90b8674fa04b292f710113004cde017a1aea303aca

Observation ae37f722-f861-4f1c-8a3f-6666fd1f5e07 · outbound

This paper cites an unresolved cited work.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.952756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.952756Z digest=sha256:5a4eeddeeceb82f16a7c4c6f0d6be3e8d70eb6f710c5caa9e2e3e1b342bb273f

Observation 4aa54054-1c99-4f75-ae2c-2baaaf0a278e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.956116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.956116Z digest=sha256:b575e7d82bee2d5ee1fa5447e7d636c016a544616276745f89d7ffc7c9326700

Observation 68f934ec-9df8-4eb0-96a6-1bab8c4dac12 · outbound

This paper cites online" 'onlinestring :=.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations online" 'onlinestring :=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.960168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.960168Z digest=sha256:79876208843633e800486a0f9d015ea8b0dce7abfb92bd9468d3fa6782827bae

Observation 6a609f6b-27fc-40a2-b100-184d6934f132 · outbound

This paper cites write newline.

When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:47.964154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:20:47.964154Z digest=sha256:202325bc860d77702e340d128ac4d7d4f3768d3f4b671c84dc69f7133a5b2e58

Pith citing papers

No inbound Pith citation observations are available.