Pith. sign in

Paper Citation Record · LEDGER

Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.18166.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.18166 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:25:30.717668Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T11:53:03.533122Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96d8eb54-9b08-4fa5-a3d1-5330150667ef · inbound

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation cites this paper.

JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:30.717668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:30.717668Z digest=sha256:bc05845590e76d3bbc3d62f8a6e685ffd8fe2fcea64e6c14653404ba56a250cd

Observation 968a4e6f-0bb3-495f-b052-a572df396cbc · inbound

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors cites this paper.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.763079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.763079Z digest=sha256:3425549532271c646a5c4da144cac78d15de2cd69e65ca0d2ad8b07191a56a07

Observation 011b081c-b633-4b13-b08b-47f08cc71277 · inbound

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace cites this paper.

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:34.599908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:34.599908Z digest=sha256:b7ad3e8269e9c8a1a0b649515c133245253bad649f3466bb2379bb64346cfe0e

Observation 3e9739da-bfe8-407d-8fbb-0b0a745ba207 · inbound

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment cites this paper.

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:53:03.534832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T11:52:36.688263Z digest=sha256:32c044be79d20258a9ed9dac67251378123fb7a2f70465ab7fbadba6165bbfc8

Observation 3b6d11f3-dffd-46e9-91d9-fd8f4ce197b0 · inbound

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles cites this paper.

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:39.775290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:39.775290Z digest=sha256:3456b4115f6d1914c7c924dd082a551b253c7cbc88c3a3d35d6dbcae3941b0e3

Observation 4e9e3dab-f6bc-4b05-83ed-17188a1e8e3f · inbound

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring cites this paper.

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.440660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:28:41.134253Z digest=sha256:d442d80379cafababbcf3b8b5a375b059a7eefb167b449f51b7b209a2b951c9e

Observation 7d3e94ee-1bdb-4344-8ae2-4535f55d15db · inbound

Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation cites this paper.

Silencing the Guardrails: Inference-Time Jailbreaking via Dynamic Contextual Representation Ablation Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:05:58.322946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:48:09.607231Z digest=sha256:95791adf42acdf33e9c49045d4048469e5268f116927568d9e4edc6358cae2a2

Observation 106f1c68-28ca-4dcc-9156-33c3e458f114 · inbound

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types cites this paper.

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:31:00.365870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:08:25.471462Z digest=sha256:42546263f2732548424b105360f3515444b352a931b6019850daa4d7bad779e2

Observation 7645e6d1-bd18-4654-b0f6-c8e614faa764 · inbound

LLM Safety From Within: Detecting Harmful Content with Internal Representations cites this paper.

LLM Safety From Within: Detecting Harmful Content with Internal Representations Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:20:23.143551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T04:33:54.058475Z digest=sha256:7c1be13e43ec3abb356127528ea7da1ec4169bfe2fd40c31fa33d54a1c79dedc