Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2402.13457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.13457 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:06.503637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T19:36:08.540565Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 96816ab0-ea0e-4e03-b2d8-bd1b244f6442 · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 201

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:47:56.160539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:513db32469397b714932bab7d606d77a47a98bc4124ec66a85e0391940ff35cc

Observation 4ec1c18d-206d-45a6-ad1b-233c7e70c5d4 · inbound

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models cites this paper.

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:06.503637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:06.503637Z digest=sha256:73d0a2d75ed0117ff684d5de306a07f75a66adc33060bb9a316227fd7cb46c01

Observation c3576dbb-ce6f-4b77-ad91-e23a35fc89f5 · inbound

Understanding the Supply Chain and Risks of Large Language Model Applications cites this paper.

Understanding the Supply Chain and Risks of Large Language Model Applications A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.912963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.912963Z digest=sha256:04da23f6f53bb73789535429be76117abc6552143df9a6eafae71332bd75885a

Observation 25f5f323-7136-4f16-b93b-e827df8a1a61 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.361494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.361494Z digest=sha256:7c57a091bc91f9745eff806a63163deb95194d920abc424edeacf01da26a4984

Observation 9c20c9f8-8f70-4da7-a200-a906231ac9c9 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:45.398903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:45.398903Z digest=sha256:1fe5e71cf50189680462c7f9e9da0fb9215eb210ffabd7d2398fc8bb937facef

Observation 5f48ed78-a3b3-4fbc-8a64-39d31cbf9d26 · inbound

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal cites this paper.

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:31:54.591943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:27:52.438709Z digest=sha256:3ea03a3054ec2c94064683eb4194f7476c4b1afcb612288b30f63c6e02c64a1d

Observation ead34801-dcd9-434d-b7cf-10580f06e3e0 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.435223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.435223Z digest=sha256:fd60f268891c8f87d589c59a554417e51095b6a4e8902b3d03d893eaab5e2f7f

Observation aa9e6cc6-b697-400f-9c2e-fd2873f19520 · inbound

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm cites this paper.

Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-04T22:33:28.168620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:33:28.168620Z digest=sha256:ef579e708da7173649fa122e9d881765bd6a5b768b3a0c26a269b9eb3224e88c

Observation f8a6ee22-da7a-4884-9f88-a3c118714fda · inbound

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying cites this paper.

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:45:57.609447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:57:59.918483Z digest=sha256:844a7cc96f7401cea1639c549fe5882b4456c92913694bde6c4f40d9c7eea638

Observation ca2742c6-95f2-45f1-8ee9-319bba9a33bb · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.374380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:3ed590a7cb67f204dad909849eabc72bea679cf525a092909f649b6a61e41712

Observation e46e7c62-8614-458f-9e27-838cb4dc1f6d · inbound

Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models cites this paper.

Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:08.542160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T22:23:00.894059Z digest=sha256:9b53e8c1acccad899691c6b5278f81b765ff85e6d2c244f827663f0b0ad70773