Pith. sign in

Paper Citation Record · LEDGER

LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2408.15221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.15221 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:32:00.935942Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T17:30:00.600665Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 00f3009b-b943-4e6c-a81e-c9876dd4efe8 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.196729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:30dc6d71ddc4b0ed737490896e1d23ff2ce479fac20ccfdaadf9e5bf7b34e823

Observation dc7aef9d-0703-43cf-9d52-107eb2ead49c · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:52.389653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:52.389653Z digest=sha256:feb5a081bdb7c32ab6cfd7ec1301e02ebcd352dec03e4003736e75bb66d01bf7

Observation 0f2e6602-08ca-43d2-9b71-c8070ef34c97 · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.687570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.687570Z digest=sha256:dd7db88f884a851e943d8c2c00d795ea3633928a9bcc6e596f1d80bf09c69020

Observation 309ec9ad-75ad-489f-8785-ee5c42109a05 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:45.154452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:45.154452Z digest=sha256:35e388bb8aab12430db3dfedb977864ccfeae78f976ea867022f91a0f59e767a

Observation 4b3e7a82-fdde-47cc-9155-3044d9651a7f · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:46.079982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:46.079982Z digest=sha256:6a95cb1b8d860d4a9427152e32d1ae13e0e523d1cd50af9dfd198e1312b72062

Observation 1a126705-4d3b-41e5-91f1-f56f557d7299 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:38.102301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:c568499904709c2cf642bcfc47b29de28257fd00d74f529ab20aee706e577730

Observation 247a5f25-22ba-4dda-8c19-878517708f42 · inbound

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense cites this paper.

TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.926434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T18:26:19.922383Z digest=sha256:501fe29c7374ffa1ea8c421a9b6368cf8fb6bd396b8494d7b473a2110d63b05d

Observation a0f878cb-8114-447b-940e-cd4f89f64fbb · inbound

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion cites this paper.

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:46:04.298032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T15:19:46.920899Z digest=sha256:ab4b8416f08249d6de24c9da62a5efb686e7c95a814a7506831c70f9757fd264

Observation 7536afa1-8f93-42f7-aebb-3233d15663a0 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.319458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:770071576cda3e7cd3babe98b715ea02a56641ceb387562c7a7a3c2ec9b37523

Observation b8f758b8-04bd-4c98-b15a-60c8b6524bd3 · inbound

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks cites this paper.

MASCing: Configurable Mixture-of-Experts Behavior via Activation Steering Masks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.234082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T05:31:47.682478Z digest=sha256:989a588fd137e430ad4c2a892c90859b16bd76db466a67979a6aabe6ae4d7669

Observation 564491a8-4c05-4997-9103-3bee6e9e20a3 · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:01.117969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:050cc99d142021ea85dc7c942eed295652506d25fb105a7cf07dcd1543d3f073

Observation dfc78e53-93fb-41db-bf6c-48a286a0781e · inbound

Evolving and Detecting Multi-Turn Deception using Geometric Signatures cites this paper.

Evolving and Detecting Multi-Turn Deception using Geometric Signatures LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:23:32.700901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T15:17:58.804973Z digest=sha256:54b87237ab71f304cb2143def8240dbc9f0b89d5cf84b51f83fa8560f6e9429d

Observation 9946a1cd-f8e2-4ad4-82e5-f7b21607fdff · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:0e0e4ee76389d9f6c88587bcaf4efae344a9ca35436ecbea5de41d639692adb2

Observation 284104f1-8e15-4221-b954-329445444012 · inbound

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models cites this paper.

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:56:20.699869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T14:51:17.667213Z digest=sha256:9e07e95bcb59827e2e24f117746100d9038d1e21ebe126c71794e3672ec7c122

Observation 3049dbee-2d34-46df-a22d-a70daf20d4b8 · inbound

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents cites this paper.

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:16:36.009627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T09:15:57.044886Z digest=sha256:7c9e0c6448f0a46a61ca4ca78669e03e6eba7b672c17a9891bce681ab7731774

Observation a3a43767-df8f-4547-8278-1b7e140699c0 · inbound

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models cites this paper.

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:47:18.134705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T21:51:34.829877Z digest=sha256:a838d5c2e15578cc715c605c2a6f61f2153a2b36c916e054dc2a3aed99c9ed5b

Observation 61326c0e-fcbf-4c6e-9818-5e9ed641c4ea · inbound

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails cites this paper.

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:58:43.534360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T04:42:05.886984Z digest=sha256:a6bb6ed9b192db3f55a220418c388c89f9430d2e71a4b5b40d3723d1c8f3bae8

Observation 69b52be5-acb9-44ee-891c-ecefc9fbcba6 · inbound

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms cites this paper.

NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:09:34.476904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T17:13:26.201462Z digest=sha256:70b64b903b7f5f898629136fb211df0beacb18d2eeeaaf6107e9aa5a893ceeae

Observation 621dde8d-66f7-41e2-8abd-fef7d8c646a7 · inbound

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models cites this paper.

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.528094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-25T23:52:02.327522Z digest=sha256:fd65978ddad36a32404120a524acb5cfae76ae877277e795801b2619d321708a

Observation 61d50880-a3b8-44fd-8e0d-baf2b1480e20 · inbound

Do Thinking Tokens Help with Safety? cites this paper.

Do Thinking Tokens Help with Safety? LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:30:00.602650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-25T23:37:49.412578Z digest=sha256:fede6a1e8fb66a987f45f34cf4a018e36159810977c4d5f0aa6dca3334052b08

Observation c78087e5-adb2-4f36-8087-e8dc34bc18db · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:54.925819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:e95188fccb78f969578b16980b26003395430a2434742accca28285bd63d1958

Observation 368139f9-4c32-473b-8a35-3f186dd78d3a · inbound

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation cites this paper.

AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T06:40:17.865408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:40:17.865408Z digest=sha256:2e2915db67e766cac9edb893f44767449f9ce4a2cfc6508c4104db957ba6ea40

Observation 52be3610-9d22-4093-b472-d455bb546b5d · inbound

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security cites this paper.

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 1939

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:08.626130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:19:08.626130Z digest=sha256:cf3d2e77c7d2fdfe64e95f4aa9d8f92fb4e74b72cbccceecd29960d2a7acf652

Observation 2b05152e-9391-4fb2-8bf2-2754d6943dd3 · inbound

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems cites this paper.

The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:54:34.458741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:54:34.458741Z digest=sha256:b61980a760e023bd9e3349c1c851bee9fd29e41405bf5e79f1d328eaa1879dd7

Observation 517c58ae-e33a-4b57-a802-dda793e6ad6d · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:24.309791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:24.309791Z digest=sha256:4441a3d3c70c9d2f93245644517af2e90330d93b309cd074fa1b871a65683c99

Observation 0067c403-d465-4008-9400-5843a1aea210 · inbound

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion cites this paper.

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:50:47.128838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:50:47.128838Z digest=sha256:8194edea7d818473aa45f365f4c099da859a4fe1295b23d2a3fea1b7b82797fc

Observation 51f6832f-3f61-4b2a-8830-0224e9affcb8 · inbound

SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks cites this paper.

SoK: Intent-Oriented Systematization of Multi-Turn LLM Jailbreaks LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T00:32:00.935942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:32:00.935942Z digest=sha256:5b93b464f12b09c3cce127a38b48b9170995c0873394b1ed78b7b02e52001f8e

Observation fb46f4d2-7cbb-4a41-ab8f-5d17ec9f892c · inbound

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation cites this paper.

Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T05:44:22.263222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:44:22.263222Z digest=sha256:080f974bf4f3c474ac7710facd19d6c081881408b63974a552f51169f010d15b