Pith. sign in

Paper Citation Record · LEDGER

How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2401.06373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06373 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.103240Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:58:59.338870Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7c3f6893-7268-4a2e-a523-6cae8af40f29 · inbound

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models cites this paper.

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T08:39:28.183330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T08:39:28.047394Z digest=sha256:e2547c54d91fae639c4da77f0ae0e5eab732b04850f8507e986b2da35ffdf22c

Observation a027ddf2-6f46-496d-8edc-38e7890ff0dd · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.888797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:6fbd0dc901667701fb7ab36d5189a66473fd8f1711fd595c6d2d1b9d0948cb66

Observation 68191e5e-ce2a-46cf-8c6a-d91f943006b8 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:08:05.541412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:d62d3e9357b48e50844b690a10f4a18abb8c333ffd404cbda4a560abe3112794

Observation 360fcd46-7fe7-4367-866f-33c3e5b1f10f · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.463792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:cdf6fa7c1176dfa2a8db3000127d4c7a2b912645c3c79bc9ffd5aca873ad2f5f

Observation 208c5995-eb14-4156-9167-17de298fe388 · inbound

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation cites this paper.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.103240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.103240Z digest=sha256:da26928215574fa006143dff4aac11780aba46c8abd0c4b94268050c24157372

Observation 3b4e3b56-076f-46b6-869b-8c5334b66cb6 · inbound

Adversarial Reasoning at Jailbreaking Time cites this paper.

Adversarial Reasoning at Jailbreaking Time How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T14:49:08.857892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:49:08.857892Z digest=sha256:e02e758b826306a10ccbeea8158c0112a0e1e8c83a6be546f2654a4a70e061a6

Observation 4443f256-d3fc-40f1-b725-0b908f096e92 · inbound

Position: Adversarial ML for LLMs Is Not Making Any Progress cites this paper.

Position: Adversarial ML for LLMs Is Not Making Any Progress How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T12:47:21.810939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:47:21.810939Z digest=sha256:87e5c2a8758871fc5adf68a9505c93156c027a7d32c890a09ada765673d3b590

Observation 2484294f-a2ce-4a99-a523-2d19d67acf6b · inbound

Safety Reasoning with Guidelines cites this paper.

Safety Reasoning with Guidelines How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T23:50:36.388248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T23:50:36.388248Z digest=sha256:7521a237e5ba4cec1a4d43eec6bda5f6defbe47e0f2f2b2227d1d62cbbb44fe5

Observation 8d14b25a-76ed-4e01-97ce-a13af257036e · inbound

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs cites this paper.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.615220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.615220Z digest=sha256:8df087f2be881667e72469e1796b0ad4e8b088a3ce103e4cfcc2adf468818e0a

Observation 7dca48dd-3bf9-4c77-8c86-b88418aa5525 · inbound

Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models cites this paper.

Mind What You Ask For: Emotional and Rational Faces of Persuasion by Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T21:41:06.430697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:41:06.430697Z digest=sha256:19a4a11dda6225af357926d9b7ffaaf198a852fbbf7dc9282cfafe25e745e802

Observation 19f96e24-f242-4c4e-9020-7f5cb9eb2ffb · inbound

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models cites this paper.

AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T15:34:57.558902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T15:33:04.194346Z digest=sha256:1a9d5b609d325c5bbc302a2a35e0fb85fcb6b7edbc951f348b3bd306af99d95e

Observation 77ebcfc6-6a88-4f9b-b1e4-354ab618fd60 · inbound

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space cites this paper.

Breaking the Ceiling: Exploring the Potential of Jailbreak Attacks through Expanding Strategy Space How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:36:00.518360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:36:00.518360Z digest=sha256:37c17dce375cb275c37622306a6ae74899b6352c1c87409f7a60ff6aaf5bf69a

Observation f94732ee-cc52-443f-8f74-e6808cdece7c · inbound

Adversarial Preference Learning for Robust LLM Alignment cites this paper.

Adversarial Preference Learning for Robust LLM Alignment How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:23.317490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:23.317490Z digest=sha256:62659b790a278c60a1ed77964f4d157ce30a569e20c993b1c12f8a3194d7b701

Observation 60366104-391d-4310-a74b-18c7f08b0d10 · inbound

HauntAttack: When Attack Follows Reasoning as a Shadow cites this paper.

HauntAttack: When Attack Follows Reasoning as a Shadow How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:33.123729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:33.123729Z digest=sha256:03b702cb21458c37fdfc0ea6596349bbf57b8a646e9551a93f7d256e0aa5c336

Observation 9b10db3e-a698-4987-a2eb-a75b8bdc3bc4 · inbound

InfoFlood: Jailbreaking Large Language Models with Information Overload cites this paper.

InfoFlood: Jailbreaking Large Language Models with Information Overload How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T01:02:31.370445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:02:31.370445Z digest=sha256:a1090258b122e5ed4f60c45a661859576918a97156b98ce9e68e35b4a2efdab6

Observation e5e45584-1a69-4502-9c71-787d370e9609 · inbound

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation cites this paper.

MGC: A Compiler Framework Exploiting Compositional Blindness in Aligned LLMs for Malware Generation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:53.371585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:53.371585Z digest=sha256:adba337ae773180b0032b8b556ded553692f30493a191e7006e67281f8cf7165

Observation 92dd1aa6-0f85-47a0-a65a-0615be0521d5 · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:17.163189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:17.163189Z digest=sha256:e5a464aa8757b398bb8dddf39f42a5288499c5b8685401d0f092f2629065a42c

Observation 4fd72723-3dd6-45be-8ceb-8ac186636391 · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:53.616114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:53.616114Z digest=sha256:0acc7f51780f5506926d779de238418993919d8af9290a14c1c2565bc8758382

Observation 8f4ce4b1-eec0-4b31-a529-0b72a386439d · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.573789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.573789Z digest=sha256:2a4636c296a838630ec14d1fb591ea73b23b21cbefba76995002cf2e521af546

Observation 0d545991-df40-428d-8eab-c84417939a76 · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:02:54.829979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:db6eaeee51ffb9bb7af6066cfe38383642118ac1e356e68f5f9a1c4610475c4e

Observation a1cd6c39-faad-47d4-9419-a2595bb633ec · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.578072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.578072Z digest=sha256:0563f39b6c5f4324bda895d95a304d2f8c155c352570a47253b3580ce1890830

Observation cb317bd1-0ce4-40ab-8e20-63a4b38816a9 · inbound

Searching for Privacy Risks in LLM Agents via Simulation cites this paper.

Searching for Privacy Risks in LLM Agents via Simulation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-05-18T22:41:53.120499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T22:40:31.407606Z digest=sha256:f5bb8d1206bc4c45a1e59370630d73cec4b02165485c4f1f61967f80c8b4bc9d

Observation bee77cb5-d75b-416c-a846-777e694ab016 · inbound

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation cites this paper.

MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T10:22:09.180107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:22:09.180107Z digest=sha256:c21d031a28d6edeb0ad8643713eaf8c997a0e544403675416510be64e4f5b3a3

Observation 0642bbd3-0f80-49d4-92d0-5955c0b8acf8 · inbound

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion cites this paper.

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T08:04:30.692057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:04:30.692057Z digest=sha256:815b757f688f9a2a9683ce18fddfb2c0039174808d6e8957275a4206dec1f9ac

Observation 18e21459-18b4-4ad5-a288-fa31eba8bb12 · inbound

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs cites this paper.

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:55:37.922855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:54:22.995178Z digest=sha256:ae3f8f7422969243f7813101bd4b982fb9050fe595fb94579a24a19f3cb13d5b

Observation a95d557b-ad7e-495a-9cd7-fbb04edb84dd · inbound

CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks cites this paper.

CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:08:00.729279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T17:06:46.917177Z digest=sha256:a7cd6d37ff793f2be87e46c29c8b8e2851ec28d5b9091e146f0ad60f824b23a1

Observation fc791e70-5c16-41dc-982d-417f1a3cede3 · inbound

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs cites this paper.

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:21:00.078211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T16:41:52.440793Z digest=sha256:9880bc0950faa0cb13057f4381130aadda94d201bc5f26a2bf780f9f3a18e4ae

Observation 516d1186-2d15-4ce0-824e-c3acbeade42e · inbound

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs cites this paper.

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:13:29.791947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T09:10:55.001543Z digest=sha256:66e433886e58351582a08828cc7d992072100cb040750ab7ccf61c3e77a51851

Observation 0c7e81cc-d3f9-4f7c-9653-cf251fcd6402 · inbound

Jailbroken Frontier Models Retain Their Capabilities cites this paper.

Jailbroken Frontier Models Retain Their Capabilities How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:09.694106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:00:08.368799Z digest=sha256:68132d4cb52d003ff98b9e69b4ecc2bf72e5dbf4f80a536e175b89a2686ec32c

Observation e53c13aa-5b6c-4bc8-a939-15c5aa77f937 · inbound

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming cites this paper.

ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:55:32.403092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:16:21.173184Z digest=sha256:848bf592da02f4a69c1bf3c7a31addafe547c136ada69dcdccb52b188df78cd3

Observation 438d1711-9f9c-4492-a0d8-2e920a86953c · inbound

Learning from Mistakes: Can LLM Self-Recover after Misalignment? cites this paper.

Learning from Mistakes: Can LLM Self-Recover after Misalignment? How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T18:51:10.298187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:51:10.298187Z digest=sha256:96196c209fb5bf90ef4a00a01b50115a8e29bb26373357427f24c93c28b495db

Observation df6f768d-85fc-4e09-8fe2-c501d508f57d · inbound

MESA: Improving MoE Safety Alignment via Decentralized Expertise cites this paper.

MESA: Improving MoE Safety Alignment via Decentralized Expertise How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:35.255305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:52:00.377915Z digest=sha256:178f74cddf81007dc1fa64da31ca2443a454f73a25988f4c08d02bffd2f3c5bb

Observation b48141fd-ac62-452d-bd44-8b6e0c66569e · inbound

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning cites this paper.

CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:55.629602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:34:26.334078Z digest=sha256:3fb0ad0fdcae5b7cd4461ed968075cc22aaec82e92572286957b65b30e2c8d68

Observation 121405e4-9bdf-4148-addf-b8270dd4f52b · inbound

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models cites this paper.

A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:58:59.340573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T23:53:24.726702Z digest=sha256:ecce18b2024851bf14a85c518d430ac5d6ce1e8667d674d1b459c4e316b7119f

Observation 835104cf-0672-487a-b873-b0aa48305821 · inbound

RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation cites this paper.

RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:37:35.234400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:37:35.234400Z digest=sha256:4ccac3eaa26b0b7671acbf4623d62ccccf6d0579fb27341bddece5414b50abf9

Observation a37eead2-fe61-4568-bb8d-e8cebb9cacac · inbound

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks cites this paper.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:10.686916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:10.686916Z digest=sha256:c2c4c33316491f5b080cd8fdb5ded6a2de7be018d6c04edaf3dd1d3241dd2a68

Observation 5eb38f5c-d248-4d90-a46d-cb6fd5a2b453 · inbound

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity cites this paper.

Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:48:49.605652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:48:49.605652Z digest=sha256:49ecac00c854af6f09b4a0354065e53de7df70e2e96f7a851585951a5ad14403

Observation ddd55fd4-c3d8-43e7-af6c-f3ee3f943be8 · inbound

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination cites this paper.

A Multimodal Automatic Redteaming Evaluation based on Atomic Jailbreak Strategy Decoupling and Combination How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:50.880396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:50.880396Z digest=sha256:2cef3ff3e53a9a4afbea24c4b7d3fbb5060808433bb1c36f85e45272d560490e