Pith. sign in

Paper Citation Record · LEDGER

Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2305.14965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14965 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:46:10.276577Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T14:47:36.164059Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 040feda9-b0a1-41c3-afa3-72e5c77baceb · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 228

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T11:17:08.651871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:c364f247795686911903889ca4d385cf96da85d6ff69b52772ed12d914ac91f1

Observation cf689c6b-47fb-41e5-9807-bc3a373ebe4b · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T02:20:44.801737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:6ac0302200a5493f0bf4f48c256c10e60109b5b1482112a4077747c6ee48456e

Observation 53501d80-3eda-4740-bb0c-5dd8c415a90c · inbound

SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation cites this paper.

SoK: The Privacy Paradox of Large Language Models: Advancements, Privacy Risks, and Mitigation Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T00:46:10.276577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:46:10.276577Z digest=sha256:2563cbd74145c6a6444eea89d6db0c5fc9906d4cb86a0f4d073e29ef28ee6485

Observation 0217de0c-6e66-48c9-9d73-5d2d29e35de2 · inbound

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring cites this paper.

JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T14:51:03.766731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:51:03.766731Z digest=sha256:9542f51d081665d7008d6ea05e415e6038967c7432ad01567c39a8cf0460041e

Observation 136d3e88-d9a6-4115-8fb7-a6fef4e511d6 · inbound

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs cites this paper.

Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T05:59:27.810979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:59:27.810979Z digest=sha256:8bc2c1612297667f6ce5c276805022b9887261b42bf6420797ec17063a8a4f63

Observation 1eb5deaa-b9c3-4003-8034-75af5f0e2349 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:52.131305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:52.131305Z digest=sha256:c595774f463aeeeb0b3a89375b29934fc9351e5d0e2fdf12f337c44ef6f80840

Observation 76ec4a41-819f-4c03-b9b3-8a24fa14a0d2 · inbound

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels cites this paper.

Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T07:06:38.076927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:06:38.076927Z digest=sha256:598c63dbf22a95d47f886e1bdb7bead0ea917a0e5312b452c16b0dd6149c37d3

Observation 1341d01f-4244-4fb1-9bc8-ebb36896b172 · inbound

AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models cites this paper.

AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T12:52:49.101462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:52:49.101462Z digest=sha256:993f1d2e321534d9589393fcc4f2c047276947cd0587417470e8cf723a26391d

Observation 2bc40bfc-1f66-47cd-8be2-7cc9e127de8f · inbound

Runtime-Structured Task Decomposition for Agentic Coding Systems cites this paper.

Runtime-Structured Task Decomposition for Agentic Coding Systems Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T14:47:36.166070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T14:46:37.166185Z digest=sha256:fb2379392e8a3b92788ee5d956089aa687f9f47a9c8efa111d94e97912179625