Pith. sign in

Paper Citation Record · LEDGER

DROJ: A Prompt-Driven Attack against Large Language Models

As of 12 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 5 inbound Pith citation observations for arXiv:2411.09125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.09125 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:05:18.310292Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:57.184523Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 84f5bb80-e193-458a-b6ad-37248f04cbf0 · outbound

This paper cites GPT-4 Technical Report.

DROJ: A Prompt-Driven Attack against Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.248906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.248906Z digest=sha256:c1b79de855ed6399e8f7acaeb6ccb4c178a7a48a868640fede37a3ccbd8626c9

Observation 499b9844-f69f-4404-bd0a-06424cd433fd · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.257404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.257404Z digest=sha256:6b1611eee2696cea7a466c7f5fb1297ccb14969d5013dc2a046f827166643574

Observation 2b5f6412-8e9b-4feb-88ad-3adbe95040a8 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

DROJ: A Prompt-Driven Attack against Large Language Models Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.272247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.272247Z digest=sha256:49a2299086701fade49396c31be02dbf95e406e210422d7f5d560a8de20cd9cc

Observation 97883a66-11dd-4ed0-9fe8-aef432d90b6c · outbound

This paper cites Baseline Defenses for Adversarial Attacks Against Aligned Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.275605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.275605Z digest=sha256:90259f17e3781d347e6136f4fa518d8ba870f5fb7ecf971e5f57718005f1a72a

Observation 55177f3b-2436-43cc-b623-78593b39ce49 · outbound

This paper cites Mistral 7B.

DROJ: A Prompt-Driven Attack against Large Language Models Mistral 7B

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.279196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.279196Z digest=sha256:54c328cba0d2b94dc41fa3b7cfec82682dac2cbcdf5e46c980b9b2bdc7364640

Observation 3748be3e-42c4-4602-8de9-3c06ec7fa0e9 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.282680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.282680Z digest=sha256:e51ec57b009240e7c7cee744cd6946b700f73adc9ff0c8df767409f492fc6563

Observation 6d0b4eaa-7234-4610-8c9b-989234954785 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.286093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.286093Z digest=sha256:9abbe1e90147155e803cf4dd3ae8a6848c0482e05915b227be434ed8be5620a8

Observation 76410ec2-e056-4511-b470-3b16c89a2efb · outbound

This paper cites ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings.

DROJ: A Prompt-Driven Attack against Large Language Models ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.293243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.293243Z digest=sha256:3781ec10741d0a7707bd702b43693fef7dbdb8dcc67aa1d9a901e1b1ef1ccc92

Observation 12eb1332-d419-40c4-bf66-1a17822cbae8 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

DROJ: A Prompt-Driven Attack against Large Language Models Finetuned Language Models Are Zero-Shot Learners

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.296806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.296806Z digest=sha256:6f1e7c94f958fcdb87168c7e9f522882d201fa950759b91458f96c3f7e2fd20d

Observation 6381b773-e8c9-48c4-92ef-f99986a5ad55 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

DROJ: A Prompt-Driven Attack against Large Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.300047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.300047Z digest=sha256:409c1597056b9dc8197ee2cc1429412d7b881dc8b79096603495bd8fdd65d32b

Observation 17bb32fe-3f83-446e-9564-6c6cd57c1830 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

DROJ: A Prompt-Driven Attack against Large Language Models GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.303373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.303373Z digest=sha256:d2117c1403b1d63175f5c8b5d937b13ba9460e36d8f28759320a85cdc5916b13

Observation ba0aabf4-8c01-4579-94de-e61e0ca19bfa · outbound

This paper cites On Prompt-Driven Safeguarding for Large Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models On Prompt-Driven Safeguarding for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.306762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.306762Z digest=sha256:5ad731d2041de07961ee5b23f5bbc01a04f750bad165fb57901accde97647bd9

Observation 77a4ff95-2fe2-4112-9023-64d5aff29bb1 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

DROJ: A Prompt-Driven Attack against Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.310292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.310292Z digest=sha256:ab0a9319eb2afc9ea169b420fd987cbe43af6aaa6648f81e031c2eec3536cc0b

Observation 1b99365c-887d-447b-9138-53138c6114d3 · outbound

This paper cites Gradient-based Adversarial Attacks against Text Transformers.

DROJ: A Prompt-Driven Attack against Large Language Models Gradient-based Adversarial Attacks against Text Transformers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.264943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.264943Z digest=sha256:e6be851f264f1bb91f210b5980f35e8a26b2299251e1037e07c9a1fb3c2fe60f

Observation 243d5db0-569a-4077-b852-58297066fc7d · outbound

This paper cites LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked.

DROJ: A Prompt-Driven Attack against Large Language Models LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.268644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.268644Z digest=sha256:2fbde4d5bab35a4277fc07cdc808b3fb76353ee7b22b6fc4c9707146805cd271

Observation 7c429825-233c-474c-bd26-21d3cf617a2e · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

DROJ: A Prompt-Driven Attack against Large Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.289592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.289592Z digest=sha256:f7c98418708a2764d6e88f060e3abaef4abf5c5b103bcf44a569d37d8303f9e0

Observation cfaa11b8-3421-401a-adc5-b3e4c3b2b82f · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

DROJ: A Prompt-Driven Attack against Large Language Models Detecting Language Model Attacks with Perplexity

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.253164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.253164Z digest=sha256:016c3244a3c2b7bd8795e1f3eb72f8e933e9d767f4bebfc08daed0173cd83872

Observation fca8b031-e28e-4df9-96ad-b4dd1afe537c · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

DROJ: A Prompt-Driven Attack against Large Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T21:05:18.260954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:05:18.260954Z digest=sha256:a40617cab5198ee154e4b133842ab3e8d6c85d5d3e69b0a4e464f11c8992d939

Pith citing papers

Observation 4340b1d7-49c0-436b-84dc-94c912fa73fd · inbound

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models cites this paper.

SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models DROJ: A Prompt-Driven Attack against Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:57.184523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:57.184523Z digest=sha256:f82d0ce2efae1acdb904ee0f7a1c627a4295e46591bedf283fa9ca5ad2fbb96c

Observation a5be9d95-d140-4449-9c6f-9a341cc5624a · inbound

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? cites this paper.

On Surjectivity of Neural Networks: Can you elicit any behavior from your model? DROJ: A Prompt-Driven Attack against Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:00:48.525273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:00:48.525273Z digest=sha256:923eaac92c3410243be063a21ae914e9f2ccb0a52f4dd7d763e20cf0b0a9adc2

Observation a308273d-44f2-4fa2-939c-cbd9a7e61673 · inbound

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems cites this paper.

LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems DROJ: A Prompt-Driven Attack against Large Language Models

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-04T17:46:19.316937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:46:19.316937Z digest=sha256:b803d64c0a9684eb30f7405b52eb8012bd266e845dbfe53d79fda3067a1d8822

Observation dab38a09-c02a-471d-ac78-c674820b964f · inbound

Probing the Difficulty Perception Mechanism of Large Language Models cites this paper.

Probing the Difficulty Perception Mechanism of Large Language Models DROJ: A Prompt-Driven Attack against Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:17:47.726898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:17:47.726898Z digest=sha256:5f88bbed8502c45c206da376ef2f4859fe11a7281b9563c99211701d9ca2d58e

Observation 5c385afd-db4c-42fe-b7bf-baee5a6e493e · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language DROJ: A Prompt-Driven Attack against Large Language Models

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.683848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:c90c436a6972202218c22100b9162ec96f6fd18a27b417046e6087448a11817f