Pith. sign in

Paper Citation Record · LEDGER

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks

As of 13 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2411.16730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16730 v4

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:17:16.019628Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1f55f04-62f1-4595-bc13-3d1dcbbada96 · outbound

This paper cites moral prompts.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks moral prompts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.132245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:17:15.991152Z digest=sha256:1a51259a9e1d9d95fd6307c9ca159bcca090f32a0132368d6bc3c74b1e046580

Observation 0fd67f83-9e59-4386-8fa3-c284181ef328 · outbound

This paper cites moral prompt.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks moral prompt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.119183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:17:15.996661Z digest=sha256:b0d92af6f250e47bc2d8d57c8ba2d8bdb90ac8fcc19d75885e15cf483d7c46e0

Observation d0d27ce1-f0e5-464b-b14c-bb5cfb238159 · outbound

This paper cites an unresolved cited work.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:17:16.105402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:17:16.001726Z digest=sha256:876d96c502238f2e95ce135bdd5a944fb77009a2838ae07f12095457fc50e41f

Observation 4a34a928-a72b-4249-aed9-d0a96d582d73 · outbound

This paper cites Derived from technical reports published by developers, it can be seen that the above-mentioned large language models adopt different architectural mechanisms.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Derived from technical reports published by developers, it can be seen that the above-mentioned large language models adopt different architectural mechanisms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.089465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:17:16.006937Z digest=sha256:3cb5f141f0b1f65ca469eae1f14f53891720a693dbda01c8bab119c3cc691c5a

Observation 4bccb044-725e-43f5-b014-0cc0db40f57c · outbound

This paper cites Grok-2 Beta The training set data can come from x.AI (X.AI, 2024) which was formerly Twitter.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Grok-2 Beta The training set data can come from x.AI (X.AI, 2024) which was formerly Twitter

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.073505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:17:16.013226Z digest=sha256:d5a51fcec84d4ec619cef4b6bef5923c8091e2286168d3bb9c6b0358d520dd2a

Observation f3fff17e-d194-4044-8d95-f14506fca272 · outbound

This paper cites Current state of LLM Risks and AI Guardrails.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Current state of LLM Risks and AI Guardrails

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:17:16.019628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:17:16.019628Z digest=sha256:513ea314820e47a67d24eeb4ca03efdc535b6326badf6dc07998f744b27e2ee3

Pith citing papers

No inbound Pith citation observations are available.