Pith. sign in

Paper Citation Record · LEDGER

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks

As of 13 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2411.16730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16730 v4

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:17:16.019628Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1f55f04-62f1-4595-bc13-3d1dcbbada96 · outbound

This paper cites moral prompts.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks moral prompts

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.132245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:17:15.991152Z digest=sha256:4fb69f7734bc98d36a29f429e3e35b3caaf472b2d14586fea4e4d956e285d7e8

Observation 0fd67f83-9e59-4386-8fa3-c284181ef328 · outbound

This paper cites moral prompt.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks moral prompt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.119183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:17:15.996661Z digest=sha256:7e964a7cd1f22622b1035ebd1586d8bb8fca55d2bd8b4143eacd3c6cfc421d0c

Observation d0d27ce1-f0e5-464b-b14c-bb5cfb238159 · outbound

This paper cites an unresolved cited work.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:17:16.105402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:17:16.001726Z digest=sha256:fa11982bdc66c09a17d4c183de04a3a6e87e60641ee03b5f2e9140ad6add484e

Observation 4a34a928-a72b-4249-aed9-d0a96d582d73 · outbound

This paper cites Derived from technical reports published by developers, it can be seen that the above-mentioned large language models adopt different architectural mechanisms.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Derived from technical reports published by developers, it can be seen that the above-mentioned large language models adopt different architectural mechanisms

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.089465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:17:16.006937Z digest=sha256:b3496dcdea94189c0b7a8bf08141dd64149884de6e26b091fb5fd0d06e23164e

Observation 4bccb044-725e-43f5-b014-0cc0db40f57c · outbound

This paper cites Grok-2 Beta The training set data can come from x.AI (X.AI, 2024) which was formerly Twitter.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Grok-2 Beta The training set data can come from x.AI (X.AI, 2024) which was formerly Twitter

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:17:16.073505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:17:16.013226Z digest=sha256:e90773df29cdc2383d4794f93248bdd13f218559df6f16d9fa38376d7cbdcfd9

Observation f3fff17e-d194-4044-8d95-f14506fca272 · outbound

This paper cites Current state of LLM Risks and AI Guardrails.

"Moralized" Multi-Step Jailbreak Prompts: Black-Box Testing of Guardrails in Large Language Models for Verbal Attacks Current state of LLM Risks and AI Guardrails

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:17:16.019628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:17:16.019628Z digest=sha256:513ea314820e47a67d24eeb4ca03efdc535b6326badf6dc07998f744b27e2ee3

Pith citing papers

No inbound Pith citation observations are available.