Pith. sign in

Paper Citation Record · LEDGER

Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.21038.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21038 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:22:49.651555Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T11:57:03.257015Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1c6d1c47-9a87-431d-b297-a701ad27e765 · inbound

Security Concerns for Large Language Models: A Survey cites this paper.

Security Concerns for Large Language Models: A Survey Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:03.546803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:03.546803Z digest=sha256:29ff1c96f2da443a1347d64004f0cb1c8124730ddfe1d57fb2565da37c2a8741

Observation 34268b1e-ffbf-4159-9ee4-6aeecca7a4bf · inbound

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures cites this paper.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.701570Z digest=sha256:10a1431abed4f53e35098ba35c03823b77635a5ce1f217b915dcf1f03668fa0e

Observation fb6e697c-c26c-404d-9cd8-67e5f1dcd6a5 · inbound

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks cites this paper.

Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:23:53.634231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T19:23:47.574214Z digest=sha256:94b22c408bf8472484301bf691890344b56d77df4a7731636bae94581506e686

Observation 3a23b19a-d3e3-43e2-8369-394bba3e2be8 · inbound

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets cites this paper.

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T11:57:03.258732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-10T11:54:32.051780Z digest=sha256:9b19b7dfb6c21e6243974570a4ae8a430390063dadbd28fd233bbcdb2a1b4c11

Observation 13b27e27-1bb9-412d-b83e-a6d2f28cbdec · inbound

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment cites this paper.

MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-14T07:16:57.009797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:16:57.009797Z digest=sha256:fc29967e22041557a02cd2c60c5dff710ff196d36f88760c60936eee896fa974

Observation 074e01f3-5f73-4ca1-bd9b-859ab30fa036 · inbound

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak cites this paper.

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T06:30:55.577560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:30:55.577560Z digest=sha256:89f5724219590ccb07d42566ec28c7a393a0b22abde47e5aefc21ab8e21c5cb2

Observation 3406f3d7-d8c9-4e94-8c11-f3745de7deff · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T18:54:57.117807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:54:57.117807Z digest=sha256:a44edeba50b9bef6be14770e7f144cbacad133e7f7e3c9926f8e6feaaa4aeb13

Observation 3af5cd8e-7a50-4e32-8e75-375ad1ea6fba · inbound

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment cites this paper.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.651555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.651555Z digest=sha256:509995fe2d2415bd130ac72835779529d62a6fe0d14ab3e29ecf92c91529bfbc