Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.158845Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2502.00580.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.158845Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-10T02:54:35.336251Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T12:51:02.591575Z
22 of 22 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0a46e19c-6fb8-45fa-9d94-1ed5cba42769 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking.SSRN Electronic Journal, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 936fb21b-621b-463d-88cc-c76fbc986007 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Automatic Jailbreaking of the Text-to-Image Generative AI Systems
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1319f387-e11a-446b-afa5-4aa9bb9c1ef5 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreakzoo: Survey, landscapes, andhorizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9286d120-4a50-4efb-b285-aeaf00e32d60 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61eb7d40-2757-4f15-aed8-30a48687bceb · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52f07ae2-6466-4927-92f7-1b219248bd14 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1313410-f756-460b-b972-a8ea4336bf16 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreaking Large Language Models with Symbolic Mathematics
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208c5995-eb14-4156-9167-17de298fe388 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c39fb8-8774-4ed4-a068-6cc17fce0e1a · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Zhang et al
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8eafc1a8-8fcd-4740-80ca-e08de5517ad0 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Li and R
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99cba573-d03e-4bad-a81b-663b460c6b30 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Defending ChatGPT against jailbreak attack via self-reminders
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e84f6c0b-5284-4b79-bec1-c8d787b250b4 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Self-Guard: Empower the LLM to Safeguard Itself
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e81fb7a-eae5-43bf-b059-9faa0bdbf9bc · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8dacb2-5f74-4406-8f01-e7cb4bdccd0f · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks.International Conference on Language Resources and Evaluation, 2023
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 654d6091-2f5b-44b0-a4fa-eee35cbe27a0 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db77dbce-d16b-4bd4-8f48-30b32cc54638 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Best-of-N Jailbreaking
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa908190-0b93-4514-bcf0-da73429615ba · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Im- proving alignment and robustness with circuit breakers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 981c3c2e-53a4-4360-9142-8192753486c3 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Using gpt-eliezer against chatgpt jailbreaking, 2022
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 87eb2491-43c1-4e75-92e0-49a52eda2327 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation chatgpt-prompt-evaluator on aligned ai’s github, 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 699a1ea2-a4ef-4991-944b-a2e38d47c0cb · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea4b7eaf-2ebd-4752-8871-5f912d20ecd2 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation The use of confidence or fiducial limits illustrated in the case of the binomial.Biometrika, 26(4):404–413, 1934
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 56058db1-e979-4b2c-862e-a32b97257ba5 · outbound
Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation AI Control: Improving Safety Despite Intentional Subversion
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f37cef00-b916-4826-81d8-595ceaa0cd8a · inbound
If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.