Pith. sign in

Paper Citation Record · LEDGER

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

As of 10 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2502.00580.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00580 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.158845Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T02:54:35.336251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T12:51:02.591575Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0a46e19c-6fb8-45fa-9d94-1ed5cba42769 · outbound

This paper cites Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking.SSRN Electronic Journal, 2023.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Pushing Boundaries or Crossing Lines? The Complex Ethics of ChatGPT Jailbreaking.SSRN Electronic Journal, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.503188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.075522Z digest=sha256:c7c57e410ed0d415c51c8ecdc9f5993192064ff28ddfcea62816311d94fcd709

Observation 936fb21b-621b-463d-88cc-c76fbc986007 · outbound

This paper cites Automatic Jailbreaking of the Text-to-Image Generative AI Systems.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Automatic Jailbreaking of the Text-to-Image Generative AI Systems

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.079617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.079617Z digest=sha256:ab7639db376b27228218e7bd12cdd494e4b1b9dbbbb43a02b01e97c8163e875b

Observation 1319f387-e11a-446b-afa5-4aa9bb9c1ef5 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, andhorizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreakzoo: Survey, landscapes, andhorizons in jailbreaking large language and vision-language models.arXiv preprint arXiv:2407.01599, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.083521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.083521Z digest=sha256:843925853e42de28e7cb0d2bb993ce92e4ae4961cb4a7208604e61248d2915c6

Observation 9286d120-4a50-4efb-b285-aeaf00e32d60 · outbound

This paper cites FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.087069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.087069Z digest=sha256:3007bb8c4a5b6df72f7f165db622e4d0d0a23555dd73b7fa2ebb01bb9641a22b

Observation 61eb7d40-2757-4f15-aed8-30a48687bceb · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.091976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.091976Z digest=sha256:acd66ba8ecb23970336aaa63f506eb97a6cf37ba4d8467ac1dc0fcd5b22d5856

Observation 52f07ae2-6466-4927-92f7-1b219248bd14 · outbound

This paper cites Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.095655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.095655Z digest=sha256:1ae6a1f5a1b9c233e8234242cf2cefe766aea65cc4f1585ada86ec7a84e6c654

Observation d1313410-f756-460b-b972-a8ea4336bf16 · outbound

This paper cites Jailbreaking Large Language Models with Symbolic Mathematics.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Jailbreaking Large Language Models with Symbolic Mathematics

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.099686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.099686Z digest=sha256:a218dee081e22253d350e299887b6406d241c70ba8efd1eb67f8de7a51025e82

Observation 208c5995-eb14-4156-9167-17de298fe388 · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.103240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.103240Z digest=sha256:7f7ffd04338adfed4bf4c9feb28e47a571e58e8c2513c27ce07aa2f4b6f6b1ad

Observation 94c39fb8-8774-4ed4-a068-6cc17fce0e1a · outbound

This paper cites Zhang et al.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Zhang et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.491451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.107167Z digest=sha256:855220fb56d686357e81d73623ad33308260353b885525b9cccd1ea1cbf5f54c

Observation 8eafc1a8-8fcd-4740-80ca-e08de5517ad0 · outbound

This paper cites Li and R.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Li and R

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.480233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.110810Z digest=sha256:85d18d09478199a09b7c14e71be11cadb4294896f799c025fe94d65047bd1343

Observation 99cba573-d03e-4bad-a81b-663b460c6b30 · outbound

This paper cites Defending ChatGPT against jailbreak attack via self-reminders.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Defending ChatGPT against jailbreak attack via self-reminders

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.468253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.114240Z digest=sha256:afd8e3ae42a210b3db9ce1165d1d30c526bad5d91824d04d31922dca8ed16e71

Observation e84f6c0b-5284-4b79-bec1-c8d787b250b4 · outbound

This paper cites Self-Guard: Empower the LLM to Safeguard Itself.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Self-Guard: Empower the LLM to Safeguard Itself

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.118069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.118069Z digest=sha256:6f436e1e4c7763d357ea5bd6cadcf79efd546937c2e3b6bf18f57f5692abf839

Observation 3e81fb7a-eae5-43bf-b059-9faa0bdbf9bc · outbound

This paper cites WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.122592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.122592Z digest=sha256:a4902eb3ef18f1f3fc280dcb34aef6fc4b538a661ae41743e95a8398dea57c16

Observation fe8dacb2-5f74-4406-8f01-e7cb4bdccd0f · outbound

This paper cites Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks.International Conference on Language Resources and Evaluation, 2023.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks.International Conference on Language Resources and Evaluation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.455148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.126856Z digest=sha256:fc9a299bc54dd092047c5078dad6ab045007728d7b9e42674a07007e00721ece

Observation 654d6091-2f5b-44b0-a4fa-eee35cbe27a0 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.130552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.130552Z digest=sha256:fa877aaab25cb4662b3e1727eb6d19745b2b3c1586680a8f3c8b4983ef0db64a

Observation db77dbce-d16b-4bd4-8f48-30b32cc54638 · outbound

This paper cites Best-of-N Jailbreaking.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Best-of-N Jailbreaking

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.135569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.135569Z digest=sha256:be3a053bb1cf4ac47dbd307cb42e8cadc450c5b124a086d0c12bb85121671f35

Observation aa908190-0b93-4514-bcf0-da73429615ba · outbound

This paper cites Im- proving alignment and robustness with circuit breakers.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Im- proving alignment and robustness with circuit breakers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.443254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.139808Z digest=sha256:8e7c1e0b894e9802a86f934ddab27c358c782becd52c5fc0552e3cbca1afefab

Observation 981c3c2e-53a4-4360-9142-8192753486c3 · outbound

This paper cites Using gpt-eliezer against chatgpt jailbreaking, 2022.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation Using gpt-eliezer against chatgpt jailbreaking, 2022

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.431110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.143480Z digest=sha256:da061137fc398fab18f1141415c482fd92c5760df20ecd596dc50d9ec493924b

Observation 87eb2491-43c1-4e75-92e0-49a52eda2327 · outbound

This paper cites chatgpt-prompt-evaluator on aligned ai’s github, 2022.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation chatgpt-prompt-evaluator on aligned ai’s github, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.419367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.147228Z digest=sha256:a40b53f14690bb683b353957ac09c6a0e5c27517aaff1387c99e4b8207d7fed7

Observation 699a1ea2-a4ef-4991-944b-a2e38d47c0cb · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.150815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.150815Z digest=sha256:843ea2b8bad72738bc99af2c563f39186605790c3cd3d33332ba55a1adaef6c0

Observation ea4b7eaf-2ebd-4752-8871-5f912d20ecd2 · outbound

This paper cites The use of confidence or fiducial limits illustrated in the case of the binomial.Biometrika, 26(4):404–413, 1934.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation The use of confidence or fiducial limits illustrated in the case of the binomial.Biometrika, 26(4):404–413, 1934

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:31:14.405865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:31:14.154989Z digest=sha256:c09c4d5fb14c6cbf94c387647dceb15299abe03d2f09f3af73de2840139229ad

Observation 56058db1-e979-4b2c-862e-a32b97257ba5 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation AI Control: Improving Safety Despite Intentional Subversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.158845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.158845Z digest=sha256:5193b7543d0fb4efd37c28a73f34fa65aecdaa3b0b3218cc718ffb2cf402cc1b

Pith citing papers

Observation f37cef00-b916-4826-81d8-595ceaa0cd8a · inbound

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems cites this paper.

If you're waiting for a sign... that might not be it! Mitigating Trust Boundary Confusion from Visual Injections on Vision-Language Agentic Systems Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:51:02.601349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:54:35.336251Z digest=sha256:6510badf6c5bdf72c75475603105f9ed9c115ca44fff9ea388133fb5e66d877c