Pith. sign in

Paper Citation Record · LEDGER

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2607.19829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19829 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:29.288152Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:49:20.302614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T10:46:11.686159Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0cd9fb51-16b4-4a21-9632-2539ec7881a4 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.308203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.308203Z digest=sha256:83dc8c634f0075030bcd2c8b6b8980aa7ca40a24693a864af05253c16bf31690

Observation d88f9d01-4cc3-4048-a3cc-ab972bdb7c47 · outbound

This paper cites Dynaguard: A dynamic guardian model with user-defined policies.arXiv preprint arXiv:2509.02563,.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Dynaguard: A dynamic guardian model with user-defined policies.arXiv preprint arXiv:2509.02563,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.702036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.702036Z digest=sha256:b8b3f8c250dd521ed30ed63b3e7cbe1ecd682eb1a6e03b8797598f7c86cfd5c9

Observation 10667c4e-b9a1-4623-8233-fb3168bb72d6 · outbound

This paper cites Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.111834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.111834Z digest=sha256:52482502b7013669d6bd099e5f6474a8c27a91da1fb332bb04e47743ff66d965

Observation 9490cd8a-54f1-426e-b140-c686b44f30d2 · outbound

This paper cites Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.258662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.258662Z digest=sha256:6b03ddeca403c086df3ea715d5dc714a1a22484460a46587a778d6edcfa19608

Observation 221abdd0-51f2-4066-af83-0b87d3abe50a · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.294658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.294658Z digest=sha256:143534d44cb8afa28901f3b435f287b7b3cc38c6518d68d5e02dd19c12a24c8c

Observation 7bee6b2c-cff6-44d5-870d-f6020aa95d81 · outbound

This paper cites Can a suit of armor conduct elec- tricity? a new dataset for open book question answering.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Can a suit of armor conduct elec- tricity? a new dataset for open book question answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.365443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.365443Z digest=sha256:61e8f85bfe052b410d85a202ed0d15955662fb1a04f295e666e60185c10f0d24

Observation 854e725e-9c82-4c49-80bd-72bfce434bb3 · outbound

This paper cites Granite Guardian.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Granite Guardian

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.436778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.436778Z digest=sha256:e9fb33e710b490f93b99b4001f6261461597cd47b23ec242068287319325d7de

Observation 9b521d1e-c292-4b2c-9a38-8ad87089f048 · outbound

This paper cites Majic: Markovian adaptive jailbreaking via iterative composition of diverse innovative strategies.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Majic: Markovian adaptive jailbreaking via iterative composition of diverse innovative strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.527425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.527425Z digest=sha256:36ec2d1ee91ee60ed2488d1d529c4ebad7247aa64a04f074691b7b4ff82ef409

Observation c5533a15-3c8c-4db2-8cf6-0ce30afd0641 · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.612489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.612489Z digest=sha256:ced84e38cab8ffbc2150ade73290a078a304f0abbb7267427f561ed22a429442

Observation 6045229e-334d-45e2-9e93-a652524a5b4c · outbound

This paper cites ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.834194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.834194Z digest=sha256:a0f463464d265016162dfd98c7fede822d4d4b5dea70252553676d20205b6f2a

Observation 8970f135-f645-4ec0-b2f6-e59684d0f5b1 · outbound

This paper cites DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.995842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.995842Z digest=sha256:3f491f02f800fbc238b98b9b1b3aac5df05fd7fe18ea478a96f691fa2e199ae0

Observation d58ce2cc-fe83-43d2-b06f-990d925cb16b · outbound

This paper cites Dynamic target attack.arXiv preprint arXiv:2510.02422,.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Dynamic target attack.arXiv preprint arXiv:2510.02422,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:28.214780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:28.214780Z digest=sha256:97ffd5e14f8f0a89e05ac3d0f0043c66627d32ff8f342ec6d7a81a15478bb497

Observation f957e25a-8ef6-4671-90fc-f4f47f355516 · outbound

This paper cites Harmmetric eval: Benchmarking metrics and judges for llm harmfulness assessment.arXiv preprint arXiv:2509.24384,.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Harmmetric eval: Benchmarking metrics and judges for llm harmfulness assessment.arXiv preprint arXiv:2509.24384,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:28.396101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:28.396101Z digest=sha256:a68d55711429e9343070e7132d3eb29bea6b9e93e2ae9951f2c7c1e2bfb11466

Observation fe54d41a-03dd-41e9-8a8b-b3d9702bcb8e · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:28.639197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:28.639197Z digest=sha256:a1ed9140abe25e596a08114e59022b8744c62d8567d902d12e8edd6ac58ff8a0

Observation 0dbe886c-34d6-43db-a75a-7fa7827063bc · outbound

This paper cites ShieldGemma: Generative AI Content Moderation Based on Gemma.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection ShieldGemma: Generative AI Content Moderation Based on Gemma

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:28.857968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:28.857968Z digest=sha256:c4d7b06b240197cf814aab535e2ee416b6b3e213f0dbd49fbacb85cc83972412

Observation d3c2feac-4084-4a7f-9b3a-cb8a85d742be · outbound

This paper cites Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Evolving Skill-Structured Attack Memory Enhances LLM Jailbreaking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:29.039928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:29.039928Z digest=sha256:23259cfea3e2e56e794eda4d31604b1bac5accf14e118689d7e449b329f8a602

Observation 15f55aeb-07aa-46dd-8b40-4b1ee674b554 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:29.288152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:29.288152Z digest=sha256:4069f08b4fe807605be4f81f12e62dbbe2dfd20e2a120a80b74625580314201b

Observation 6d40fd64-6589-481e-9e53-08b0119ad890 · outbound

This paper cites Qwen3Guard Technical Report.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Qwen3Guard Technical Report

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:29.208494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:29.208494Z digest=sha256:79ded632170d4e98078c85efe149b118db4e27f5b54b75472b8ba8e92721f3b1

Observation 395d7eb0-19b4-49a6-8ddb-268b82d5a129 · outbound

This paper cites YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.996387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.996387Z digest=sha256:8aa57441e15597b6b2e6e0cb19347226c44ccfaa58729fc4e4a50d2bfc83c021

Observation 5ba3c237-a16f-4fd8-8a2c-e98f075c15ce · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.463838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.463838Z digest=sha256:b760e8420f726261f37def156c5774677f14d67efc197e5525c4cece1a35f437

Observation 1f6a45b4-2ba2-4116-afa2-68f9d5e552bf · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.383606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.383606Z digest=sha256:fca04e42b33866a9d8f743eb6bb7c414b1ea2d7c1cf8690844a4644bdc6c8797

Observation 13b206d3-c391-4ac3-ba2d-c0fd199f33f4 · outbound

This paper cites Race: Large-scale reading comprehension dataset from examinations.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Race: Large-scale reading comprehension dataset from examinations

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.932655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.932655Z digest=sha256:d1c6f36e5bd432416606df0020cb64242aa3fddfd9fe9d15150b9038a5bb8377

Observation e6f17d31-b5a2-4056-926d-74a0b212b90c · outbound

This paper cites A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.561319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.561319Z digest=sha256:d4f3946ea4cc6e3ae84313cfbb159f37a490842d6f4fd82fcc2c9fb436dc173b

Observation c8def011-a8e3-4757-8c17-ccc5f25fcd72 · outbound

This paper cites Autodan: Generating stealthy jailbreak prompts on aligned large language models.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Autodan: Generating stealthy jailbreak prompts on aligned large language models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:27.181672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:27.181672Z digest=sha256:526ea0e79a4d0336f83920722f24a06e0442fb4e6f998284dd6f906f81ba3f52

Observation a3fe96a3-5ae9-4aa7-a047-2c6670dbf6b5 · outbound

This paper cites an unresolved cited work.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.624887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.624887Z digest=sha256:f9f1739df7f6894b672bc0bc3e3232b260695cd6118efc269965a46f5c30a27d

Observation 70930a7a-5005-46e0-ae70-a2fe46ce03fa · outbound

This paper cites Dualbreach: Efficient dual-jailbreaking via target-driven initialization and multi- target optimization.arXiv preprint arXiv:2504.18564,.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection Dualbreach: Efficient dual-jailbreaking via target-driven initialization and multi- target optimization.arXiv preprint arXiv:2504.18564,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.781514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.781514Z digest=sha256:2a16cb136f64240bc595d1ae5ce267b4628038145fd535fd45f571b61d7d4744

Observation e8aaf750-6a18-4fec-9160-935d9589216e · outbound

This paper cites NonTextual Target Attack.

DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection NonTextual Target Attack

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T11:40:26.849024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:40:26.849024Z digest=sha256:4ed5c6ff5f2ae9a7f7ff3584a6d63672b6e8a85e5f51da3f9cbfa4058d26bd2d

Pith citing papers

Observation afe0bdc5-8d00-441b-b918-89d842d108e4 · inbound

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning cites this paper.

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Reference 97

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T10:46:11.691488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T10:46:10.892459Z digest=sha256:b642f8df49e4428e374b8b3ffb126ddf663203a36de439ca6f67dd505c327828

Observation 15258ecd-328c-47e5-99ee-d10dba8ba14f · inbound

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions cites this paper.

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:49:20.302614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:49:20.302614Z digest=sha256:c0b0343919c87f04e91d651bb846ebce4b195d398b8d52572aa26e6954b4757a