Pith. sign in

Paper Citation Record · LEDGER

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

As of 22 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 3 inbound Pith citation observations for arXiv:2502.05223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05223 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.629797Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 459f049e-e32d-4f03-bf86-078caf999575 · outbound

This paper cites write newline.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.402335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.402335Z digest=sha256:8721dd107b2fd098f53c6c272aa4598279a4973ae57fa0ffe7a0cea0c5aba770

Observation 6e7713c9-9c8c-4d40-a39c-56ef3f94ee16 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.409603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.409603Z digest=sha256:9245269562cbcd64744d6035c1f7889b73671e1d8a087474ee9bac06029add8d

Observation 191688ba-f10c-44cf-82ca-eb966100d94a · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.415583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.415583Z digest=sha256:2b25ee8f4a5439cf00d12133640385697a1c87b887dfb76f5b5682f3a26256f3

Observation 5088ebba-7f09-473d-87a0-af1d184e04fd · outbound

This paper cites Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.421184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.421184Z digest=sha256:af60be4031116b14ccc60d782505e12d7e5c66d5534863538840071617cea31a

Observation 80cdee47-7175-435b-8cd9-458f16846005 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.426038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.426038Z digest=sha256:4430a99483dd7d01f4fb27faad294c05f9e592c8911f98049d09e9b62a3c0aa7

Observation 64be1b15-dd59-433c-845e-99f6590d888e · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.430779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.430779Z digest=sha256:6732ed91629b29e48b1cec5d5964cdafecd03dc03cfae7db50e701df7e1b4efe

Observation 61743f0c-2d37-4243-8c85-ec7bf8150e34 · outbound

This paper cites The Llama 3 Herd of Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.435398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.435398Z digest=sha256:b7bb4128fc35751b0343747ac33b3a95321b117795dbf9d92f9dd4e3d735f3bd

Observation 278ecdb3-bffe-4857-8bc4-04b004ce9f9c · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.440887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.440887Z digest=sha256:efa22dc7cea072c0debfe9dbe233b4acc4ff981c1efd3772222d851de4c9fe9c

Observation e68ec673-131f-4676-807a-597eee54ec68 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.445532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.445532Z digest=sha256:71539aa68ec8e1e0f834b80336b583014528e4a232283fde534796c59941c9e3

Observation bea59369-e07a-43b2-8280-ba77c93b01e9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.450515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.450515Z digest=sha256:17b7795a4717d21af8d225810ab56ecc981d583272c506a9ffad6f64c5c957bc

Observation 89ba909c-73d6-423b-9a4f-9a419dcf7c97 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.455592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.455592Z digest=sha256:cba12b339973bd3a081127a3fc3a13cd4cafc97d67d6fdca45e514fa79726d19

Observation c3e539d8-921e-48b6-944e-6adec3c70b88 · outbound

This paper cites ChatGPT for good? On opportunities and challenges of large language models for education.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ChatGPT for good? On opportunities and challenges of large language models for education

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.460678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.460678Z digest=sha256:8d8dffbc5106ee956c32beefd27a91c16ef1ccc8062ca34dc5789ed66817d609

Observation d9172f6b-7dbe-4588-9038-23b5edfb9671 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.465363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.465363Z digest=sha256:4b579fb288e34bbc32a6b45b753f52b70bac32623d8e1c9d9432f005622f9a35

Observation 69c29d1a-8897-4304-9e63-78f6bb46e2ba · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.471038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.471038Z digest=sha256:d44bb9ad6212ff9756a2f1b4984fd434902fccae2b1645a604ff162e2d83b992

Observation da506306-78e8-410a-8b25-a5fda1296988 · outbound

This paper cites Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.476145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.476145Z digest=sha256:34b7e7295e90f1b62e123ad121856c9bbcb7eed5c6783b2fcd9ea5781b21bbee

Observation 0c83e0ae-e84b-4d66-aa77-d32bf75a8b52 · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.481354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.481354Z digest=sha256:381472d9ce9e4681e59fe9b089188e33b484a412bf76afb552089dbb6bf6d683

Observation ad97aa6d-bf26-4bda-b5a3-3bb411b7fb77 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.486463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.486463Z digest=sha256:077281f2636255976a117eed8290ad58592841b67b273cf4f9e7b7c3cf8632bb

Observation 8499386b-27f7-4af3-91a4-39fc6e9bf8ff · outbound

This paper cites Optimization and Optimizers for Adversarial Robustness.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Optimization and Optimizers for Adversarial Robustness

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.491594Z digest=sha256:a6565332548f862d26567f6ee9dd8273fbff6631cfd468bf49691deed831780d

Observation 2e445c13-8ceb-4f88-a004-b19a9c5dedd1 · outbound

This paper cites Implications of Solution Patterns on Adversarial Robustness.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Implications of Solution Patterns on Adversarial Robustness

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.496465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.496465Z digest=sha256:ffe39a5f841e2006a365349be74e23c3ac159d7e4085eab77c647b4afbda5d49

Observation d30824fb-2002-4268-a59c-558d8e1eb226 · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.501027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.501027Z digest=sha256:8849243f20199f061ee98867299069ecfbc5d9dee349f8705c84d0bf04b4fb0f

Observation 0b1b1209-5caa-4588-a0ba-a542ca732beb · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.505854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.505854Z digest=sha256:65f99c647b41c449208e6683aeb105c2ee998984ebf2306a7399145ded37fce1

Observation 7a6bcbc2-39d1-4672-879c-75f27e4eba40 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.510546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.510546Z digest=sha256:a471821ee56ec56b4e23f1765ed1ee30dde34f21a1caaabf6469b327a723c2c3

Observation 8c63baa5-456a-4ad6-93e7-871219d91a5d · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.515272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.515272Z digest=sha256:cb2a04d4879359be464c524846cc1d24e2baaf97550364b74b9d41b086c6801a

Observation b43f2b73-5929-4512-a73f-49059039e97a · outbound

This paper cites PaCE: Parsimonious Concept Engineering for Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PaCE: Parsimonious Concept Engineering for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-09T04:22:01.131533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.520024Z digest=sha256:b3c1d943fd718f9124a5cda7bffe2e0c5f3d7d3434aff168d7ac952499a2981c

Observation 0471f4b8-458c-4256-81a0-ff339a0b1d2e · outbound

This paper cites CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.524554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.524554Z digest=sha256:61b1aacc1715d07960abf6f0e49331eb055b02a487ee3772b6795e0027a5cb29

Observation fac10459-bf1d-4c13-b110-7e175e3c6f89 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.529225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.529225Z digest=sha256:62b042fe7855234b46ab8a8202e5fa6ee39a6f19e489fdaa9d0846d44a7cf963

Observation 9651c7b5-3875-4322-ac62-594d9e7c45e9 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.533797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.533797Z digest=sha256:7e4b562b492d9ca3e689146a37b4d9783752a9ee37d2014c91b288c363dc6be4

Observation 8fcf32ac-8614-446d-8a65-abb1eb2646e6 · outbound

This paper cites GPT-4 Technical Report.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.538369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.538369Z digest=sha256:70b95121802f5bb6d0e4acf556043a162f5e8c85357536b39e1dd47f31951f0f

Observation 33f586e0-896b-443f-bb8a-6ccef5b4c4ea · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.543215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.543215Z digest=sha256:e84bf61596e39b1b038bdcf82169304640596b298c04e8585703b991cdeb3c1c

Observation 3271049b-af78-410a-bcea-5b81b63e4ee6 · outbound

This paper cites Language Models are Unsupervised Multitask Learners.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Language Models are Unsupervised Multitask Learners

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:22:01.613807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.548025Z digest=sha256:3b8b490f321354d09f7de2fb4e464041bc5075d3f2148f9b3e26828575f03d79

Observation f0f9d0b6-7cde-43f0-b938-de0eabe047be · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.552452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.552452Z digest=sha256:a6901f58675d5812f6817dda5b3c429b31076c000eb7e2dcd476c91a8e545fa4

Observation aa4e9fcd-d4da-4313-baff-dddabc045772 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.557214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.557214Z digest=sha256:f0d04db7c7cc8cd65fb06ad089bb8937128d87a253689e7622b661d5feb96238

Observation 8b8c48c7-1360-46d9-8230-a3935d895ddc · outbound

This paper cites Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.561961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.561961Z digest=sha256:9dfa8a7c2fb3d59a8376df3173d3c6e76626b122966737643db42e135f57ad74

Observation 8654f3bc-4971-45fd-a9a5-6f8bbe711918 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.566995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.566995Z digest=sha256:a4dc31d7dfcb517306e4e4f47b2a01e40c88af2ce4506eec3c1abbb34efe15ce

Observation 7f5b80dd-b813-4747-acfb-b42906214420 · outbound

This paper cites PAL: Proxy-Guided Black-Box Attack on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PAL: Proxy-Guided Black-Box Attack on Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.571836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.571836Z digest=sha256:d0aa0b8c987707c16d149082a8466d8af558c58e178158ba871ed146ee36c846

Observation 1314c283-1c55-4387-9791-ae36b421ab10 · outbound

This paper cites Fine-tuning large neural language models for biomedical natural language processing.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Fine-tuning large neural language models for biomedical natural language processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.576698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.576698Z digest=sha256:4f48ea6a2daee57969d515071c29ebb1120aaccaebf0cfe2ddd3cea3d3eed9c9

Observation 25260112-4486-4903-ae24-22b8dc441103 · outbound

This paper cites DAN is my new friend, 2022.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DAN is my new friend, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:22:01.598827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.581386Z digest=sha256:077fb9d0e22f1936ac7eb1274e4b25a4f2ca1ffbde7d273fd7f340b959b9b46a

Observation ca68ceff-f619-4d82-9cd6-59b7506b6773 · outbound

This paper cites ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.586129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.586129Z digest=sha256:bcd660ac2b8ef983834a00562dfba15a2f3d775712a8909000e12f82f3d9c0eb

Observation 1ad1b49f-6085-4866-85ba-4e0be4fb408b · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detoxifying Large Language Models via Knowledge Editing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.590879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.590879Z digest=sha256:eac5f8f968498b667b1fdd05a87033ff0742f248e0f8a39f8754cfe779e58694

Observation 13bd01fd-1ea4-4ffc-9c38-f7a871a497ed · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.595645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.595645Z digest=sha256:b5fe7d336fddcd79152ddc1357fb280e0cd67352a714f63d82e72b8bf49b281b

Observation 03d864db-e3d3-4784-b34c-74591dfc4cfd · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BloombergGPT: A Large Language Model for Finance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.600486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.600486Z digest=sha256:4e604904696a36fda480a897713fc716b177cdc28cf43fae15ffc0d16289d799

Observation c6ca5c90-bcec-4041-8a42-ba1e7e58f2ab · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Low-Resource Languages Jailbreak GPT-4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.605573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.605573Z digest=sha256:df2cb51a7e8ca6b301dd143b0ac8cf85eb4188a19cd1d97555f803a2aaa23bee

Observation f7c3d5e1-2813-4608-9a57-e73e14d2ca90 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.610406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.610406Z digest=sha256:173c822c9c3ad505b4d87a4d0c171a0e8b4b92801596016c4aa96e2c59cf0ece

Observation 8d14b25a-76ed-4e01-97ce-a13af257036e · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.615220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.615220Z digest=sha256:11f5821350bbba70a262389ec3a1d7ce3bae826fd2aa137d9e2ad9c342dfbda2

Observation bee0ddb6-61b0-492d-8ef8-40e55289de2d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.620241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.620241Z digest=sha256:2ae9042b0ac47602984e443bb89426d1488470a8d4293a7a7737cd11c49cca07

Observation 10abcab6-f732-4f54-aa9c-412cd4b2c790 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.624850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.624850Z digest=sha256:0c9d42f5604af622984c24c7dd8c33a72782d8743e03c40d2979b2e64a32c39a

Observation be48752f-c814-4ba7-ac7a-f9fa692eedd2 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.629797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.629797Z digest=sha256:8fbfd927acd9a170f664ff7ed2fe1352d2d3c56ca0bdae18c779bedba9abf886

Pith citing papers

Observation aee82b5c-2cb2-4aca-8879-ad5fd2588850 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:28:48.757064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:f2ff7db1e01799b9bd29de7bb00cee451a9fca282ad1ef388705f615707ed23b

Observation aea6f80c-12a2-4d8b-a998-5451fc2c2e92 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 143

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:56.857636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:96c04398ab16571f29908d6bea4ce8b0368e9328b2234b78fc0b2e78a6037548

Observation 069a6103-2a68-4c7d-8e54-236b5a44bc91 · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.845932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:c9abfdf2cecfb6072182062654f0a3b3a3a3858757d74ea2d874ffbacdecaf14