Pith. sign in

Paper Citation Record · LEDGER

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 3 inbound Pith citation observations for arXiv:2502.05223.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05223 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.629797Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 459f049e-e32d-4f03-bf86-078caf999575 · outbound

This paper cites write newline.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.402335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.402335Z digest=sha256:e5e025697790daa50cdc873f3341d7f9ce6780e9715600862920c697259ff7ed

Observation 6e7713c9-9c8c-4d40-a39c-56ef3f94ee16 · outbound

This paper cites Detecting Language Model Attacks with Perplexity.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detecting Language Model Attacks with Perplexity

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.409603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.409603Z digest=sha256:9df3026d5a6e651964cb58f798109d93eb6109a84ca349977d5535e5ec2a3257

Observation 191688ba-f10c-44cf-82ca-eb966100d94a · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.415583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.415583Z digest=sha256:f6d475b8eab61792cf2cda39359709e897e4ac71817a5b64ef150e85540ee670

Observation 5088ebba-7f09-473d-87a0-af1d184e04fd · outbound

This paper cites Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.421184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.421184Z digest=sha256:8ac9d1761e8491f8759253f45f307926d073d91e4138d16d8783782f6b902d35

Observation 80cdee47-7175-435b-8cd9-458f16846005 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.426038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.426038Z digest=sha256:d1ab60c7e3a1bf2abb49e100603e6b659b407570a2aa1af752f1edd09d89035f

Observation 64be1b15-dd59-433c-845e-99f6590d888e · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.430779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.430779Z digest=sha256:9347bdedc65a2f472eaeea2348fe06ed52bc1180da994599b0cbed75bbe918c2

Observation 61743f0c-2d37-4243-8c85-ec7bf8150e34 · outbound

This paper cites The Llama 3 Herd of Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.435398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.435398Z digest=sha256:bf912605fe74e7a97d1f44cf3097a99a7888226f12a93cb0e80613ba24d5f2e8

Observation 278ecdb3-bffe-4857-8bc4-04b004ce9f9c · outbound

This paper cites BERTopic: Neural topic modeling with a class-based TF-IDF procedure.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BERTopic: Neural topic modeling with a class-based TF-IDF procedure

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.440887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.440887Z digest=sha256:aa34fdbfd44a6a3576fcf53214b29abaaad7415ef3151b44ba74f020f2fc6c35

Observation e68ec673-131f-4676-807a-597eee54ec68 · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.445532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.445532Z digest=sha256:daae345056b13007fcef67a07c8f980fbe45e545155c3a03d151ed5918b86bf6

Observation bea59369-e07a-43b2-8280-ba77c93b01e9 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.450515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.450515Z digest=sha256:308be5ee1177f1453255d69c554ee51e4c737c5548de4cee11ba88831442ec03

Observation 89ba909c-73d6-423b-9a4f-9a419dcf7c97 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.455592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.455592Z digest=sha256:650cc3ea55e6da04ebad85d1d758303a0702e47fdb22b4b1df1b39d85d09e295

Observation c3e539d8-921e-48b6-944e-6adec3c70b88 · outbound

This paper cites ChatGPT for good? On opportunities and challenges of large language models for education.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ChatGPT for good? On opportunities and challenges of large language models for education

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.460678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.460678Z digest=sha256:2e53a982d4563a1ff59248e9b3ae6753dfb71193614a627ae8311f59f0b9bc5e

Observation d9172f6b-7dbe-4588-9038-23b5edfb9671 · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.465363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.465363Z digest=sha256:7d8eda04f71043123dbdb7ce5071d8b704e97404837eba2df01802b125fd2809

Observation 69c29d1a-8897-4304-9e63-78f6bb46e2ba · outbound

This paper cites Multi-step Jailbreaking Privacy Attacks on ChatGPT.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.471038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.471038Z digest=sha256:d08120413eac8e2127d84d8b9d19a84657bd8ea89cba48c5fb31f21525d1c2fa

Observation da506306-78e8-410a-8b25-a5fda1296988 · outbound

This paper cites Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.476145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.476145Z digest=sha256:266f2a2fef3311cb617d7d47c4257124b4b1f2ecb3745caf93699c66704c8264

Observation 0c83e0ae-e84b-4d66-aa77-d32bf75a8b52 · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.481354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.481354Z digest=sha256:8a467abc2d8f92e1bc7d06623ad755a259db4463c97d89a692e02246fc514287

Observation ad97aa6d-bf26-4bda-b5a3-3bb411b7fb77 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.486463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.486463Z digest=sha256:1cf6202ab628c534865f5efdb3009b17e64e098c2d1ed8d92dca0d80a005b61d

Observation 8499386b-27f7-4af3-91a4-39fc6e9bf8ff · outbound

This paper cites Optimization and Optimizers for Adversarial Robustness.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Optimization and Optimizers for Adversarial Robustness

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.491594Z digest=sha256:3807e07e19e7661a73be6582533b9aa954fdb77314047bfcc1a05c7e499746cd

Observation 2e445c13-8ceb-4f88-a004-b19a9c5dedd1 · outbound

This paper cites Implications of Solution Patterns on Adversarial Robustness.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Implications of Solution Patterns on Adversarial Robustness

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.496465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.496465Z digest=sha256:4c811368b6c7d6069db6dcac305745cec7a6a8d9c8c24f4ff042ee6d37b5e54c

Observation d30824fb-2002-4268-a59c-558d8e1eb226 · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.501027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.501027Z digest=sha256:7909fefd8ebac16c32c58b10880427c2159c357aaa4bd734e5943aec6ab63ce6

Observation 0b1b1209-5caa-4588-a0ba-a542ca732beb · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.505854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.505854Z digest=sha256:86efa338c2ddda4a2af805929767518969f7107acfbc2018110973f41c511633

Observation 7a6bcbc2-39d1-4672-879c-75f27e4eba40 · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.510546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.510546Z digest=sha256:f30b36394b189eea227a1f8262e99101b515209b54dddc86335d660742f65394

Observation 8c63baa5-456a-4ad6-93e7-871219d91a5d · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.515272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.515272Z digest=sha256:746f1e14864583de006f71cd75dd8052dd709d66ee3d8fb85892c94ac459383f

Observation b43f2b73-5929-4512-a73f-49059039e97a · outbound

This paper cites PaCE: Parsimonious Concept Engineering for Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PaCE: Parsimonious Concept Engineering for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-09T04:22:01.131533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.520024Z digest=sha256:9115ae5f6ead68de0263d240489c8004dd86497728ce7c6c502919e0103b344d

Observation 0471f4b8-458c-4256-81a0-ff339a0b1d2e · outbound

This paper cites CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.524554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.524554Z digest=sha256:5a887ac50c7fdb5939367d2c6db28b09af76c60a7ebb2ba18aabcee56c74b5d9

Observation fac10459-bf1d-4c13-b110-7e175e3c6f89 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.529225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.529225Z digest=sha256:aefdef41390d6f4f5271cb6d821841f1836ca10a38f8ecfe5089647858651e7b

Observation 9651c7b5-3875-4322-ac62-594d9e7c45e9 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.533797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.533797Z digest=sha256:bfca76d229729e622bbf64ee5497bdd628c0adba3a02a081ebbb9897fbb16102

Observation 8fcf32ac-8614-446d-8a65-abb1eb2646e6 · outbound

This paper cites GPT-4 Technical Report.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPT-4 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.538369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.538369Z digest=sha256:aba172669dad237340c84186d6bfff0f53b09327568941cf142933fad6868c44

Observation 33f586e0-896b-443f-bb8a-6ccef5b4c4ea · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.543215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.543215Z digest=sha256:54782a57d7f345d1aff6ba85f4ccb6af8e82f21e44bef34c53fe595d4edf2e33

Observation 3271049b-af78-410a-bcea-5b81b63e4ee6 · outbound

This paper cites Language Models are Unsupervised Multitask Learners.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Language Models are Unsupervised Multitask Learners

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:22:01.613807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.548025Z digest=sha256:1ab3126385b4edb51a5a3dcb6faf16763f3a28e3125e038a9ef4fd709285f004

Observation f0f9d0b6-7cde-43f0-b938-de0eabe047be · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.552452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.552452Z digest=sha256:0f3caea9fe165c4ce04e1ce61bb5e6876e2d1743cbd4a265e483e385bb091d3c

Observation aa4e9fcd-d4da-4313-baff-dddabc045772 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.557214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.557214Z digest=sha256:a5a135cef5b1e1161f9aeaeb47cc92a059891ae27c6ef57eafa5ac80ac2c7ef8

Observation 8b8c48c7-1360-46d9-8230-a3935d895ddc · outbound

This paper cites Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.561961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.561961Z digest=sha256:0c287bbf91800798835fbe2c526911d66657e97b45601a48b312b417c272a7c9

Observation 8654f3bc-4971-45fd-a9a5-6f8bbe711918 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.566995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.566995Z digest=sha256:3e81117d87bc263fe0785c7471e3556c9efd68369785a4370020e9d8f39b6372

Observation 7f5b80dd-b813-4747-acfb-b42906214420 · outbound

This paper cites PAL: Proxy-Guided Black-Box Attack on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PAL: Proxy-Guided Black-Box Attack on Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.571836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.571836Z digest=sha256:bbb5b9e022a4e811c6d0c25be76c6fd400052b7368ab3097f730a32624c6702b

Observation 1314c283-1c55-4387-9791-ae36b421ab10 · outbound

This paper cites Fine-tuning large neural language models for biomedical natural language processing.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Fine-tuning large neural language models for biomedical natural language processing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.576698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.576698Z digest=sha256:e2c2c80015769ed7f57c26d2a9880e17ea23dd4316bc31495f81f5da12a3da1f

Observation 25260112-4486-4903-ae24-22b8dc441103 · outbound

This paper cites DAN is my new friend, 2022.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DAN is my new friend, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T04:22:01.598827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T04:22:00.581386Z digest=sha256:977a47152aef1dc6ae8ad99c52f6711fb491820260037c20e516581132b0185c

Observation ca68ceff-f619-4d82-9cd6-59b7506b6773 · outbound

This paper cites ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.586129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.586129Z digest=sha256:2ecc52dcf672a71de6bf8df9f9e80d852c27a4193dad8b5f213bd0148beccdb5

Observation 1ad1b49f-6085-4866-85ba-4e0be4fb408b · outbound

This paper cites Detoxifying Large Language Models via Knowledge Editing.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detoxifying Large Language Models via Knowledge Editing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.590879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.590879Z digest=sha256:d9678d2ec2fb16320bf06c2579a2c688fbee2a67446a454879bb10585fca55ea

Observation 13bd01fd-1ea4-4ffc-9c38-f7a871a497ed · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbroken: How Does LLM Safety Training Fail?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.595645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.595645Z digest=sha256:bbc489b0ea922b715824793da5ab4490c1e3db8350b3e5e9a23d355d89d5f99a

Observation 03d864db-e3d3-4784-b34c-74591dfc4cfd · outbound

This paper cites BloombergGPT: A Large Language Model for Finance.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BloombergGPT: A Large Language Model for Finance

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.600486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.600486Z digest=sha256:48ef121ee3330a8e37c3bfb4b162702621f51d6691df8f7d0d7f86805bbf1c68

Observation c6ca5c90-bcec-4041-8a42-ba1e7e58f2ab · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Low-Resource Languages Jailbreak GPT-4

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.605573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.605573Z digest=sha256:543616f4cafb049cb6ae2656c8777642a399623791a6bf185b3660cb6907f802

Observation f7c3d5e1-2813-4608-9a57-e73e14d2ca90 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.610406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.610406Z digest=sha256:ace0cd3a5752033f98fd2ce855cc306af1bba894f3d8731a9485b85a32090fdb

Observation 8d14b25a-76ed-4e01-97ce-a13af257036e · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.615220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.615220Z digest=sha256:8df087f2be881667e72469e1796b0ad4e8b088a3ce103e4cfcc2adf468818e0a

Observation bee0ddb6-61b0-492d-8ef8-40e55289de2d · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.620241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.620241Z digest=sha256:e4cb1ec420f13aff8c97eb1ce8419db319e99972744434acd1ce8095ba99c2d1

Observation 10abcab6-f732-4f54-aa9c-412cd4b2c790 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.624850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.624850Z digest=sha256:162d899ba9e8b02923b97c9b16ab2fdff046ad1c447ad891ddc72fb4c49456d4

Observation be48752f-c814-4ba7-ac7a-f9fa692eedd2 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.629797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.629797Z digest=sha256:23e455569c776ca859ce8b303ab3c19d6607150d7df2480ebeac750deb5f0816

Pith citing papers

Observation aee82b5c-2cb2-4aca-8879-ad5fd2588850 · inbound

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety cites this paper.

Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:28:48.757064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T01:27:16.967080Z digest=sha256:52aefdc6b7c7c694b2f6757c5b22348eab8f0e92c3514ad814d81bfc1ef68be4

Observation aea6f80c-12a2-4d8b-a998-5451fc2c2e92 · inbound

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations cites this paper.

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 143

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:56.857636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T20:13:10.814899Z digest=sha256:7c8ff56687eea6e83e8e7f766773ea8b562d823d077f232e9108051158b785d8

Observation 069a6103-2a68-4c7d-8e54-236b5a44bc91 · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs

Reference 200

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.845932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:5d948dafbb2a6f65d5daa43f7dc3e38d4bea85cd8ba9d46815f2d1b8d33ac553