Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.629797Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 3 inbound Pith citation observations for arXiv:2502.05223.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.629797Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
47 of 47 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 459f049e-e32d-4f03-bf86-078caf999575 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e7713c9-9c8c-4d40-a39c-56ef3f94ee16 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detecting Language Model Attacks with Perplexity
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191688ba-f10c-44cf-82ca-eb966100d94a · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5088ebba-7f09-473d-87a0-af1d184e04fd · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80cdee47-7175-435b-8cd9-458f16846005 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64be1b15-dd59-433c-845e-99f6590d888e · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61743f0c-2d37-4243-8c85-ec7bf8150e34 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 278ecdb3-bffe-4857-8bc4-04b004ce9f9c · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BERTopic: Neural topic modeling with a class-based TF-IDF procedure
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68ec673-131f-4676-807a-597eee54ec68 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea59369-e07a-43b2-8280-ba77c93b01e9 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs LoRA: Low-Rank Adaptation of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ba909c-73d6-423b-9a4f-9a419dcf7c97 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e539d8-921e-48b6-944e-6adec3c70b88 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ChatGPT for good? On opportunities and challenges of large language models for education
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9172f6b-7dbe-4588-9038-23b5edfb9671 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c29d1a-8897-4304-9e63-78f6bb46e2ba · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Multi-step Jailbreaking Privacy Attacks on ChatGPT
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da506306-78e8-410a-8b25-a5fda1296988 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c83e0ae-e84b-4d66-aa77-d32bf75a8b52 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad97aa6d-bf26-4bda-b5a3-3bb411b7fb77 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8499386b-27f7-4af3-91a4-39fc6e9bf8ff · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Optimization and Optimizers for Adversarial Robustness
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e445c13-8ceb-4f88-a004-b19a9c5dedd1 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Implications of Solution Patterns on Adversarial Robustness
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d30824fb-2002-4268-a59c-558d8e1eb226 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1b1209-5caa-4588-a0ba-a542ca732beb · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a6bcbc2-39d1-4672-879c-75f27e4eba40 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c63baa5-456a-4ad6-93e7-871219d91a5d · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b43f2b73-5929-4512-a73f-49059039e97a · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PaCE: Parsimonious Concept Engineering for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0471f4b8-458c-4256-81a0-ff339a0b1d2e · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac10459-bf1d-4c13-b110-7e175e3c6f89 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9651c7b5-3875-4322-ac62-594d9e7c45e9 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fcf32ac-8614-446d-8a65-abb1eb2646e6 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPT-4 Technical Report
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f586e0-896b-443f-bb8a-6ccef5b4c4ea · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3271049b-af78-410a-bcea-5b81b63e4ee6 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Language Models are Unsupervised Multitask Learners
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0f9d0b6-7cde-43f0-b938-de0eabe047be · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4e9fcd-d4da-4313-baff-dddabc045772 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Code Llama: Open Foundation Models for Code
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b8c48c7-1360-46d9-8230-a3935d895ddc · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8654f3bc-4971-45fd-a9a5-6f8bbe711918 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f5b80dd-b813-4747-acfb-b42906214420 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs PAL: Proxy-Guided Black-Box Attack on Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1314c283-1c55-4387-9791-ae36b421ab10 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Fine-tuning large neural language models for biomedical natural language processing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25260112-4486-4903-ae24-22b8dc441103 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs DAN is my new friend, 2022
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ca68ceff-f619-4d82-9cd6-59b7506b6773 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad1b49f-6085-4866-85ba-4e0be4fb408b · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Detoxifying Large Language Models via Knowledge Editing
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13bd01fd-1ea4-4ffc-9c38-f7a871a497ed · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Jailbroken: How Does LLM Safety Training Fail?
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d864db-e3d3-4784-b34c-74591dfc4cfd · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs BloombergGPT: A Large Language Model for Finance
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ca5c90-bcec-4041-8a42-ba1e7e58f2ab · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Low-Resource Languages Jailbreak GPT-4
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c3d5e1-2813-4608-9a57-e73e14d2ca90 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d14b25a-76ed-4e01-97ce-a13af257036e · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee0ddb6-61b0-492d-8ef8-40e55289de2d · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10abcab6-f732-4f54-aa9c-412cd4b2c790 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be48752f-c814-4ba7-ac7a-f9fa692eedd2 · outbound
KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee82b5c-2cb2-4aca-8879-ad5fd2588850 · inbound
Response-Based Knowledge Distillation for Multilingual Jailbreak Prevention Unwittingly Compromises Safety KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation aea6f80c-12a2-4d8b-a998-5451fc2c2e92 · inbound
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
Reference 143
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 069a6103-2a68-4c7d-8e54-236b5a44bc91 · inbound
Prompt Governance? On Governing Technologies Governed by Natural Language KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs
Reference 200
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.