Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:47.341663Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.06043.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:47.341663Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:08:04.673125Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T21:08:06.388576Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 20ab7275-4286-4557-9065-dbc99fd0f870 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca1a89b-0b75-4bdf-9f69-555eb9834939 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05791a07-f465-41f7-bb98-f0b59ef38b68 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9f5339-3746-4348-96d9-378f28bdd6ab · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f211f1dd-e80a-4b48-9029-1116493f6a2a · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 155f482f-eb0c-4b09-930b-c5236f1f9e56 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9a0a10f-545e-408e-a876-ba2746afa4df · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Recent Advances in Attack and Defense Approaches of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2213e70-c872-4395-afc6-7cd63728a498 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations DeepSeek-V3 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a459e4-4b2b-4c9f-9f7c-cc657cb21a57 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90cbfe8c-bc9c-412d-934d-4bf66a9f80dd · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a34fef4-bc96-4448-b0cc-ee6154146cec · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14825689-1475-4b1b-92ab-944630b02ac0 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Cai, James Wexler, Fernanda B
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2baaf5bd-1d05-4f98-852d-db0b98c6f7f7 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee87e432-a3ea-48b7-8126-f346b8db9b00 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4e61275-a567-406e-9d36-772ed10f644d · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ffdae9da-ff4b-4453-b3b6-53c2e2961b41 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fcaa8e-59eb-4e78-9b0e-88e058a1dce9 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92ec32f8-01f8-42ac-9f86-3b9e0d7f35e5 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75817f1-eb76-4a4d-ad33-e6183d1af717 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa26c873-6fc4-4540-aa2b-924dd1d18413 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations GPT-4 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a67b6306-0fed-4213-9137-12a769f00d0f · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a238d2-a2d6-4640-99b8-806ff726846c · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Qwen2.5 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39ceda5e-54fd-49f1-883c-711e3adea093 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26780fa-9390-45b8-bf79-0ac77c907ffc · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b578c107-79d4-4f86-b42f-35891a32855f · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1715f98-6308-4bfa-a4e3-663672797a83 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5284303c-4ce3-43b0-a30c-41b934161202 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations A StrongREJECT for Empty Jailbreaks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8f141d-59c6-49ae-ba2b-348faa9e9ffc · outbound
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 758c82ee-0cc8-4b74-a19b-d0aaa8d1712a · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce1bd3a-1c87-4860-93ae-a6b77a32d2df · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbroken: How Does LLM Safety Training Fail?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a191a1b-a55e-4e05-8ab0-ed4870e40c1e · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Dai, and Quoc V Le
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f7b9bf5-41b1-4bcc-b490-cae65e440cd3 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Ethical and social risks of harm from Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2c71f7-cde5-4944-bc27-36a99e3ed40e · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c38dfd-c68d-406b-834d-8f98bd11a601 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bca0902-16d6-4263-b795-1ee6f7ec13bb · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84632e61-765a-4cf8-b5b5-9911eab0c50f · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70640b84-e213-4601-b2c0-aa0b955aa74c · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Low-Resource Languages Jailbreak GPT-4
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87b57c8e-c12e-49b7-a269-950a04fd0d57 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations How Do Large Language Models Capture the Ever-changing World Knowledge? A Review of Recent Advances
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36385948-cc43-4a26-8b82-35896d8f38de · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66a5804c-728e-4412-a414-3aaff9b71d58 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3429be-0dbf-4ac4-8253-5162e656584a · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Representation Engineering: A Top-Down Approach to AI Transparency
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86dd1667-75f1-4985-ba75-b247edf29fbc · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations online" 'onlinestring :=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13d529b9-ac61-4a1e-8e94-43b2e793f351 · outbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations write newline
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1330a7-e8ce-46f9-b491-83cbe88844a7 · inbound
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.