Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:20.752585Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.07402.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:20.752585Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:03.319920Z
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7ac0cda6-c77c-48c6-99e7-985e17627304 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal language model fine-tuning for text classifica- tion, 2018
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a8c6123-ac07-45b4-aaef-ea22fed4a683 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Training language models to follow instructions with human feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7e06a19-25a4-49ea-9cf4-46033b91e68d · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Scaling instruction-finetuned language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f687c9cc-53af-4811-acf1-5026fc9bf588 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Fine-Tuning Language Models from Human Preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6938c8e7-dd8a-473b-a01a-e7b0717c189e · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Direct preference optimization: Your language model is secretly a reward model
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be8403d9-f7ee-4916-a14e-cf1608a7fa44 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreakchat.com, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6ff2025-0476-4e55-9c57-d49f5a91cb13 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1a9f34-036d-4335-8a32-9f3ab8954e01 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36:61478– 61500, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd3b9ee6-56ab-4793-ba5a-dc13d7fd6fcf · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures do anything now
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1753045-f0d9-442a-aecf-64ed9abd11ff · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don’t listen to me: understanding and exploring jailbreak prompts of large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac5bca8e-c73d-4f9e-90ca-fa8bed0495f8 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36c6d1dc-e66c-4e72-868b-e30305efdd7a · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d463ea2-dc58-4f1d-b807-d3b19f063b03 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff06be60-b1ab-4ac7-b4c7-f3d678f2af6b · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa7ba33-3fe7-41de-877d-2461d41a0ce3 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d202259b-976c-48e2-8129-af12e45802bc · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3d46d1-a7b7-4829-a1a1-a80d0775e3ca · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd526e5-5d8d-4e82-91ab-2b9b4db91aa4 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 595452c4-6357-4760-8f52-30d432d8a57b · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81904555-64f5-41c7-a0fe-bfecc3e67c03 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72aa392a-d53a-4eb3-8989-88f9158796ad · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ba783a-5b5a-4037-a109-1035b69cdfe8 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don't Say No: Jailbreaking LLM by Suppressing Refusal
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95cbf8a9-6f0a-4ea8-b67d-d55cec96b58f · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37d45666-6cd8-4238-a472-ce1bc8adba7e · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d542fe5-5475-47de-b4ff-013c8b89ec90 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcc24c3-5661-475e-b745-7e0a568132f8 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Llm jailbreak attack versus defense techniques–a comprehensive study
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec8ff256-1793-4bab-b989-42bd5d49e3b0 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures LLM-Safety Evaluations Lack Robustness
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2eaf344-2622-4f66-b217-4850c0eb5904 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f20d886e-3a37-4af6-864a-5e093bcc8e5e · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebeab4e4-5f8f-4310-a1ba-9de7ff330d25 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c51fb7-d940-44ea-83c4-8445e1beb86d · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f30fa75-a85a-42b8-b79d-0bb40852e600 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Clas 2024: The competition for llm and agent safety
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7db82a51-965f-4deb-8037-686fccefef3d · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures A StrongREJECT for Empty Jailbreaks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 606f5adb-80f0-41a5-8703-0e69f61b6ff6 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The art of saying no: Contextual noncompliance in language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 931ffb9d-d955-4d82-b723-9b4705eebfc9 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0dc3ecf-570a-4d1e-ae01-7785e1e52f36 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking large language models with symbolic mathematics, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b995f12-018a-4ea8-83a6-6fbb92a7a99d · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34268b1e-ffbf-4159-9ee4-6aeecca7a4bf · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37041393-4ed4-4b92-a263-8c859f7bd858 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Rethinking How to Evaluate Language Model Jailbreak
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a7751e3-04d6-4088-b57d-1229ae8704ae · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1070a30-5b1e-4ab6-992e-ec2f7ad6ff83 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26835b23-1641-4939-ab42-09bdb134a4a5 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2912aaeb-2854-4b65-8fa6-47bfc05b8e3b · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multimodal Situational Safety
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b139af91-eb49-41c9-8e53-aad5732b7f87 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b5f8c9-896d-4480-b287-6ebff5ed5fad · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e565ca-a31a-48a8-acc5-b88382c9daf0 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ee1f79-47cb-4735-aa43-bfa24bca4413 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multilingual Jailbreak Challenges in Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c549df12-5d4b-49ff-9913-808afde46e72 · outbound
Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 394b7429-0807-402b-917e-b5967f811225 · inbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.