Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:29:05.726369Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.09542.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T15:29:05.726369Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cf6cb120-67b1-41bf-b019-e9111a4386c9 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab2919dd-806c-424e-8956-62f3c213d7e9 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A General Language Assistant as a Laboratory for Alignment
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8973b0c1-811c-4423-a888-9e549fe628b3 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8eafbb4-fbd9-46fd-be47-5de215094557 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Constitutional AI: Harmlessness from AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3577d0b8-ad5e-4563-a364-055d8427f460 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking black box large language models in twenty queries
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4453f3f4-4ef9-402b-a0d5-644ec403240b · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f37533c3-1181-45cd-9bc5-2bd0e96a6551 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4829fe29-b551-41ab-b263-4392f680078b · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c5cfd8-6d5f-40e7-af5b-ef6cfff3b8ba · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The Capacity for Moral Self-Correction in Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae3ab0f-536a-4a82-82db-72bc43301046 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389448d2-09d8-404b-888c-ad0c5895393f · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The False Promise of Imitating Proprietary LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8115d964-3cb9-4c0f-a0fd-fa6779c1d80d · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23ddf8f7-5158-4d8a-bc66-00780fae46cd · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Large Language Models Cannot Self-Correct Reasoning Yet
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a95ddca6-7952-4e9a-89db-a1c1b4f2b944 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9b3ed27-3258-4243-9144-7c54b0806b15 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91400847-e9e4-49f0-b266-c5cb5e21c91b · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs OpenAI o1 System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cda88f7-9338-4117-b3ca-9c98239c7471 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safepath: Preventing harmful reasoning in chain-of-thought via early alignment.arXiv preprint arXiv:2505.14667, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4555ac79-d927-4960-9761-1d98ea0474a9 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safechain: Safety of language models with long chain-of-thought reasoning capabilities
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95743acf-a563-4a7c-bc33-c13c6bd14715 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 955979d0-3bb4-4a34-9266-9cf5feca45ce · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs THINKSAFE: Self-Generated Safety Alignment for Reasoning Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff13057-31e3-4ca6-b906-011d7846bc84 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Let’s verify step by step
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56bea3d0-9a7f-48da-b88d-8094e11eb082 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4117ab0-b475-4371-a632-f47124876341 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7027b879-9fbc-468f-afe9-8723c8f02100 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Self-refine: Iterative refinement with self-feedback.Advances in neural information processing systems, 36:46534–46594, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258ede9a-b766-4929-9d77-441f447059f5 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5eec18c-e10c-43ae-b116-8a077603f519 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c74a11-f904-4edb-81bd-d61f6d5d1930 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d512dfe9-cd42-4b0b-b54f-d0216ac93074 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Red teaming language models with language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11466c5b-2fad-48d7-94d1-5240e4b3b6d5 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bc502a-9eb2-4b34-8964-9fa8fb60fa39 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32b7d0d-edee-42d5-9854-cb5845ba7514 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs do anything now
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89750b60-a096-4b78-8a8f-97026e5107e0 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7cc5a2-bb57-4719-b43c-f7815da885e3 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 622f2d09-2564-4895-8193-33b64ca35895 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a0c044-44ae-4807-9bc3-520d5269b55b · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5adeab22-8ba0-45ea-ad91-cfe5b8accf32 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Star-1: Safer alignment of reasoning llms with 1k data
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f52b90-8304-45fc-b58f-a16cff7ec945 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Chain-of-thought prompting elicits reasoning in large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc3de7ab-d19e-47cd-9ba2-aaa3f829d9fa · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Qwen3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35fe2477-ad33-4155-8580-d005670e1473 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb879f7-34b4-403d-aa6f-5c4a2055947a · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs American invitational mathematics examination (aime) 2024, 2024
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072bc497-905b-4f94-bc85-03700599a989 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llamafactory: Unified efficient fine-tuning of 100+ language models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7891c688-afcf-4cb3-a6b5-cc2bb17480c6 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b63c6a9b-9028-4c02-a10a-4aa916a71e91 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 23b43115-e760-4bb9-acd8-c3f554392c80 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Escape any double quotes in strings
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d48f1e6-5549-438b-a06f-b8f566f3bcf9 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b183074e-209c-4edd-a5df-01b135f99a57 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you can autonomously refine the prompt based on feedback (e.g., minor rewording, removing trigger words) to improve the score, do so directly via`query_target`
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 33c8cf29-4510-432b-bc10-33df6bd63a97 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you have just called strategies twice, your next action MUST be a`query_target`to test the prompt
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5cb3ef28-c24b-43cd-81cf-2f8ecd840b56 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Instead, refine the prompt further based on the scorer’s feedback to push it over the success threshold (>8.5)
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3d21896-1deb-4502-8584-b820de347158 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scorer prompt.The scorer model evaluates the harmfulness and actionability of the target model’s response and returns both a scalar score and refinement feedback for the attacker
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b8842433-a570-4bde-8154-81d9908cfa00 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e64c0545-1d7a-4130-a918-1d53812aff59 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ca019b35-dfa4-42f8-8e12-ef9403170870 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b2b06089-9a3f-48bd-a7a0-543ae0bd301a · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs score": <float 1.0–10.0>,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0a960579-ef65-4a63-ac3a-6b6a8c68de57 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ac214206-d5fd-4dab-964b-b80725fa80f8 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e0658f8d-8380-4456-823d-38966a0ffdc1 · outbound
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Response:
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.