Pith. sign in

Paper Citation Record · LEDGER

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

As of 20 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.09542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09542 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:29:05.726369Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved48
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf6cb120-67b1-41bf-b019-e9111a4386c9 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.507886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.507886Z digest=sha256:8ab857937393b3a142f357ac019398a65cfc637a62cef0a5d127b466b9127e6b

Observation ab2919dd-806c-424e-8956-62f3c213d7e9 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A General Language Assistant as a Laboratory for Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.512848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.512848Z digest=sha256:9a4d4d2453ee9b49218adbeafab268a9e51e59e36f18f8d16a9a1ab4a9242331

Observation 8973b0c1-811c-4423-a888-9e549fe628b3 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.518062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.518062Z digest=sha256:f5b13137a3445ff504c74019e54b93dae81ebd0996e51e8b55b5047312586dd1

Observation e8eafbb4-fbd9-46fd-be47-5de215094557 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.523115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.523115Z digest=sha256:38b66e37585be2837722c5c566533bca3b5996d7bcbf876b4ff6743943780014

Observation 3577d0b8-ad5e-4563-a364-055d8427f460 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking black box large language models in twenty queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.527631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.527631Z digest=sha256:ab85ea83920f45246d7dc4d95bbe1673062fdb570970df1e8c99f3450ada3571

Observation 4453f3f4-4ef9-402b-a0d5-644ec403240b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.531641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.531641Z digest=sha256:588edaf2d5431a496276ae8f66d9a68716537a64e3cfdfa492191a8c064b6931

Observation f37533c3-1181-45cd-9bc5-2bd0e96a6551 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.536341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.536341Z digest=sha256:7a39da0a92e2b9476b72f89d93cfa7c4b67530cd8282d816cb69f70d5cd08746

Observation 4829fe29-b551-41ab-b263-4392f680078b · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.540704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.540704Z digest=sha256:0f630ad3f7ccba55ea27d9ecfe94e8294d875e925bc6d828629f4e3a7142926d

Observation 11c5cfd8-6d5f-40e7-af5b-ef6cfff3b8ba · outbound

This paper cites The Capacity for Moral Self-Correction in Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The Capacity for Moral Self-Correction in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.544671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.544671Z digest=sha256:50c5559850e06c6b208e8adb6ce0d9f9ed209ae4ac5498b2ceb03a7dc2798fe0

Observation fae3ab0f-536a-4a82-82db-72bc43301046 · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.548662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.548662Z digest=sha256:056464813329195c720cbbf55681e1c2c1458923df0170a2c3d2060bbf66da9e

Observation 389448d2-09d8-404b-888c-ad0c5895393f · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The False Promise of Imitating Proprietary LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.552430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.552430Z digest=sha256:22107b409584f3c1473ed0f48ca66792157c18275dc4110fff724681738e4a57

Observation 8115d964-3cb9-4c0f-a0fd-fa6779c1d80d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.556472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.556472Z digest=sha256:7719ba2140e2d9897e494b7a383720d332d03051e786506254f35722b0a9d811

Observation 23ddf8f7-5158-4d8a-bc66-00780fae46cd · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Large Language Models Cannot Self-Correct Reasoning Yet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.560233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.560233Z digest=sha256:4e5abf1ba3ddd5930af2ef5cd9392affc5cb3d50592896369685e3effdded86b

Observation a95ddca6-7952-4e9a-89db-a1c1b4f2b944 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.564052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.564052Z digest=sha256:84bdb434dd701ff4db58dc4c404b593d813afdaa5bc90773d3695bf4d8c93714

Observation a9b3ed27-3258-4243-9144-7c54b0806b15 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.567489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.567489Z digest=sha256:6de674cdbbeb8e516bf6cfbc315897471818d2e7e524aff08227a8f505accc7d

Observation 91400847-e9e4-49f0-b266-c5cb5e21c91b · outbound

This paper cites OpenAI o1 System Card.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.571092Z digest=sha256:7828a72764f546154df0ddbc5fe4c58a3dc9828effdabeed0be709a40595394d

Observation 5cda88f7-9338-4117-b3ca-9c98239c7471 · outbound

This paper cites Safepath: Preventing harmful reasoning in chain-of-thought via early alignment.arXiv preprint arXiv:2505.14667, 2025.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safepath: Preventing harmful reasoning in chain-of-thought via early alignment.arXiv preprint arXiv:2505.14667, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.574479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.574479Z digest=sha256:15d063ff7e6e77b116a56c14833a88cb84f87d6e9fb99c2ea19709e894fd549d

Observation 4555ac79-d927-4960-9761-1d98ea0474a9 · outbound

This paper cites Safechain: Safety of language models with long chain-of-thought reasoning capabilities.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safechain: Safety of language models with long chain-of-thought reasoning capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.577483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.577483Z digest=sha256:e8db96f720d0de5f5de217dadbe93540b43d1c175820b252d30485e185204fb3

Observation 95743acf-a563-4a7c-bc33-c13c6bd14715 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.580752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.580752Z digest=sha256:66a16bdac1526fc594eef071b714b142aebb07ffcf21a049e6711223a6f456d3

Observation 955979d0-3bb4-4a34-9266-9cf5feca45ce · outbound

This paper cites THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.584328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.584328Z digest=sha256:9936acc6ff073f3c367fb3f1fd121307186118f0edd39b5cb948de5f16de80a4

Observation 2ff13057-31e3-4ca6-b906-011d7846bc84 · outbound

This paper cites Let’s verify step by step.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Let’s verify step by step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.588431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.588431Z digest=sha256:5318c6d3ca865f458d5f6baf51ca6feb09dcc726ddf452196f6252f038d4413a

Observation 56bea3d0-9a7f-48da-b88d-8094e11eb082 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.592292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.592292Z digest=sha256:0c741b290b37c38ab5970aa51508fa71d03fbef3a06f4d8ca0e3baac804bfe04

Observation d4117ab0-b475-4371-a632-f47124876341 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.596816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.596816Z digest=sha256:05004b87268689fa24dc49b25c24253e233c2581ad3b91390f5e828e30435d41

Observation 7027b879-9fbc-468f-afe9-8723c8f02100 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in neural information processing systems, 36:46534–46594, 2023.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Self-refine: Iterative refinement with self-feedback.Advances in neural information processing systems, 36:46534–46594, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.600844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.600844Z digest=sha256:10344fbf17987b670b14f903ae5e68cfb50f16bec964f0aabbdd9b60692137dc

Observation 258ede9a-b766-4929-9d77-441f447059f5 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.604723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.604723Z digest=sha256:bd8d6b51ff465a559d8131b2fe97881db93245ed9e42da1cfc1782521c3f9bcb

Observation b5eec18c-e10c-43ae-b116-8a077603f519 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.608669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.608669Z digest=sha256:e45ff22cff26d1f40573b88bb5408db5100107295caaefe6e0a98145afb50c2a

Observation f0c74a11-f904-4edb-81bd-d61f6d5d1930 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.612455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.612455Z digest=sha256:8e14716883bc7f7ac13f5b2af21e4c25de0e56cc5e5bdc543d9cd8403c541bc2

Observation d512dfe9-cd42-4b0b-b54f-d0216ac93074 · outbound

This paper cites Red teaming language models with language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Red teaming language models with language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.616067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.616067Z digest=sha256:2383bb952559b0bc191e1d708fd36cb02abd04f928db22d10fc50fedbf246a34

Observation 11466c5b-2fad-48d7-94d1-5240e4b3b6d5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.620082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.620082Z digest=sha256:35e3a156a842a49d4bcb261d0189267fc60d3fa94f51d5219bd67bf8d9811277

Observation 98bc502a-9eb2-4b34-8964-9fa8fb60fa39 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.623979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.623979Z digest=sha256:9f7ae0b716666eaef3217fba91d123d57a6ee94283255e51efc6d748cf69df56

Observation d32b7d0d-edee-42d5-9854-cb5845ba7514 · outbound

This paper cites do anything now.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs do anything now

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.627954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.627954Z digest=sha256:2c51ecbf32803a70562c16887e877757bd627bafaa5ad3cbb25d1596bce4f93d

Observation 89750b60-a096-4b78-8a8f-97026e5107e0 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.631560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.631560Z digest=sha256:06c215582d20de23a2dd5a9c31f4c9f36fd4111cc1d23bae9340f5b4cbe67cd6

Observation 9f7cc5a2-bb57-4719-b43c-f7815da885e3 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.635580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.635580Z digest=sha256:10cf081465c9768657523e88bef042f644c950ffa88e78177ea56da43c6cf136

Observation 622f2d09-2564-4895-8193-33b64ca35895 · outbound

This paper cites A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.639525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.639525Z digest=sha256:0fffb8f0c151c90d4e41e739eef54c9e2ec90b87d512983961484f9fb2627b4b

Observation d5a0c044-44ae-4807-9bc3-520d5269b55b · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.643492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.643492Z digest=sha256:0a78f31e0c5d12656d8f4dd12291a261f823e186eeee90f129b0e531ad9da570

Observation 5adeab22-8ba0-45ea-ad91-cfe5b8accf32 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Star-1: Safer alignment of reasoning llms with 1k data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.647519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.647519Z digest=sha256:2461d2a3a66e4616e69692ee6713c5fee15e19c5b7c5a8936decff46b1810e6a

Observation b7f52b90-8304-45fc-b58f-a16cff7ec945 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.652296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.652296Z digest=sha256:4ae1e35461a3e1a22b02ed11ecd9182045c1cdae4eb1adee1e38498b3ec03aab

Observation fc3de7ab-d19e-47cd-9ba2-aaa3f829d9fa · outbound

This paper cites Qwen3 Technical Report.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.656483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.656483Z digest=sha256:a2954149478c2bf9db2abe4bb859be16d64c3cfc09a1dff3fe5bd68ee9efebad

Observation 35fe2477-ad33-4155-8580-d005670e1473 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.661372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.661372Z digest=sha256:988c121e680c07f455290e49ab36a72a9cf556d7ada816c92e2c8c79bc8c8d50

Observation 0eb879f7-34b4-403d-aa6f-5c4a2055947a · outbound

This paper cites American invitational mathematics examination (aime) 2024, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs American invitational mathematics examination (aime) 2024, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.664759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.664759Z digest=sha256:90590463b75f368803bee3426fe3e230834757f84f5bcd162a6a7a6a5ce7203d

Observation 072bc497-905b-4f94-bc85-03700599a989 · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.668040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.668040Z digest=sha256:9b751beec5c5a5f7b5fe475a90bf3ef1836c94eeb0b7ee7e98d930621ffd34cd

Observation 7891c688-afcf-4cb3-a6b5-cc2bb17480c6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.671368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.671368Z digest=sha256:17aad72d05bd47142f8a7f40fe91cfa8252d0c8ae68d9e404be1bd66c2fe357d

Observation b63c6a9b-9028-4c02-a10a-4aa916a71e91 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.266808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.675575Z digest=sha256:2222041b8e7c85d56cd3d03114053fef7a42451173812c9bc425e00c331f09db

Observation 23b43115-e760-4bb9-acd8-c3f554392c80 · outbound

This paper cites Escape any double quotes in strings.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Escape any double quotes in strings

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.254399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.679000Z digest=sha256:44e78d40460b221c9a29870fb1365d44d285f1067656c6667b614ec94be84bfc

Observation 6d48f1e6-5549-438b-a06f-b8f566f3bcf9 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.242875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.682219Z digest=sha256:2e0869b2fa5bb904f99d943e7b781b53d0ed3db845fab025af239c35d5093fb4

Observation b183074e-209c-4edd-a5df-01b135f99a57 · outbound

This paper cites If you can autonomously refine the prompt based on feedback (e.g., minor rewording, removing trigger words) to improve the score, do so directly via`query_target`.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you can autonomously refine the prompt based on feedback (e.g., minor rewording, removing trigger words) to improve the score, do so directly via`query_target`

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.231930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.685646Z digest=sha256:0ed7e5349af94828c7da01ed8d4384d0ac8a81931c162497a6de442633b2c4af

Observation 33c8cf29-4510-432b-bc10-33df6bd63a97 · outbound

This paper cites If you have just called strategies twice, your next action MUST be a`query_target`to test the prompt.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you have just called strategies twice, your next action MUST be a`query_target`to test the prompt

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.220163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.689222Z digest=sha256:c55c6e3c334c9af9ab38d92f19e05ff9c3e60d6f25096a0e476a592eecfd26c9

Observation 5cb3ef28-c24b-43cd-81cf-2f8ecd840b56 · outbound

This paper cites Instead, refine the prompt further based on the scorer’s feedback to push it over the success threshold (>8.5).

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Instead, refine the prompt further based on the scorer’s feedback to push it over the success threshold (>8.5)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.207614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.693143Z digest=sha256:2ee9b1d1761c22bbd12fd52443c9583338b00dd37e2ddc887c2db19f1923c8a1

Observation a3d21896-1deb-4502-8584-b820de347158 · outbound

This paper cites Scorer prompt.The scorer model evaluates the harmfulness and actionability of the target model’s response and returns both a scalar score and refinement feedback for the attacker.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scorer prompt.The scorer model evaluates the harmfulness and actionability of the target model’s response and returns both a scalar score and refinement feedback for the attacker

Reference 49

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.195755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.696936Z digest=sha256:6f39a5f53ad407310597ae5122fe82e6cce4055baf65943ce8cbaca2bfad3ba2

Observation b8842433-a570-4bde-8154-81d9908cfa00 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.183160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.702180Z digest=sha256:f95bb8152931bfcbcd5cb48615e582ab42652a22d1db2acf669013e87270df91

Observation e64c0545-1d7a-4130-a918-1d53812aff59 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.171110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.706534Z digest=sha256:9e04ec6c1102097f4f33050c53f4d98f51230e37df56ea5e8fbc833bf3d317f9

Observation ca019b35-dfa4-42f8-8e12-ef9403170870 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.159760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.710395Z digest=sha256:0834f2381cc1a9f9971560a182becadeb1473777b2c22a9347f9d19fa94a90cf

Observation b2b06089-9a3f-48bd-a7a0-543ae0bd301a · outbound

This paper cites score": <float 1.0–10.0>,.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs score": <float 1.0–10.0>,

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.148618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.714281Z digest=sha256:5b719c816cd6fc8322c81209d17e9fd2db33d85d0b27007e25fc5ab707c550a5

Observation 0a960579-ef65-4a63-ac3a-6b6a8c68de57 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.135615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.718307Z digest=sha256:e34bc26f7bdbec78ec554d3f9cc3d50db945998fdb448c80cd8a5161285f07ad

Observation ac214206-d5fd-4dab-964b-b80725fa80f8 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.122950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.722216Z digest=sha256:884f1db0264c3b31eac209cb99c1332960fee86cd71403451531c4736feeb75a

Observation e0658f8d-8380-4456-823d-38966a0ffdc1 · outbound

This paper cites Response:.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Response:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.106220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T15:29:05.726369Z digest=sha256:04fc6c7a2f324ae71f2d3d768316fc4948a0777a68a76aff11a1988a8512d248

Pith citing papers

No inbound Pith citation observations are available.