Pith. sign in

Paper Citation Record · LEDGER

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.09542.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09542 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:29:05.726369Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved48
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf6cb120-67b1-41bf-b019-e9111a4386c9 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.507886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.507886Z digest=sha256:e7895a82041f94c82658a06ed6e3a866a32006a6098e15eded639e033239761c

Observation ab2919dd-806c-424e-8956-62f3c213d7e9 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A General Language Assistant as a Laboratory for Alignment

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.512848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.512848Z digest=sha256:92f784379924400889c7ba714ee3b83a2f7f7f976ac0bf855350377f18b9a9ac

Observation 8973b0c1-811c-4423-a888-9e549fe628b3 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.518062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.518062Z digest=sha256:517a8ad6c2d67027210c778cfbf1fe521c4f39536869e0752c29f6f1ff80e21b

Observation e8eafbb4-fbd9-46fd-be47-5de215094557 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Constitutional AI: Harmlessness from AI Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.523115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.523115Z digest=sha256:229326dc58eaa63a32772c795deddad67ad5e95195b2907f9f9e8f46b5337563

Observation 3577d0b8-ad5e-4563-a364-055d8427f460 · outbound

This paper cites Jailbreaking black box large language models in twenty queries.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Jailbreaking black box large language models in twenty queries

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.527631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.527631Z digest=sha256:13ba7e72326993bbd89002b68f6c544621d8f495047470b5a46ef906c1fae898

Observation 4453f3f4-4ef9-402b-a0d5-644ec403240b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.531641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.531641Z digest=sha256:fe7d2921a301cd98248de9d6dd2cc5fd19e1dc20d7887cceeae35e22fc90c124

Observation f37533c3-1181-45cd-9bc5-2bd0e96a6551 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.536341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.536341Z digest=sha256:cef94cccbd73e57e3d9db0ca04be449f4045d8c9ca2a84eb72b429f84b833fc2

Observation 4829fe29-b551-41ab-b263-4392f680078b · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.540704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.540704Z digest=sha256:311b31212a1d56ddb11b161c62f8c896f2f10e09305f2be752c3730ae792ebab

Observation 11c5cfd8-6d5f-40e7-af5b-ef6cfff3b8ba · outbound

This paper cites The Capacity for Moral Self-Correction in Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The Capacity for Moral Self-Correction in Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.544671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.544671Z digest=sha256:a89d3f16747068c29173e9c9b045ddeca5a1c42b56f7f2a6e02251a19c90dfe5

Observation fae3ab0f-536a-4a82-82db-72bc43301046 · outbound

This paper cites Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665–673, 2020

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.548662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.548662Z digest=sha256:9967b3f388edbde3688208b62dc71de3a6ffbcd056dd7339eac77223e124be0b

Observation 389448d2-09d8-404b-888c-ad0c5895393f · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs The False Promise of Imitating Proprietary LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.552430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.552430Z digest=sha256:3b0c2e9d4aaa777b386a5f77a1a862c99ea9f6099b4c44e6e1708791b2b4c08e

Observation 8115d964-3cb9-4c0f-a0fd-fa6779c1d80d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.556472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.556472Z digest=sha256:0586405bc0f15bb19bb47b4abaf4b4aa79fb2600fb742479d6cf05d245e9608a

Observation 23ddf8f7-5158-4d8a-bc66-00780fae46cd · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Large Language Models Cannot Self-Correct Reasoning Yet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.560233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.560233Z digest=sha256:5576c909045009b10feaddff5ded8caaa36392153c259e768bf3e57185a8837a

Observation a95ddca6-7952-4e9a-89db-a1c1b4f2b944 · outbound

This paper cites Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.564052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.564052Z digest=sha256:78807332b604dacf160c294c2245cc3e271dea4a6a41369563e27ef0d552aa69

Observation a9b3ed27-3258-4243-9144-7c54b0806b15 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.567489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.567489Z digest=sha256:69ccb781aff372b829552627d98742cadcff7361b0bc112399f60af5ed332f93

Observation 91400847-e9e4-49f0-b266-c5cb5e21c91b · outbound

This paper cites OpenAI o1 System Card.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.571092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.571092Z digest=sha256:5b4a465528800464ba5261184751d39d1d03371ea8e56d2023d9fdffc506cce4

Observation 5cda88f7-9338-4117-b3ca-9c98239c7471 · outbound

This paper cites Safepath: Preventing harmful reasoning in chain-of-thought via early alignment.arXiv preprint arXiv:2505.14667, 2025.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safepath: Preventing harmful reasoning in chain-of-thought via early alignment.arXiv preprint arXiv:2505.14667, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.574479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.574479Z digest=sha256:89d87b44d59e5c9c033c0b9b0ae0196bd7b6d7af4bda879e62dd1c14bb5b9a74

Observation 4555ac79-d927-4960-9761-1d98ea0474a9 · outbound

This paper cites Safechain: Safety of language models with long chain-of-thought reasoning capabilities.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Safechain: Safety of language models with long chain-of-thought reasoning capabilities

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.577483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.577483Z digest=sha256:5bc131b9ffe904c0c1e6b3f5721aba82749a4d6e1507d63cee87d6067057b7d6

Observation 95743acf-a563-4a7c-bc33-c13c6bd14715 · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.580752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.580752Z digest=sha256:5914989a9b1b4b5b54c9bf085f74ac33e29c802372f553a5a5d87729ef79f614

Observation 955979d0-3bb4-4a34-9266-9cf5feca45ce · outbound

This paper cites THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.584328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.584328Z digest=sha256:239d34dbc6b64f65c7445216e190e6ce37fa37a3bb0f6501ccc3171da75ebefa

Observation 2ff13057-31e3-4ca6-b906-011d7846bc84 · outbound

This paper cites Let’s verify step by step.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Let’s verify step by step

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.588431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.588431Z digest=sha256:3921398e1848aa264a7d3c54e15bbb9a64c90b3b862145ef6a6e53658ba4a045

Observation 56bea3d0-9a7f-48da-b88d-8094e11eb082 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.592292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.592292Z digest=sha256:f5b3f55e6c78bf75c30641fc3fea7b23c1d7b7fb2daef4d7e9bb51685e6e2998

Observation d4117ab0-b475-4371-a632-f47124876341 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.596816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.596816Z digest=sha256:11d2db777c46b7e4412b3d1cbc38d9cc794a8b8b552f89866f7dfa85ce927233

Observation 7027b879-9fbc-468f-afe9-8723c8f02100 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.Advances in neural information processing systems, 36:46534–46594, 2023.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Self-refine: Iterative refinement with self-feedback.Advances in neural information processing systems, 36:46534–46594, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.600844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.600844Z digest=sha256:89ebf66bc68c5cee1cd6faf810319870c00291fa55b95d2abe36789e0cd91591

Observation 258ede9a-b766-4929-9d77-441f447059f5 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.604723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.604723Z digest=sha256:f77389b2d68480efa0d540e6123d2b6d9f192a9f47add40ff1ea435a74d42ed9

Observation b5eec18c-e10c-43ae-b116-8a077603f519 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.608669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.608669Z digest=sha256:659650d98d50cd47ae9ac945e4f865ad583fd8e7bb2c7624581157217a5f5a93

Observation f0c74a11-f904-4edb-81bd-d61f6d5d1930 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.612455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.612455Z digest=sha256:86d28e735ef6e48b306cd0c3eeaf172fd4dbcf791ec30441970ca9425b041217

Observation d512dfe9-cd42-4b0b-b54f-d0216ac93074 · outbound

This paper cites Red teaming language models with language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Red teaming language models with language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.616067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.616067Z digest=sha256:57b08cc0dd0cd083b830d7b988b3d5ac577d171f55d9bb7792c6d521726e91a8

Observation 11466c5b-2fad-48d7-94d1-5240e4b3b6d5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.620082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.620082Z digest=sha256:ed6dc43e06bb2cc9ca667b9890406670af28c732170cea88315a76973c25df72

Observation 98bc502a-9eb2-4b34-8964-9fa8fb60fa39 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.623979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.623979Z digest=sha256:7f0eb928ca98c862f2598140fd60855d1c8ecc289df29e2f482aa8b745d47329

Observation d32b7d0d-edee-42d5-9854-cb5845ba7514 · outbound

This paper cites do anything now.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs do anything now

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.627954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.627954Z digest=sha256:57d0837f8151f1bc5306664654024772b4fba55119e3d1e05af61082e2d1f80f

Observation 89750b60-a096-4b78-8a8f-97026e5107e0 · outbound

This paper cites Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Reflexion: Language agents with verbal reinforcement learning.Advances in neural information processing systems, 36:8634–8652, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.631560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.631560Z digest=sha256:8f060cc41364a6eed64df5c04cd72ae0d1696c151b031b759c5a2cc87d922b86

Observation 9f7cc5a2-bb57-4719-b43c-f7815da885e3 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.635580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.635580Z digest=sha256:ed74bdaaaae0b9850461b568d8355a283e11d76ab0aef958dbecb929ee4d5e05

Observation 622f2d09-2564-4895-8193-33b64ca35895 · outbound

This paper cites A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs A strongreject for empty jailbreaks.Advances in Neural Information Processing Systems, 37:125416–125440, 2024

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.639525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.639525Z digest=sha256:7642c28f25e779492b66a74473a8f5cb37415b682534cb002a569ae9e508b997

Observation d5a0c044-44ae-4807-9bc3-520d5269b55b · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.Advances in Neural Information Processing Systems, 37:95266–95290, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.643492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.643492Z digest=sha256:ba63f85917acef763b2c79f74b1cc3625b041b8c44bc65f81809d9736520cc04

Observation 5adeab22-8ba0-45ea-ad91-cfe5b8accf32 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Star-1: Safer alignment of reasoning llms with 1k data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.647519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.647519Z digest=sha256:73d347dc7a1e7e32c9fa8bd45fd7c53f7939850d305673e4ea94abeed369d1f3

Observation b7f52b90-8304-45fc-b58f-a16cff7ec945 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.652296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.652296Z digest=sha256:b51cd80eb407997ce6407df0ed7f27203b1a357513a04a6b62407dce801e0969

Observation fc3de7ab-d19e-47cd-9ba2-aaa3f829d9fa · outbound

This paper cites Qwen3 Technical Report.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Qwen3 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.656483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.656483Z digest=sha256:4d5f0aa47a3cb43246d5322198f027eab2c08d59f6f9d3379166a208e840330a

Observation 35fe2477-ad33-4155-8580-d005670e1473 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.661372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.661372Z digest=sha256:6ce755738ac97a2fef46b51afc7047f5cce388a794a25bfcbe0e27053232d435

Observation 0eb879f7-34b4-403d-aa6f-5c4a2055947a · outbound

This paper cites American invitational mathematics examination (aime) 2024, 2024.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs American invitational mathematics examination (aime) 2024, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.664759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.664759Z digest=sha256:433fcae77a48bcb6302de12c533a25f6342251df604ea25bd447e62eea4ba447

Observation 072bc497-905b-4f94-bc85-03700599a989 · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.668040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.668040Z digest=sha256:bfe765ecaf47bf110a36fdb67af2ea80b1f6949f9a2929ece92b165943812493

Observation 7891c688-afcf-4cb3-a6b5-cc2bb17480c6 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T15:29:05.671368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:29:05.671368Z digest=sha256:795156b5c5dc5308271f2ce8b64f8d3249dec036bbfea7cad9c51e9c8a901a91

Observation b63c6a9b-9028-4c02-a10a-4aa916a71e91 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.266808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.675575Z digest=sha256:913e3894f436e4955f167ad8717db2d37fff5a7ef4e35ca6e45938968d94a28c

Observation 23b43115-e760-4bb9-acd8-c3f554392c80 · outbound

This paper cites Escape any double quotes in strings.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Escape any double quotes in strings

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.254399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.679000Z digest=sha256:b2fa6cf5e4d32dd2a114617432914c571a1048aaee0eebdbea243115ac3349ae

Observation 6d48f1e6-5549-438b-a06f-b8f566f3bcf9 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.242875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.682219Z digest=sha256:b962750c249c4a45ef40d8bbd0c9c8b88e479568e861422361352f60e83cf148

Observation b183074e-209c-4edd-a5df-01b135f99a57 · outbound

This paper cites If you can autonomously refine the prompt based on feedback (e.g., minor rewording, removing trigger words) to improve the score, do so directly via`query_target`.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you can autonomously refine the prompt based on feedback (e.g., minor rewording, removing trigger words) to improve the score, do so directly via`query_target`

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.231930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.685646Z digest=sha256:bce3cf25ce950e8ad3553246f07d0de532029c0a8940b062e0e8b2938a9867e8

Observation 33c8cf29-4510-432b-bc10-33df6bd63a97 · outbound

This paper cites If you have just called strategies twice, your next action MUST be a`query_target`to test the prompt.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs If you have just called strategies twice, your next action MUST be a`query_target`to test the prompt

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.220163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.689222Z digest=sha256:5f77a6bf7a5a48650a8463252691eda764061e28ee405746a30567eba11ff512

Observation 5cb3ef28-c24b-43cd-81cf-2f8ecd840b56 · outbound

This paper cites Instead, refine the prompt further based on the scorer’s feedback to push it over the success threshold (>8.5).

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Instead, refine the prompt further based on the scorer’s feedback to push it over the success threshold (>8.5)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.207614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.693143Z digest=sha256:70e2400f7ec0c3385151b0c0cd2b9eee4c53653f20f72e34517054091fd813ef

Observation a3d21896-1deb-4502-8584-b820de347158 · outbound

This paper cites Scorer prompt.The scorer model evaluates the harmfulness and actionability of the target model’s response and returns both a scalar score and refinement feedback for the attacker.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Scorer prompt.The scorer model evaluates the harmfulness and actionability of the target model’s response and returns both a scalar score and refinement feedback for the attacker

Reference 49

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.195755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.696936Z digest=sha256:07e534e70e91cef564e9e39e908b7a5f9b8d0ac1ef13af0098ebd839092c2bd5

Observation b8842433-a570-4bde-8154-81d9908cfa00 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.183160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.702180Z digest=sha256:662bcd01c5ddc26adec5fde2b406fee188a04adc1c84856d28699fe0cd7efa8e

Observation e64c0545-1d7a-4130-a918-1d53812aff59 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.171110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.706534Z digest=sha256:772700bbfc7da8b19fecc25b70715f496d9cd67ae06651b6236d397dc92828f9

Observation ca019b35-dfa4-42f8-8e12-ef9403170870 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.159760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.710395Z digest=sha256:11683a9c50e2383a7618245614f5d50fa0c7bc73bc1c0d10807831b819e8d5bd

Observation b2b06089-9a3f-48bd-a7a0-543ae0bd301a · outbound

This paper cites score": <float 1.0–10.0>,.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs score": <float 1.0–10.0>,

Reference 53

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T15:29:06.148618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.714281Z digest=sha256:ff24e1059fbd3cae942ee7ac1aadae65a1d2e952e45076cbec03af4df7310a4e

Observation 0a960579-ef65-4a63-ac3a-6b6a8c68de57 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.135615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.718307Z digest=sha256:c8835c8dbce7735a7d1753275473486f8cebb4d192ca4a96ba251aea4575827f

Observation ac214206-d5fd-4dab-964b-b80725fa80f8 · outbound

This paper cites an unresolved cited work.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:29:06.122950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.722216Z digest=sha256:f8f1c712738093aded8a285804b6b16855d0caafec86666b53e1abd34ad4d802

Observation e0658f8d-8380-4456-823d-38966a0ffdc1 · outbound

This paper cites Response:.

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs Response:

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:29:06.106220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T15:29:05.726369Z digest=sha256:47e7b01e0f14ab5674469b73ee903a7b4f1650074e815ef5e0d038bb40ecad14

Pith citing papers

No inbound Pith citation observations are available.