Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T07:53:06.928500Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2605.05630.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T07:53:06.928500Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-03T20:38:16.610310Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T20:38:54.897445Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9f5c8bdc-7917-4bd1-a9c0-c5ad354326d6 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 75916d8e-9faf-45f0-ba37-6e3740459295 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Universal jailbreak backdoors in large language model alignment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 89fca9b5-dff1-44c3-b1f4-b4a2b595b5b8 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Benchmarking Misuse Mitigation Against Covert Adversaries
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f78f678a-2c6d-4ebe-a168-2cb0d8d6e3bc · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ebfb0e98-c9e9-4ab2-ab72-55cff24707fb · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9db8f46f-217b-4543-a34b-6eecedbc27b8 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily.North American Chapter of the Association for Computational Linguistics
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3857975-e69f-46fd-8694-7540a07137d1 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Attacks, defenses and evaluations for llm conversation safety: A survey
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 740b670b-74a2-40ad-8cec-4fe7e4eb5134 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Mtsa: Multi-turn safety alignment for llms through multi-round red-teaming
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 36150d7d-2fb5-4921-943f-46c053fd6e4f · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Harmful prompt classification for large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3c4d917c-6b4a-485c-a223-a63e1132a81d · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0959dda4-fb4d-4767-9404-06500551d91b · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d9d7d05f-fe85-4681-a27b-535e3a84c147 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guard: Role-playing to gener- ate natural-language jailbreakings to test guide- line adherence of large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e02d5534-5ac1-45f0-a59b-e841c7e3c46a · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Mdagents: An adaptive collaboration of llms for medical decision-making.Advances in Neural Information Processing Systems, 37:79410–79452
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0299db71-2a34-44ef-ac66-3c611e44fb7b · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Drattack: Prompt de- composition and reconstruction makes powerful llms jailbreakers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06c9a443-6878-4a9d-bbee-fd494ae0ca8a · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bdf3ec43-17a0-4847-beac-d294515311dc · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9ed64086-8e57-4d7a-9161-0749add5e415 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue MALicious INTent dataset and inoculating LLMs for enhanced disinformation detection
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0dd0f639-04ef-4738-86a4-6e6dbe736f4b · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Helping big language models protect themselves: An enhanced filtering and summariza- tion system.arXiv preprint arXiv:2505.01315
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f85ff8e8-28ef-4466-9364-237936a1a29c · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 713e0187-0d59-4fe0-84d2-a0ab68997008 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Under- standing and mitigating overrefusal in llms from an unveiling perspective of safety decision boundary
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b901169b-d343-4807-a2ae-f999b4192b6d · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e0e18198-5df1-4b24-a777-3fb36a7faae7 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e906ecb5-360e-4186-a871-b4f3fba228e8 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Derail yourself: Multi-turn llm jailbreak attack through self- discovered clues
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 12e04150-aa1e-4b4c-ae7d-b674abde3880 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llms know their vulnerabilities: Uncover safety gaps through natural distribution shifts
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0036137a-b777-4228-b85a-73bf3fa6172c · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Xstest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01eb5443-436e-4bd2-8f16-e1629e75369a · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ee96c0bf-1e18-471d-836e-b1aa4f4c3912 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llms in software security: A survey of vulnerability detection techniques and insights.ACM Computing Surveys, 58(5):1–35
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c76ea605-4480-48fc-8fd8-f6a6cae98125 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Safe in isolation, dangerous together: Agent-driven multi- turn decomposition jailbreaks on LLMs
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 243e972b-6898-4523-b033-64922b8fe830 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd9a4b83-5810-431f-ba96-96acd2a4195d · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue RoleBreak: Character hallucination as a jailbreak attack in role-playing systems
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 75345be8-5153-4eac-b542-c5cd256fb9b1 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a4363bf8-9826-416b-81c3-532ecf4351c3 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Do llms really forget? evaluating unlearning with knowledge correlation and confidence awareness
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ea262b3d-35de-4b08-a360-b0c8b4fd9177 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue The trojan knowledge: Bypassing commercial LLM guardrails via harmless prompt weaving and adaptive tree search
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e07299f-5b16-43b2-8667-2130f86a1661 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c3dcbaeb-85bb-4a43-a5f8-3ee2b02e116f · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Chain of attack: Hide your intention through multi-turn interrogation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b6d226c8-474d-485d-90d1-cdd3e69bc946 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Low-Resource Languages Jailbreak GPT-4
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf6a401c-29f4-4f08-8bf5-5470b59922ae · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 341417bf-ea3e-42fd-848f-2951adf7dca5 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue DAMON: A dialogue-aware MCTS framework for jailbreaking large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1e128392-699d-478e-ab93-85e4a71fffb3 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Intention analysis makes llms a good jailbreak defender
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b0573178-2b7b-428d-b7be-9ecaef243605 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dfe079d6-5def-4232-8c0e-88f2c01d0992 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Qwen3Guard Technical Report
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a39de01f-bfe8-44e8-a071-c5bf51f894c7 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue How alignment and jailbreak work: Explain llm safety through intermediate hidden states
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9478d0c-3645-425a-a4ed-d18a15886c46 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Improving alignment and robustness with circuit breakers.Advances in Neural Information Processing Systems, 37:83345– 83373
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98d77232-8e32-4708-a8e0-2d753668b23d · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5aea485c-9738-47f5-9723-071b12d54600 · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 327ab0f2-d5ec-49bb-b311-9e03c137afbb · outbound
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue index" (0-based)
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c0df8b7e-8afc-435d-917b-790524933707 · inbound
Investigating and Alleviating Harm Amplification in LLM Interactions One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca7c218f-c25e-460d-b32d-cbeab855beed · inbound
Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.