Pith. sign in

Paper Citation Record · LEDGER

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

As of 12 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2605.05630.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05630 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T07:53:06.928500Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:38:16.610310Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T20:38:54.897445Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact17
  • verified fuzzy25
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9f5c8bdc-7917-4bd1-a9c0-c5ad354326d6 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:57:31.632110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:0f79cef3a60e4e91be7aa6d7e5a40bba1713c2c25deae4c500888c0413dc4b39

Observation 75916d8e-9faf-45f0-ba37-6e3740459295 · outbound

This paper cites Universal jailbreak backdoors in large language model alignment.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Universal jailbreak backdoors in large language model alignment

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.927124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:f2d4072234e0162aec7685647765a4264b7d4f59e4b566712dc4f451ca7b5e5e

Observation 89fca9b5-dff1-44c3-b1f4-b4a2b595b5b8 · outbound

This paper cites Benchmarking Misuse Mitigation Against Covert Adversaries.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Benchmarking Misuse Mitigation Against Covert Adversaries

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:57:31.664087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:4e63d26989c47dc2cee1d62b2ea1379ac8bb7f3b1469cb8e5b3c6fc4690f726b

Observation f78f678a-2c6d-4ebe-a168-2cb0d8d6e3bc · outbound

This paper cites When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.942266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:805aeafb3bf8b40bc5ecaf971f072bd88b0f7d212ef305be6fb4ae9b48c41dc9

Observation ebfb0e98-c9e9-4ab2-ab72-55cff24707fb · outbound

This paper cites Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.645187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:cce23fd975fd2d8b364f458363dfe5980e6706c93d8726f364fb56525d814cf7

Observation 9db8f46f-217b-4543-a34b-6eecedbc27b8 · outbound

This paper cites A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily.North American Chapter of the Association for Computational Linguistics.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily.North American Chapter of the Association for Computational Linguistics

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.934902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:a640a1a2ffd8f84f4edc0971034b4c11da31be68b3c11299d25ebbf66edc2cf8

Observation b3857975-e69f-46fd-8694-7540a07137d1 · outbound

This paper cites Attacks, defenses and evaluations for llm conversation safety: A survey.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Attacks, defenses and evaluations for llm conversation safety: A survey

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.930908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:6334b61376e34873fc6ee00b33ce995e9f0e93757a597af503cb40b68f14f160

Observation e132bfc6-f458-4968-bb02-2054c8aeab10 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:57:31.670860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:afb9898acd08c914c26fb46be7e50cd23f779183c1df535cfe31f7ca7e66a324

Observation 740b670b-74a2-40ad-8cec-4fe7e4eb5134 · outbound

This paper cites Mtsa: Multi-turn safety alignment for llms through multi-round red-teaming.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Mtsa: Multi-turn safety alignment for llms through multi-round red-teaming

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.923013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:707275b93c223dd7ee14400e30e820a8cb8b2377682b48b27a32a9902d211e21

Observation 36150d7d-2fb5-4921-943f-46c053fd6e4f · outbound

This paper cites Harmful prompt classification for large language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Harmful prompt classification for large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.919510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:fc92f8625ed55e90b632a5d6ba4a47fd660a1b19dec219a2e55fc3f50eb49df1

Observation 3c4d917c-6b4a-485c-a223-a63e1132a81d · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:57:31.591542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:cbc742d457d8b8d1bc09f8a6de83ee5e2ab368536e79f6622ca2e7dee8f38801

Observation 0959dda4-fb4d-4767-9404-06500551d91b · outbound

This paper cites Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.915722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:5a74c6e0198dd301d1b1a5d4f5ab3222350bb795f254cc213b093b72dad014c3

Observation d9d7d05f-fe85-4681-a27b-535e3a84c147 · outbound

This paper cites Guard: Role-playing to gener- ate natural-language jailbreakings to test guide- line adherence of large language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guard: Role-playing to gener- ate natural-language jailbreakings to test guide- line adherence of large language models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.609669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:b36d4c07d96f03fbb7b5709210eedba79c03dc880f0ae7bc444a402e47b849dd

Observation e02d5534-5ac1-45f0-a59b-e841c7e3c46a · outbound

This paper cites Mdagents: An adaptive collaboration of llms for medical decision-making.Advances in Neural Information Processing Systems, 37:79410–79452.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Mdagents: An adaptive collaboration of llms for medical decision-making.Advances in Neural Information Processing Systems, 37:79410–79452

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.900019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:f233ff1a9f9b0bf4eff3cc0a124bae6a2815e01e9100a282cc16656ca27bbba6

Observation 0299db71-2a34-44ef-ac66-3c611e44fb7b · outbound

This paper cites Drattack: Prompt de- composition and reconstruction makes powerful llms jailbreakers.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Drattack: Prompt de- composition and reconstruction makes powerful llms jailbreakers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.903591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:f21e02db8d7f4f46ae50752626ee3745ed516e4d00fd1f639fa30b1d5e944335

Observation 06c9a443-6878-4a9d-bbee-fd494ae0ca8a · outbound

This paper cites AutoDAN: Generating stealthy jailbreak prompts on aligned large language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue AutoDAN: Generating stealthy jailbreak prompts on aligned large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.896605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:cd13d64c9141677b72234f6c90135b26501bbd8634021ec4a99fdaacb66e4384

Observation bdf3ec43-17a0-4847-beac-d294515311dc · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T07:57:31.578976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:91fef131ee988ea4d6e72fff6c961363d22fd4a2872ccdfc7c3573195f842918

Observation 9ed64086-8e57-4d7a-9161-0749add5e415 · outbound

This paper cites MALicious INTent dataset and inoculating LLMs for enhanced disinformation detection.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue MALicious INTent dataset and inoculating LLMs for enhanced disinformation detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.906987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:a2e329caadd9423823169e6bee0d9e67ee93a134517a1e1331939944607df463

Observation 0dd0f639-04ef-4738-86a4-6e6dbe736f4b · outbound

This paper cites Helping big language models protect themselves: An enhanced filtering and summariza- tion system.arXiv preprint arXiv:2505.01315.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Helping big language models protect themselves: An enhanced filtering and summariza- tion system.arXiv preprint arXiv:2505.01315

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.567279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:ba6f50aae2a921a7fab45d766c160f0a1cdee7661c7ee0efcae3897554d21c63

Observation f85ff8e8-28ef-4466-9364-237936a1a29c · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.892763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:81245bcbc46f2a6e83a51a34c2423a91f989b639263056491ef921d32c82760e

Observation 713e0187-0d59-4fe0-84d2-a0ab68997008 · outbound

This paper cites Under- standing and mitigating overrefusal in llms from an unveiling perspective of safety decision boundary.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Under- standing and mitigating overrefusal in llms from an unveiling perspective of safety decision boundary

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.888996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:26c6c0777c7b2f09c834f821da5c8034d242a81995f1d60069f74fc27e4ad79a

Observation b901169b-d343-4807-a2ae-f999b4192b6d · outbound

This paper cites Automated Red Teaming with GOAT: the Generative Offensive Agent Tester.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.620956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:fa1d36af9b16895ba4ebc596e1f85c8a446503cebe54bb9309c0a95047bb2870

Observation e0e18198-5df1-4b24-a777-3fb36a7faae7 · outbound

This paper cites X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.626795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:ba83db9c01570cd124cd8bb99d3946c0132b3d155298acea7a56f6608e9a2cc2

Observation e906ecb5-360e-4186-a871-b4f3fba228e8 · outbound

This paper cites Derail yourself: Multi-turn llm jailbreak attack through self- discovered clues.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Derail yourself: Multi-turn llm jailbreak attack through self- discovered clues

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.678460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:9ede05cfa0ec80a2fd1f7c32381997418ba47e4f7dbd49be95cde982a1aa2172

Observation 12e04150-aa1e-4b4c-ae7d-b674abde3880 · outbound

This paper cites Llms know their vulnerabilities: Uncover safety gaps through natural distribution shifts.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llms know their vulnerabilities: Uncover safety gaps through natural distribution shifts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.881502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:8cbb772dd9c2c4bc4e75421ef1fbf6140bcb41e0767ef98033e8c136f1e07433

Observation 0036137a-b777-4228-b85a-73bf3fa6172c · outbound

This paper cites Xstest: A test suite for identifying exaggerated safety behaviours in large language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Xstest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.873952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:b8749f6a9d803bde04b3359e1ef00a1778b85826726d28cf02ac09d1e2b426fc

Observation 01eb5443-436e-4bd2-8f16-e1629e75369a · outbound

This paper cites Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Great, now write an article about that: The crescendo {Multi-Turn}{LLM} jailbreak attack

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.878068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:2c0a50aef3ec0234f0b6670036ede31d1e11457f439e9c622ab70716ebdab882

Observation ee96c0bf-1e18-471d-836e-b1aa4f4c3912 · outbound

This paper cites Llms in software security: A survey of vulnerability detection techniques and insights.ACM Computing Surveys, 58(5):1–35.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Llms in software security: A survey of vulnerability detection techniques and insights.ACM Computing Surveys, 58(5):1–35

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.885365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:11f0a8397533f40abebcfd6d53e7a786e998a569514159b0fd89d2e6dbd81536

Observation c76ea605-4480-48fc-8fd8-f6a6cae98125 · outbound

This paper cites Safe in isolation, dangerous together: Agent-driven multi- turn decomposition jailbreaks on LLMs.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Safe in isolation, dangerous together: Agent-driven multi- turn decomposition jailbreaks on LLMs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.869745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:5c2bdb133c6c31dc7e19728e8145cfa3e5ccac2b2e58ee1f7ad84058b7337269

Observation 243e972b-6898-4523-b033-64922b8fe830 · outbound

This paper cites an unresolved cited work.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-13T07:57:32.864569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:76be9e65bc7961bd48f79151f4f34a4fa846e4bb6178b967aa1753d6da730b5d

Observation bd9a4b83-5810-431f-ba96-96acd2a4195d · outbound

This paper cites RoleBreak: Character hallucination as a jailbreak attack in role-playing systems.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue RoleBreak: Character hallucination as a jailbreak attack in role-playing systems

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.860681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:337afb157e16bfca8f5cd1886e32c10426e05a6f9e10858ec94f9912dc5ce2e2

Observation 75345be8-5153-4eac-b542-c5cd256fb9b1 · outbound

This paper cites Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Prompt, Divide, and Conquer: Bypassing Large Language Model Safety Filters via Segmented and Distributed Prompt Processing

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.638947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:d0d930cb319c83d34754fc0bba88ac9fd69f6b65421c78f14da33abe901849e5

Observation a4363bf8-9826-416b-81c3-532ecf4351c3 · outbound

This paper cites Do llms really forget? evaluating unlearning with knowledge correlation and confidence awareness.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Do llms really forget? evaluating unlearning with knowledge correlation and confidence awareness

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.615277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:e60ab2a3623524c33252c7a202eba9868b702973dcf33e88625f77fe0dac960c

Observation ea262b3d-35de-4b08-a360-b0c8b4fd9177 · outbound

This paper cites The trojan knowledge: Bypassing commercial LLM guardrails via harmless prompt weaving and adaptive tree search.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue The trojan knowledge: Bypassing commercial LLM guardrails via harmless prompt weaving and adaptive tree search

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.598074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:d092d476fd8c66ad550bbce26fad21ed5ee31b3175e9f71a9644a317802e5def

Observation 6e07299f-5b16-43b2-8667-2130f86a1661 · outbound

This paper cites RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:57:31.572737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:f1670a4bba08192ab47df021310e0eaf8e6c38ffc1d4e214d196d6234d21113e

Observation c3dcbaeb-85bb-4a43-a5f8-3ee2b02e116f · outbound

This paper cites Chain of attack: Hide your intention through multi-turn interrogation.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Chain of attack: Hide your intention through multi-turn interrogation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.856337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:36bb897a56ed863216b1512f1524029e2feef006745832274fe8e00f4aff633b

Observation b6d226c8-474d-485d-90d1-cdd3e69bc946 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Low-Resource Languages Jailbreak GPT-4

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:24:14.203470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:68ea089297aaa2a94984cf81ba5d5549a595fb23c5c0bd07c1b7c12912e66686

Observation bf6a401c-29f4-4f08-8bf5-5470b59922ae · outbound

This paper cites Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:57:31.586472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:53970ca743c59cef44ffa08b4435520ead90b0c8f0edd4b3fac86da4051cc02a

Observation 9c3d6022-748a-4c66-b3d8-d7f6c01daa7a · outbound

This paper cites Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-29T00:24:32.402891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:1b1b03829000ade424acb4652d8e9dfc8dfca95fa0d995cbff7ad49f0770b5cd

Observation 341417bf-ea3e-42fd-848f-2951adf7dca5 · outbound

This paper cites DAMON: A dialogue-aware MCTS framework for jailbreaking large language models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue DAMON: A dialogue-aware MCTS framework for jailbreaking large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.848385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:eab50e3365c4c7db0a9639dbccc620cadefaed554c58a548e731f7ecf68d3e8a

Observation 1e128392-699d-478e-ab93-85e4a71fffb3 · outbound

This paper cites Intention analysis makes llms a good jailbreak defender.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Intention analysis makes llms a good jailbreak defender

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.844494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:afdc44b86fc16ef5b0bfe455bd1492b6f64138457f680f96bed3fa9c89f1961f

Observation b0573178-2b7b-428d-b7be-9ecaef243605 · outbound

This paper cites an unresolved cited work.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-13T07:57:32.851910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:892e5c904780d069e82e649aa2c6fba635caab06402b2ee3473026dbbbe2a260

Observation dfe079d6-5def-4232-8c0e-88f2c01d0992 · outbound

This paper cites Qwen3Guard Technical Report.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Qwen3Guard Technical Report

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:33:37.835217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:a991687f5af558f3c159907b87b8eae450f96de9ad45fd4ebb42479ddd3bff80

Observation a39de01f-bfe8-44e8-a071-c5bf51f894c7 · outbound

This paper cites How alignment and jailbreak work: Explain llm safety through intermediate hidden states.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue How alignment and jailbreak work: Explain llm safety through intermediate hidden states

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.911223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:dd7aef09ac9fcc73f27ad2695ea682d1b735eed843be73acf53b722e9ba50614

Observation b9478d0c-3645-425a-a4ed-d18a15886c46 · outbound

This paper cites Improving alignment and robustness with circuit breakers.Advances in Neural Information Processing Systems, 37:83345– 83373.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Improving alignment and robustness with circuit breakers.Advances in Neural Information Processing Systems, 37:83345– 83373

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.835812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:8852738158080ef263dbb750df6e1bc44caded1fd819aeab811d51357d3a8a17

Observation 98d77232-8e32-4708-a8e0-2d753668b23d · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:57:31.603876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:e27759b8a0fea21c546525e55e5adc0d5c30f69e5ddabf33002b6ecff3f7dd19

Observation 5aea485c-9738-47f5-9723-071b12d54600 · outbound

This paper cites an unresolved cited work.

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-13T07:57:32.831853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:b9e0da3beb0168d058530ca004a42fb786ca129c961fde3bcd240b6046d72ba5

Observation 327ab0f2-d5ec-49bb-b311-9e03c137afbb · outbound

This paper cites index" (0-based).

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue index" (0-based)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T07:57:32.840740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T07:53:06.928500Z digest=sha256:5093dcc561a52e61c2fc9340ea7fc756629cea8332d959263df7d6ae660a0fe1

Pith citing papers

Observation c0df8b7e-8afc-435d-917b-790524933707 · inbound

Investigating and Alleviating Harm Amplification in LLM Interactions cites this paper.

Investigating and Alleviating Harm Amplification in LLM Interactions One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:16:24.014149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T14:31:52.027889Z digest=sha256:ad1b65f8082735683960e361dae9c897599b97f3d2b5ac115b03666faef8cd03

Observation ca7c218f-c25e-460d-b32d-cbeab855beed · inbound

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety cites this paper.

Cognitive Firewall: A Proactive, Zero-Trust, Multi-Gate Framework for LLM Safety One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:38:54.898605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T20:38:16.610310Z digest=sha256:2e6ed052a3f74765be51f1368aa12a9eb3d2c5e8e7d3ccaf500c55b6bbe0843c