Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:31.644019Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2505.18556.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:31.644019Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T13:18:40.623601Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T17:51:41.885862Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b431b791-0fd2-4ac2-90f8-61f550819cd4 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3172e170-78cf-451b-a001-86059b6107a0 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41aa9c5-5712-43fd-88d4-784ad7b7964d · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b126d2e1-42ec-4643-b165-ef19f7956778 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09c0e544-7a5d-481a-80b9-12e2c9a810fa · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ee2aaab-b997-4149-aaac-e5d97dde2032 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506c0bed-abf0-423f-8d97-e31c7977d43e · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0ffcb1-bb92-4f76-a247-4ef24ff1250b · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Z-BERT-A: a zero-shot Pipeline for Unknown Intent detection
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation de1eedfa-e02d-4ed1-9394-bf001ae8372d · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c92bd8b-b5ce-4772-aa71-25501a945267 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2b84897-36d2-4d2b-84ef-9eaac98ee83e · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83ece4b4-ec12-41c9-9c8b-c471793193ec · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31a650b5-5135-4391-bedc-1ad88c054921 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a8a5e0b-ea4e-4b82-bd59-b5d088003154 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882510a1-c0c5-42ec-88dd-476aca8704e5 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation OpenAI o1 System Card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f64d5f7-277b-49ed-9b32-797732415312 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df9b1b16-5935-4d78-86de-5a319a58074c · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e6f370-574e-401e-8292-f543d35becb0 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 81e35d6b-b051-47dc-b881-75a029c5c93b · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3251d485-49a5-4b94-835f-62a66192a6df · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 242c95c8-92d6-4104-b7ba-e7860536e439 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation DeepSeek-V3 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f4dd05-014a-4240-89c4-2321628b6b58 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation FlipAttack: Jailbreak LLMs via Flipping
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d61b577-7030-4b17-a34d-833315d29f99 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5991bb0f-c38b-479b-86b3-021d59dccfc8 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e1c95f-3cdb-4dd7-939e-16de1224f06f · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4601b9e5-8120-418b-a1fd-0461eb6ece5c · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd04e60b-495f-4d0b-859a-a2c0b99cf229 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 48f16cf2-8790-4d24-b2d3-f2a00cc4603e · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f88d5ee-0bcf-4695-9c52-275fba0d9acb · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation IntentGPT: Few-shot Intent Discovery with Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ce1dea-3956-4aaf-bb0c-193073a598e0 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3d3b302f-04de-4aeb-ad6b-60458330f346 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 003315bc-28a4-43f6-b6d9-af64111e6ab6 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef5bb73a-5f1c-4140-ab2f-96bfc7a71cc3 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61305ab5-bfa8-431c-a5a8-6c4fadff7ba4 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50dc41cf-1536-4e7b-869a-64ad6d1d0615 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf8c749-40c7-4757-a814-50bc26cf75e1 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 739e6a25-7f9e-4c5f-9925-4b25d9d16e52 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 213c098c-63ae-4d11-aee0-962ef541656e · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 791bdf0f-7e6b-4e1e-bcd8-d243c270fae0 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fcdcaf-b1b8-419a-b03b-8dc884c67fa1 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041362cf-344d-431d-9f65-b2f644b51495 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4cf9ed77-4cdd-4c78-ba1c-76616ac9bf66 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deb7a500-333c-4e9e-af3a-52a49d38e0eb · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7549e762-cf0c-45fa-9cd5-6c51e9a911ec · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cafffd9b-ee29-49b4-81c8-a09b699277a3 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 306dfaa0-926f-4e67-9edb-c8e95cb94cd8 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 363920a5-7472-43d6-bfb9-c5b14637ff60 · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c3068a1-aab5-4bc1-8738-4a37c628852f · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Autodan: Interpretable gradient-based adversarial attacks on large language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06bbe084-a062-490a-81fd-b9a1ac80b7ae · outbound
Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a48bd81-4ad5-4d9e-b690-0bbcb0aab7c9 · inbound
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 49805767-a29c-4379-acc2-cebec45087db · inbound
Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d40ec3d7-7faa-417d-b17c-e93b52d8736f · inbound
Incomplete Prompt Jailbreaks in Large Language Models Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.