Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2309.02705.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:05:54.141304Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T10:26:11.131634Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 631bfaad-3fa2-4fae-ac76-5779a5c67dd1 · inbound
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Certifying LLM Safety against Adversarial Prompting
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 16e16c7d-57e5-4dd9-933b-72ca8b2c757f · inbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ed3bd17f-e334-4801-aab6-521ee2ee6ffa · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey Certifying LLM Safety against Adversarial Prompting
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 30c23f4e-9525-4707-a82a-3fed24844c62 · inbound
Trustworthiness in Retrieval-Augmented Generation Systems: A Survey Certifying LLM Safety against Adversarial Prompting
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 067b97c0-090d-4a19-94ae-c5e60c6c44d9 · inbound
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts? Certifying LLM Safety against Adversarial Prompting
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9687f5e-8ce3-42c1-9d87-e8cecd757124 · inbound
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation Certifying LLM Safety against Adversarial Prompting
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63f34c02-d67b-46db-a2b7-3cc9c6b7e774 · inbound
Large Language Model Safety: A Holistic Survey Certifying LLM Safety against Adversarial Prompting
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8e2bfd-f142-4877-9547-c03c44f799ff · inbound
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1d2ffc2-ae93-4667-bcc7-2d7e1f5abd91 · inbound
On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Certifying LLM Safety against Adversarial Prompting
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dccb1ad-7a00-4758-8163-e94637e9e88d · inbound
Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Certifying LLM Safety against Adversarial Prompting
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba88c8c3-1e8b-4b93-8a26-bc931f53e8bf · inbound
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints Certifying LLM Safety against Adversarial Prompting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc496c7-288c-4817-adef-d181b7edf0c2 · inbound
Smoothed Embeddings for Robust Language Models Certifying LLM Safety against Adversarial Prompting
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3d688a-2342-46c6-ab35-c42f30f790ec · inbound
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Certifying LLM Safety against Adversarial Prompting
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7368f9-3cd2-407b-a5fb-7a0e6b7d2f2e · inbound
Training Users Against Human and GPT-4 Generated Social Engineering Attacks Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9828062f-5849-40d1-9b97-d663117bbc92 · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Certifying LLM Safety against Adversarial Prompting
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 38238e41-52d7-45dd-abf9-04724b17db53 · inbound
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Certifying LLM Safety against Adversarial Prompting
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fa7f62-9199-4a79-b8a9-dcac1e75db20 · inbound
Towards Copyright Protection for Knowledge Bases of Retrieval-augmented Language Models via Reasoning Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddb39bc-b53c-44c9-9040-4e0465eff828 · inbound
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression Certifying LLM Safety against Adversarial Prompting
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7e5b90-202f-4a04-ab93-b32cd90ddba8 · inbound
Security of Internet of Agents: Attacks and Countermeasures Certifying LLM Safety against Adversarial Prompting
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d8eba85-300b-4352-9b4e-9cf80184c1d6 · inbound
Adversarial Suffix Filtering: a Defense Pipeline for LLMs Certifying LLM Safety against Adversarial Prompting
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f750fb9a-4f20-4a8f-b8de-afa25a802156 · inbound
Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks Certifying LLM Safety against Adversarial Prompting
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dffa9100-9c3a-4ca8-84fa-6049a977eb8f · inbound
LLM-Powered AI Agent Systems and Their Applications in Industry Certifying LLM Safety against Adversarial Prompting
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f3a082e9-bc92-40ec-abb7-0c054b0d020d · inbound
Superplatforms Have to Attack AI Agents Certifying LLM Safety against Adversarial Prompting
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3eb8416-463e-4425-b4f4-9836979c048e · inbound
Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05916946-9f08-4808-8aed-d87770cbb58d · inbound
Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Certifying LLM Safety against Adversarial Prompting
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96294aca-ada6-4830-a85d-48d6c525f494 · inbound
Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis Certifying LLM Safety against Adversarial Prompting
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee86101-436a-4575-adaa-a245aed87060 · inbound
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques Certifying LLM Safety against Adversarial Prompting
Reference 187
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f793b9-b949-4dce-9e30-e3e82aa31e24 · inbound
Strategic Deflection: Defending LLMs from Logit Manipulation Certifying LLM Safety against Adversarial Prompting
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · inbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c02eec3d-2f04-4f2e-9a7c-0af33298520a · inbound
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Certifying LLM Safety against Adversarial Prompting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b89c6b03-55ef-4ed4-9aa2-6f69f0c2266d · inbound
CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection Certifying LLM Safety against Adversarial Prompting
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72fec629-b593-4885-80d1-c63beb3e8e61 · inbound
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain Certifying LLM Safety against Adversarial Prompting
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 758bf927-a815-4e7b-a2c5-a5e6843da869 · inbound
AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents Certifying LLM Safety against Adversarial Prompting
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 039e0742-22f3-4a98-9c64-534d84230ae3 · inbound
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 510fa07b-1d80-43a3-9c68-e3914854dc9c · inbound
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Certifying LLM Safety against Adversarial Prompting
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7068a9da-7e48-4fb7-af66-9367cc3fce90 · inbound
BEAVER: An Efficient Deterministic LLM Verifier Certifying LLM Safety against Adversarial Prompting
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 742d023c-c22e-41d8-9422-6fb186a319dc · inbound
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents Certifying LLM Safety against Adversarial Prompting
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 67eb77b1-719c-4bfc-958c-616813410f69 · inbound
ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying Certifying LLM Safety against Adversarial Prompting
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f77ade4b-66b0-431f-b006-0d2153c9f251 · inbound
Towards Understanding the Robustness of Sparse Autoencoders Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f4f54fa7-3020-4fb1-9dbb-b50c0d9f1373 · inbound
Ethics Testing: Proactive Identification of Generative AI System Harms Certifying LLM Safety against Adversarial Prompting
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8e02b80d-2e64-4ee1-a04f-f5a2a82b6032 · inbound
Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing Certifying LLM Safety against Adversarial Prompting
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a1572414-af7d-48a5-8cd2-ba636ab52bfa · inbound
Green Shielding: A User-Centric Approach Towards Trustworthy AI Certifying LLM Safety against Adversarial Prompting
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a48b12bb-7580-43b9-b7f4-649c91f08507 · inbound
Attention Is Where You Attack Certifying LLM Safety against Adversarial Prompting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 476d33b0-1915-4757-8903-a0c31063cdda · inbound
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Certifying LLM Safety against Adversarial Prompting
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4508b4c4-07f6-454d-b379-b70604ee8b4d · inbound
How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Certifying LLM Safety against Adversarial Prompting
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed216ea4-a6eb-4ff1-8647-fc9c1db40474 · inbound
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Certifying LLM Safety against Adversarial Prompting
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 64ca20fe-1ed2-4694-9d5e-249a255c72dc · inbound
Re-Triggering Safeguards within LLMs for Jailbreak Detection Certifying LLM Safety against Adversarial Prompting
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 13365dc7-0a85-4bc8-9853-39a7b5d1634b · inbound
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models Certifying LLM Safety against Adversarial Prompting
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 57000007-7040-4aba-8176-8cb457f56708 · inbound
Adversarial Reframing: A Framework for Targeted Generation in Language Models Certifying LLM Safety against Adversarial Prompting
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bc19f71c-d222-44f1-8b1b-e02761f36024 · inbound
Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents Certifying LLM Safety against Adversarial Prompting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation feb1401b-66ba-44d2-8ea8-a30a62a14e04 · inbound
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Certifying LLM Safety against Adversarial Prompting
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7fcce9d9-0859-4eb1-94ae-65d4e74112df · inbound
Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions Certifying LLM Safety against Adversarial Prompting
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6ca8f49f-5d36-498f-aaf9-9beb9e4c8a40 · inbound
MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 771bd9f5-e888-44a8-a10d-77ae5df37a46 · inbound
Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Certifying LLM Safety against Adversarial Prompting
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.