Pith. sign in

Paper Citation Record · LEDGER

Certifying LLM Safety against Adversarial Prompting

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2309.02705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.02705 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:05:54.141304Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T10:26:11.131634Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 631bfaad-3fa2-4fae-ac76-5779a5c67dd1 · inbound

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks cites this paper.

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Certifying LLM Safety against Adversarial Prompting

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:11:00.805894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-14T17:11:00.639293Z digest=sha256:2be6006a2719d044c90c86eebaa88c26509f1a6b45fbd6c3b41a6ac14cc3bf88

Observation 16e16c7d-57e5-4dd9-933b-72ca8b2c757f · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Certifying LLM Safety against Adversarial Prompting

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:08:05.632731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:d29e00aac9435cf7dc7c7831fc2fca1de3d9f9be275f12e9c240dea19661ed02

Observation ed3bd17f-e334-4801-aab6-521ee2ee6ffa · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Certifying LLM Safety against Adversarial Prompting

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.701189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:940febd38b2a333ebf79d8bed3f6c1a00bba21bd8025b3f97e515c9b579d07a3

Observation 30c23f4e-9525-4707-a82a-3fed24844c62 · inbound

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey cites this paper.

Trustworthiness in Retrieval-Augmented Generation Systems: A Survey Certifying LLM Safety against Adversarial Prompting

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:08:25.977185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T21:08:11.787013Z digest=sha256:b8684ab02ba92ded4acc3932a535118297e3b0299f46f2f91210d8f74c03d565

Observation 067b97c0-090d-4a19-94ae-c5e60c6c44d9 · inbound

Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts? cites this paper.

Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts? Certifying LLM Safety against Adversarial Prompting

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T22:41:27.545838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:41:27.545838Z digest=sha256:07de7859f5be8dee9308224f5a3aea2d40afa6104720953fc4d19b43d7209810

Observation b9687f5e-8ce3-42c1-9d87-e8cecd757124 · inbound

Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation cites this paper.

Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation Certifying LLM Safety against Adversarial Prompting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:57:06.202073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:57:06.202073Z digest=sha256:135b1dce55249a9347e68419841fe94a1ac946e5cf63fb9819a1a9fa9d0fc157

Observation 63f34c02-d67b-46db-a2b7-3cc9c6b7e774 · inbound

Large Language Model Safety: A Holistic Survey cites this paper.

Large Language Model Safety: A Holistic Survey Certifying LLM Safety against Adversarial Prompting

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-11T05:19:34.409300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:19:34.409300Z digest=sha256:44d7d70d5688d652b43834969b4f9fb8a1f420fc3b65b5db55612df2b50f4428

Observation 1c8e2bfd-f142-4877-9547-c03c44f799ff · inbound

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models cites this paper.

Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models Certifying LLM Safety against Adversarial Prompting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:26.657326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:26.657326Z digest=sha256:b86b287c514b1beebd604e86d7c2827257736ac8336157955a594d3797feeadc

Observation b1d2ffc2-ae93-4667-bcc7-2d7e1f5abd91 · inbound

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs cites this paper.

On the Validity of Traditional Vulnerability Scoring Systems for Adversarial Attacks against LLMs Certifying LLM Safety against Adversarial Prompting

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:18.863343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:40:18.863343Z digest=sha256:7ca7b954d6dc09411e964d77f5cd0a41407043f9a3a517a8c886d43bae7dc8b7

Observation 8dccb1ad-7a00-4758-8163-e94637e9e88d · inbound

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense cites this paper.

Layer-Level Self-Exposure and Patch: Affirmative Token Mitigation for Jailbreak Attack Defense Certifying LLM Safety against Adversarial Prompting

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T22:12:00.047226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:12:00.047226Z digest=sha256:dd346cf0af84caf32df4759173bbfdc5e49bd456caaf77af7fd241d7f231a383

Observation ba88c8c3-1e8b-4b93-8a26-bc931f53e8bf · inbound

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints cites this paper.

Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints Certifying LLM Safety against Adversarial Prompting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:36.508096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:34:36.508096Z digest=sha256:17453c6f9ad1d8504ee522aa5de2343beb5ee3ea1b4a42c51ef41adfa082ebf2

Observation edc496c7-288c-4817-adef-d181b7edf0c2 · inbound

Smoothed Embeddings for Robust Language Models cites this paper.

Smoothed Embeddings for Robust Language Models Certifying LLM Safety against Adversarial Prompting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T12:52:40.928821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:52:40.928821Z digest=sha256:61fbc8c0207ace481d938d2de0883a4c84a612ad4ede69061035a2a489c65be9

Observation da3d688a-2342-46c6-ab35-c42f30f790ec · inbound

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models cites this paper.

Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models Certifying LLM Safety against Adversarial Prompting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T00:12:04.836347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:12:04.836347Z digest=sha256:bb588318b95ecc4c9e91d90e58fd4f5810e0fa46dc1ddf570b1c4c2e2a807887

Observation 5a7368f9-3cd2-407b-a5fb-7a0e6b7d2f2e · inbound

Training Users Against Human and GPT-4 Generated Social Engineering Attacks cites this paper.

Training Users Against Human and GPT-4 Generated Social Engineering Attacks Certifying LLM Safety against Adversarial Prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T14:37:39.773175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T14:37:39.773175Z digest=sha256:2ed57af396e8e882e0b7a6af322c8793549a7c37bbb73d04acaeb47f315c2602

Observation 9828062f-5849-40d1-9b97-d663117bbc92 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Certifying LLM Safety against Adversarial Prompting

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:42:33.899921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:a9d89eed44bb77f1bc992c3a8883205033434c5652611ca91dfba8c1a40db2f1

Observation 38238e41-52d7-45dd-abf9-04724b17db53 · inbound

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences cites this paper.

Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences Certifying LLM Safety against Adversarial Prompting

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T10:23:06.144068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:23:06.144068Z digest=sha256:1e02f27048751e4bf97ce1b4f9f6519499a8d1a66e1f47a7bd658ce6db1cce20

Observation 42fa7f62-9199-4a79-b8a9-dcac1e75db20 · inbound

Towards Copyright Protection for Knowledge Bases of Retrieval-augmented Language Models via Reasoning cites this paper.

Towards Copyright Protection for Knowledge Bases of Retrieval-augmented Language Models via Reasoning Certifying LLM Safety against Adversarial Prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:16:48.271400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:16:48.271400Z digest=sha256:b09f876006784263594fd82ab7d91a5dcf375a663a578141b406e4547e66afa0

Observation fddb39bc-b53c-44c9-9040-4e0465eff828 · inbound

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression cites this paper.

PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression Certifying LLM Safety against Adversarial Prompting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:05:54.141304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:05:54.141304Z digest=sha256:9943a6f3c764b3b6c6dd88fb6116d8936da17952f6dfcae0199e98d9926ebcaa

Observation 5a7e5b90-202f-4a04-ab93-b32cd90ddba8 · inbound

Security of Internet of Agents: Attacks and Countermeasures cites this paper.

Security of Internet of Agents: Attacks and Countermeasures Certifying LLM Safety against Adversarial Prompting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:26:17.933498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:26:17.933498Z digest=sha256:382911600ab1a9799caffbadd0ca7904bf40031786d105c98abc109d7f5c7c44

Observation 7d8eba85-300b-4352-9b4e-9cf80184c1d6 · inbound

Adversarial Suffix Filtering: a Defense Pipeline for LLMs cites this paper.

Adversarial Suffix Filtering: a Defense Pipeline for LLMs Certifying LLM Safety against Adversarial Prompting

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:32:35.011638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:32:35.011638Z digest=sha256:cfccc7518957367d66e7cda861d45aeefc5bbcb922a0ba04cba8bd8e24a061b3

Observation f750fb9a-4f20-4a8f-b8de-afa25a802156 · inbound

Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks cites this paper.

Investigating the Vulnerability of LLM-as-a-Judge Architectures to Prompt-Injection Attacks Certifying LLM Safety against Adversarial Prompting

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:09.836301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:09.836301Z digest=sha256:8b708a2671e10908dfd160e11f5f77e199911ea3ff7a3fc188c451418db1fe29

Observation dffa9100-9c3a-4ca8-84fa-6049a977eb8f · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry Certifying LLM Safety against Adversarial Prompting

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:06:37.893296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:45c8a70e7ec224a043a472548da0b1dc843d7280a048bd0b8538c9b597aa7205

Observation f3a082e9-bc92-40ec-abb7-0c054b0d020d · inbound

Superplatforms Have to Attack AI Agents cites this paper.

Superplatforms Have to Attack AI Agents Certifying LLM Safety against Adversarial Prompting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:42:37.234824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:42:37.234824Z digest=sha256:8c9338f95b4ea99535e2726a0876bfc84f26b9c3d9829ff77fc08a26264da3cc

Observation d3eb8416-463e-4425-b4f4-9836979c048e · inbound

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models cites this paper.

Test-Time Immunization: A Universal Defense Framework Against Jailbreaks for (Multimodal) Large Language Models Certifying LLM Safety against Adversarial Prompting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:21.099918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:21.099918Z digest=sha256:972cf65b9d63f2294630adcee779fd9b2f93b51ebb9e1f785df088dde4c0b306

Observation 05916946-9f08-4808-8aed-d87770cbb58d · inbound

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap cites this paper.

Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap Certifying LLM Safety against Adversarial Prompting

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:21.159164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:37:21.159164Z digest=sha256:e9809783a6db818291afdd2ee703b214c6f10e900a4f483768f20c5b881b7cfe

Observation 96294aca-ada6-4830-a85d-48d6c525f494 · inbound

Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis cites this paper.

Safe and Performant Deployment of Autonomous Systems via Model Predictive Control and Hamilton-Jacobi Reachability Analysis Certifying LLM Safety against Adversarial Prompting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:50:27.924372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:50:27.924372Z digest=sha256:bae3a62b912b1a99113a783a3898daffbac4f61345ceb5809ff591806875b307

Observation 0ee86101-436a-4575-adaa-a245aed87060 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques Certifying LLM Safety against Adversarial Prompting

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:30.518383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:30.518383Z digest=sha256:5865145c4151c0e8c7d913339afed936ac74b3306ffccfbd01b622d0e99d87e4

Observation 15f793b9-b949-4dce-9e30-e3e82aa31e24 · inbound

Strategic Deflection: Defending LLMs from Logit Manipulation cites this paper.

Strategic Deflection: Defending LLMs from Logit Manipulation Certifying LLM Safety against Adversarial Prompting

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.541606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.541606Z digest=sha256:6822d3b126e2453d1960e65d9bce9f1fcf9fbacaf22fec7f49323869ca9c35ea

Observation df7a80e7-73ef-4c35-b707-6f4833e9d6b6 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Certifying LLM Safety against Adversarial Prompting

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T23:12:59.802317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:12:59.802317Z digest=sha256:8f2faef0c72a037f89d53503029cd2759241bf1c858aa4889036c1210e291d7c

Observation c02eec3d-2f04-4f2e-9a7c-0af33298520a · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Certifying LLM Safety against Adversarial Prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.403420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.403420Z digest=sha256:3f7944b391f2ce9160540881b4938e06ea68b77e003f9004d86212c1238ac8ac

Observation b89c6b03-55ef-4ed4-9aa2-6f69f0c2266d · inbound

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection cites this paper.

CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection Certifying LLM Safety against Adversarial Prompting

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:18:24.068355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:18:24.068355Z digest=sha256:d3a7ee3f6f7a5f37a0166531a285a2d93e9e2a8c9d29302a075355c697fedcbe

Observation 72fec629-b593-4885-80d1-c63beb3e8e61 · inbound

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain cites this paper.

Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain Certifying LLM Safety against Adversarial Prompting

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:44:10.487753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:44:10.487753Z digest=sha256:6c502a51321080e3ec5285433c0e819fe4badabccd721b8224d7ccc3cadb15b2

Observation 1eebe795-ea34-4255-a580-d7a0b4b0d66b · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.473339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.473339Z digest=sha256:23ddf4171720ca0bd92631dfda14d11bf585b70f6d9c870c1f8a1db33586dfad

Observation 758bf927-a815-4e7b-a2c5-a5e6843da869 · inbound

AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents cites this paper.

AgentSentinel: An End-to-End and Real-Time Security Defense Framework for Computer-Use Agents Certifying LLM Safety against Adversarial Prompting

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T21:44:12.318888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:44:12.318888Z digest=sha256:bc4c80286891c915a1307ab3c37d71108c08fad3159182c6161843b74aea7a1d

Observation 039e0742-22f3-4a98-9c64-534d84230ae3 · inbound

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models cites this paper.

When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models Certifying LLM Safety against Adversarial Prompting

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:31:11.646688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T09:29:14.842228Z digest=sha256:65f0d8df029b4c1acc7a306b682794ed8a64a4a4933c206c5f7bc1e5e491ae60

Observation 510fa07b-1d80-43a3-9c68-e3914854dc9c · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Certifying LLM Safety against Adversarial Prompting

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:45.397309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:45.397309Z digest=sha256:04dbee18ff22e95febc23e377e82d05f1a8013cbdc476d5ee78cef1641fa6b4d

Observation 7068a9da-7e48-4fb7-af66-9367cc3fce90 · inbound

BEAVER: An Efficient Deterministic LLM Verifier cites this paper.

BEAVER: An Efficient Deterministic LLM Verifier Certifying LLM Safety against Adversarial Prompting

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.336475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T01:58:44.719715Z digest=sha256:3b8a94ef28e7697f84b3a0146e2d6f2ab986c488d59f44e106e795c9e0b144ec

Observation 742d023c-c22e-41d8-9422-6fb186a319dc · inbound

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents cites this paper.

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
verified exact
orphan_title_repair, observed 2026-05-13T17:16:32.546647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T17:16:17.464937Z digest=sha256:aae787a679d175314c3cce01a04d03d8c4064909155796bc6e9d292d4227c497

Observation 67eb77b1-719c-4bfc-958c-616813410f69 · inbound

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying cites this paper.

ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying Certifying LLM Safety against Adversarial Prompting

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:45:57.667592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:57:59.918483Z digest=sha256:e7c0b2a7b8bec8cbb3671798738a3f0d285b5de22b5f3aa974fdf52a671961f8

Observation f77ade4b-66b0-431f-b006-0d2153c9f251 · inbound

Towards Understanding the Robustness of Sparse Autoencoders cites this paper.

Towards Understanding the Robustness of Sparse Autoencoders Certifying LLM Safety against Adversarial Prompting

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:10.241956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T05:44:24.532548Z digest=sha256:420c5d44a6252597fe64fabf5f904d461689985f6ce69a0ac4f5c650027e8013

Observation f4f54fa7-3020-4fb1-9dbb-b50c0d9f1373 · inbound

Ethics Testing: Proactive Identification of Generative AI System Harms cites this paper.

Ethics Testing: Proactive Identification of Generative AI System Harms Certifying LLM Safety against Adversarial Prompting

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:56:08.057350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T20:49:22.147548Z digest=sha256:c4f76b6385264aeff2908c603d7803c9d54b97e6bb86aa5c9e4f70e00937fa95

Observation 8e02b80d-2e64-4ee1-a04f-f5a2a82b6032 · inbound

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing cites this paper.

Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing Certifying LLM Safety against Adversarial Prompting

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:25.686124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:09:43.879809Z digest=sha256:a4a96373ce39968d45318f9452542278c63a6d1a710927d167d5ac10e17946a1

Observation a1572414-af7d-48a5-8cd2-ba636ab52bfa · inbound

Green Shielding: A User-Centric Approach Towards Trustworthy AI cites this paper.

Green Shielding: A User-Centric Approach Towards Trustworthy AI Certifying LLM Safety against Adversarial Prompting

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:23.340402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T03:43:54.896449Z digest=sha256:cf9b9673775fc374d35a9853339ebd0c77e9fa3ffdf5f58005329439b1a17af8

Observation a48b12bb-7580-43b9-b7f4-649c91f08507 · inbound

Attention Is Where You Attack cites this paper.

Attention Is Where You Attack Certifying LLM Safety against Adversarial Prompting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:05.235919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T19:54:41.445447Z digest=sha256:9652698701a9b82b7c7c8a1d4e7ee7b479fecc3b40cc5eec8fe313318ba1058c

Observation 476d33b0-1915-4757-8903-a0c31063cdda · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Certifying LLM Safety against Adversarial Prompting

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:07.800202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T12:13:16.896486Z digest=sha256:3d9e94c8e7e18cc928621831d26ad86218fdce481f382cf8beb89d559f2a0461

Observation 4508b4c4-07f6-454d-b379-b70604ee8b4d · inbound

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation cites this paper.

How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation Certifying LLM Safety against Adversarial Prompting

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T14:49:14.882245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:49:14.882245Z digest=sha256:a745d33eb31b443c8c477d8ecfe21763e4a77c77f7de6f9752e7e23952e6be17

Observation ed216ea4-a6eb-4ff1-8647-fc9c1db40474 · inbound

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability cites this paper.

Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability Certifying LLM Safety against Adversarial Prompting

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:41:21.790576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:41:15.286881Z digest=sha256:d41f55c36126b6940b772702e0392a5ee7bcc485d8aaec0e04026a13c9893e8d

Observation 64ca20fe-1ed2-4694-9d5e-249a255c72dc · inbound

Re-Triggering Safeguards within LLMs for Jailbreak Detection cites this paper.

Re-Triggering Safeguards within LLMs for Jailbreak Detection Certifying LLM Safety against Adversarial Prompting

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:06:25.074198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T04:37:29.075413Z digest=sha256:d6573af90044bd3b8808d75c40080e85342cf27456eb89f11745b1faa7958a92

Observation 13365dc7-0a85-4bc8-9853-39a7b5d1634b · inbound

Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models cites this paper.

Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models Certifying LLM Safety against Adversarial Prompting

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:38:05.831828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T06:33:30.647965Z digest=sha256:313ed16ff120020fd76f2463a07f126ca1202f3d62369d86e2ae7068aa654b3a

Observation 57000007-7040-4aba-8176-8cb457f56708 · inbound

Adversarial Reframing: A Framework for Targeted Generation in Language Models cites this paper.

Adversarial Reframing: A Framework for Targeted Generation in Language Models Certifying LLM Safety against Adversarial Prompting

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:36:21.028156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T09:35:51.862736Z digest=sha256:f0c1257697583d0b77d4daaa6fe23024dbda4d4512d6eac0a20e83349eff8dcf

Observation bc19f71c-d222-44f1-8b1b-e02761f36024 · inbound

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents cites this paper.

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents Certifying LLM Safety against Adversarial Prompting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:16:36.016810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T09:15:57.044886Z digest=sha256:e4ceba65234c5419ecd501da5872940e55b43b9adef8f5f5c00d3a73402225aa

Observation feb1401b-66ba-44d2-8ea8-a30a62a14e04 · inbound

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks cites this paper.

SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks Certifying LLM Safety against Adversarial Prompting

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:26:59.323331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T01:16:07.252429Z digest=sha256:4072171b360ab53376518adecbca995024867adb6cd3b57178573924eaa4bbb4

Observation 7fcce9d9-0859-4eb1-94ae-65d4e74112df · inbound

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions cites this paper.

Mitigating Taint-Style Vulnerabilities in MCP Servers via Security-Aware Tool Descriptions Certifying LLM Safety against Adversarial Prompting

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-09T10:26:11.133047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-09T10:22:23.782469Z digest=sha256:29edfaec2946aa65c49b4b8079ab954a41b4f74698d67e05a886d78e2726e5a8

Observation 6ca8f49f-5d36-498f-aaf9-9beb9e4c8a40 · inbound

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents cites this paper.

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents Certifying LLM Safety against Adversarial Prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T01:34:11.280395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:34:11.280395Z digest=sha256:bd334406252af4da0429ce74e323ba7e8f97f4037ef14a62c09d8800db00179e

Observation 771bd9f5-e888-44a8-a10d-77ae5df37a46 · inbound

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks cites this paper.

Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Turn Attacks Certifying LLM Safety against Adversarial Prompting

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T01:16:13.014747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:16:13.014747Z digest=sha256:58c23c7b16fbc0448351246b2ad71390877fd71e9f01275db4c6b872fbfd7921