Pith. sign in

Paper Citation Record · LEDGER

Universal and Transferable Adversarial Attacks on Aligned Language Models

As of 6 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 100 inbound Pith citation observations for arXiv:2307.15043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.15043 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T07:42:09.112946Z

measured 129 of 129 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 723 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.774858Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact23
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

187
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ca4e928a-9d59-41b2-b973-a170c0b100b1 · outbound

This paper cites Generating Natural Language Adversarial Examples.

Universal and Transferable Adversarial Attacks on Aligned Language Models Generating Natural Language Adversarial Examples

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T07:44:08.458367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:80ee8b8ccbac0647b98ee3e7135088a229a5a18981da4f01f200592558308bdf

Observation b2f5e9bd-290b-470c-9202-5eb960eee6c9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Universal and Transferable Adversarial Attacks on Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.438344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:8cd43e3a1053cd2fa34f8a80b047d4c412d05cce7802ff8f7a4d5f447f438762

Observation 6a9e9bee-8dd3-4138-a99f-f2f458f35de7 · outbound

This paper cites Evasion attacks against machine learning at test time.

Universal and Transferable Adversarial Attacks on Aligned Language Models Evasion attacks against machine learning at test time

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:44:08.931997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:4dda221969363ea0da692eb0256463ce3aee185c05e9c6792822a8ea67622ae5

Observation 70b52c00-5d1e-4019-8c6b-dbede855f858 · outbound

This paper cites Are aligned neural networks adversarially aligned?.

Universal and Transferable Adversarial Attacks on Aligned Language Models Are aligned neural networks adversarially aligned?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.583092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:610be1d6ad1e110fe60da4d29e9b127f988161ff1fcb7eed064b75bf428e503a

Observation 4a777885-6e35-43da-8749-bb209fb2e591 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Universal and Transferable Adversarial Attacks on Aligned Language Models QLoRA: Efficient Finetuning of Quantized LLMs

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.576786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:4e0edd6889319e3d814bc6d565902f12e3d57c1b27b1b4d3909fef282fb44ed5

Observation 97d41f36-9e04-4305-898e-8d65624e5c6f · outbound

This paper cites HotFlip: White-Box Adversarial Examples for Text Classification.

Universal and Transferable Adversarial Attacks on Aligned Language Models HotFlip: White-Box Adversarial Examples for Text Classification

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.570834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:448515657fab6c17104870921c6f8248a5b58ca97fc802940b17886480aa1aa9

Observation d5dc8f3f-5cae-44a5-9567-4dfa8e5f5e6e · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

Universal and Transferable Adversarial Attacks on Aligned Language Models Improving alignment of dialogue agents via targeted human judgements

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.427372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:13010e4d1cfdd4cd7eb2cdc89bf39c66e5359c989f57dba53f1f40f8659399c4

Observation f7530c69-8ed4-4384-bf38-f51aeeed4585 · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Universal and Transferable Adversarial Attacks on Aligned Language Models Explaining and Harnessing Adversarial Examples

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.565407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:84b19cd0b80c212badd8da94e902f84e105673661f59a73b799475601c6fc514

Observation d85c7d76-6cfc-4435-86f8-85b94843a510 · outbound

This paper cites Gradient-based Adversarial Attacks against Text Transformers.

Universal and Transferable Adversarial Attacks on Aligned Language Models Gradient-based Adversarial Attacks against Text Transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.560184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:edf1c89c3e6b829f7a0ef19ef3b1d71bf5d06a405f3b8075ce34356d757a9862

Observation e297ad03-7c85-42a5-bc48-d7e4bd71d09e · outbound

This paper cites Adversarial Examples for Evaluating Reading Comprehension Systems.

Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial Examples for Evaluating Reading Comprehension Systems

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.553572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:9313de1fd1db73b545b65705214923340645271eaf5f16db1a5b1ad11e7541f4

Observation 2057dccc-700a-4cf2-9f2f-cc98e76fe724 · outbound

This paper cites Automatically Auditing Large Language Models via Discrete Optimization.

Universal and Transferable Adversarial Attacks on Aligned Language Models Automatically Auditing Large Language Models via Discrete Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.444210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:24102beca6d346752ebf7b93cdf97719c0cfa398688f523b0c3db99a1cf61577

Observation 321cb681-049f-486e-96dc-72cd871218ab · outbound

This paper cites Scalable agent alignment via reward modeling: a research direction.

Universal and Transferable Adversarial Attacks on Aligned Language Models Scalable agent alignment via reward modeling: a research direction

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.548111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:4cb58d6d6018a1a335b811fd5fd0bcb366d89fda1413d2ece57e15efc0666af5

Observation 0f77806a-b0a6-430c-a32d-59bb0a883152 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Universal and Transferable Adversarial Attacks on Aligned Language Models The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.432848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:8a4ed31d0317e2bec9783b209edbd51bec79a115d19bbd1aabb033afbc368ccd

Observation 1ecd6f6f-ef44-4f02-8e17-ff3599895285 · outbound

This paper cites Sok: Certified robustness for deep neural networks.

Universal and Transferable Adversarial Attacks on Aligned Language Models Sok: Certified robustness for deep neural networks

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T07:44:08.935695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:4e3e13943e7bb663ebd971f99a0cef925a97a13ebe3b47c67554bf6a8ccb1e9a

Observation b979029f-6eec-4452-a10f-2985c4c8fd37 · outbound

This paper cites Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models.

Universal and Transferable Adversarial Attacks on Aligned Language Models Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.543081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:49bdc776898ae59309fee3e71545069e4ffb07928d0ea53fee36d45be7288b12

Observation d76ba06f-d1c6-48a0-b92b-75a91a292791 · outbound

This paper cites Black Box Adversarial Prompting for Foundation Models.

Universal and Transferable Adversarial Attacks on Aligned Language Models Black Box Adversarial Prompting for Foundation Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.537226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:26e51c372657bb3b557351d6b1d24306198fe77d4356e61f8b8a203a83b638a0

Observation ac6388fc-d75b-4f5b-a80b-472482296aca · outbound

This paper cites Universal Adversarial Perturbations for Speech Recognition Systems.

Universal and Transferable Adversarial Attacks on Aligned Language Models Universal Adversarial Perturbations for Speech Recognition Systems

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.531019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:4e4a8373174c4d2bb455c5772d3d5e9b61e0f0981ff838cf8bcbe4d59eddffff

Observation d83fd42c-7ee9-4e4a-9ba9-b8cea034257a · outbound

This paper cites Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples.

Universal and Transferable Adversarial Attacks on Aligned Language Models Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.523743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:dea16647aa7ef727a0011998104c00443406add152412ec35ce5e1061374471b

Observation c0389a04-54c2-4631-9fdd-c8aa7ef0fd9c · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Universal and Transferable Adversarial Attacks on Aligned Language Models The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T07:44:08.464630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:d5d353ea61ca2fccbfa5abe5be78093f7fa917ff27c84bad0081f84786d90b83

Observation 1045f988-1058-450c-9c07-c2c184a9cd74 · outbound

This paper cites AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.

Universal and Transferable Adversarial Attacks on Aligned Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-24T07:44:08.517831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:a80ced927f5e433a55ad6b7e3cd61fcd75c2b23a982b0d1d577153c7108a42b9

Observation 3831fdf3-9c73-4317-93af-53f49c229acb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Universal and Transferable Adversarial Attacks on Aligned Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T07:44:08.512307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:a146a6525b6e60993c6ea20dd159431de827ee4803eabb66df5178f20047b1cb

Observation 023b5a70-651b-4c14-b264-82fadcd3dcd6 · outbound

This paper cites The Space of Transferable Adversarial Examples.

Universal and Transferable Adversarial Attacks on Aligned Language Models The Space of Transferable Adversarial Examples

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.506484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:632c83cba108d7d5edaf293c63d7d27d52690a9bd4a39ef8e38b511fa0b80084

Observation 51751e4c-a6b8-468f-8be4-42ff42147cec · outbound

This paper cites Universal Adversarial Triggers for Attacking and Analyzing NLP.

Universal and Transferable Adversarial Attacks on Aligned Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.501178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:b6e04a9e2c7eaa2288f146d3490b4720272f7b2b3ccecad135eada12b0e6ab51

Observation 299597fc-8011-4ba3-900f-bd62627e7964 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.495253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:d284fdf3e059a244820aaf6ecadb02e85bf7b2bb241c84b0da1caca876035c3e

Observation 3bd9b335-1537-4360-a4d6-67d4c015a45d · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Universal and Transferable Adversarial Attacks on Aligned Language Models Jailbroken: How Does LLM Safety Training Fail?

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.488843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:afe719c254bfbd2e3c7c60cffe33f028aeb0fa9608975b8a26c4cd3d468b6b94

Observation e62e4a33-f511-4137-b300-e7a0247455d2 · outbound

This paper cites Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery.

Universal and Transferable Adversarial Attacks on Aligned Language Models Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.477006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:ebb4c473de21f0445976f4c6f498d764dcc40f0403f58a7e33b58a1c083afd9d

Observation 3efae1f4-54d7-4687-aafa-7b082e8c0467 · outbound

This paper cites Fundamental Limitations of Alignment in Large Language Models.

Universal and Transferable Adversarial Attacks on Aligned Language Models Fundamental Limitations of Alignment in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.470910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:a940c7a02e0effc7e92a6969b1656f921c7a5fb1a041c8ac33750731e716b69e

Observation cebe60e5-147d-4845-8fa1-b8dd4a3d0d07 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Universal and Transferable Adversarial Attacks on Aligned Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:44:08.450809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:d2d4834f1e21139a3029a574197e4d176034fd2c53d517306a0b5d8621ac348e

Observation 2b2e6116-13c7-4016-8e76-f738ee64014d · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

Universal and Transferable Adversarial Attacks on Aligned Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T07:44:08.483210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:42:09.112946Z digest=sha256:1e9e7270be216b4f53908bbfbf2005d41878f6a2045e4b251a69383b1a51664e

Pith citing papers

Observation 689f133a-5324-4b83-8b9e-c4eea52b2441 · inbound

XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models cites this paper.

XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:51:50.807968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:51:50.771676Z digest=sha256:8da8b0006545e8a7ed8455e9e3781ce85c5115936e0508c1be584a5986b7aa51

Observation 7ec25416-37d4-4a6f-826e-2f9e0880554d · inbound

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models cites this paper.

"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 95

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T08:39:28.195931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T08:39:28.047394Z digest=sha256:9e74566e8867748033b49a8a65d36a99a097f9a7df99ebe2ba90215cfd562b41

Observation acaf8eee-4c15-4042-b03f-d453f71624a3 · inbound

Baseline Defenses for Adversarial Attacks Against Aligned Language Models cites this paper.

Baseline Defenses for Adversarial Attacks Against Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T23:24:39.984464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T23:24:39.835347Z digest=sha256:0c6a0e4e7de95b8e96a918e6dfad9ab628672882e5d373cb871ad835ff85df1b

Observation cfc55ff5-8613-4f87-b8a0-da5af497d12f · inbound

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts cites this paper.

GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:25:21.174508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T06:25:20.966510Z digest=sha256:ce7b77f2ef18de34ff8296a7e929bc45b035dc709e534c3a24add1d3b8c800c7

Observation 8bc99e31-5022-4f5d-8b1b-659d0ef8fd2e · inbound

Low-Resource Languages Jailbreak GPT-4 cites this paper.

Low-Resource Languages Jailbreak GPT-4 Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-17T09:24:14.046876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T09:24:13.911401Z digest=sha256:011c0e3d69bfd021ca9fc44ec95d68e6183dad38116397373202bfc8ca3c6e3d

Observation 9f633d5a-b017-40db-b951-b5cc550752bf · inbound

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks cites this paper.

SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-14T17:11:00.729569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T17:11:00.639293Z digest=sha256:be6a66939f265b53b447a2472f2ea0e44b949916624a9f5f8e8714e4551b1df2

Observation cb3c49d1-813d-41b1-94d9-da2e6c6bff3e · inbound

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation cites this paper.

Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:00:51.530893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T22:00:51.487120Z digest=sha256:5fd6ceaf5b22baa297b264739f92696047397e29e0e0fd7eb644add21f853174

Observation 9079ed20-f698-4c86-9b07-489d1b1d3578 · inbound

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting cites this paper.

Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T01:59:51.414513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T01:59:51.326493Z digest=sha256:ee17c8a6d454a1c85d4c348ad493b6ab3eb9324b6c33b4bf7e0eb4da8cdb3859

Observation 19f7b153-cf69-4d84-a46e-93245620482a · inbound

AI Alignment: A Comprehensive Survey cites this paper.

AI Alignment: A Comprehensive Survey Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T14:28:49.032977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T14:28:48.987140Z digest=sha256:883ed65e370cbdb071dffd03eac826091f0cc31bb35d7c68b333fca215c51e20

Observation 58a9aa42-db13-4505-a043-5befa3274470 · inbound

Scalable Extraction of Training Data from (Production) Language Models cites this paper.

Scalable Extraction of Training Data from (Production) Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:00:46.810082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T19:00:46.708242Z digest=sha256:ea45c09b5d1787a1c97ce2a2ef17eaf476567fc26e29abc64bfa207f59ae395e

Observation 8e84bd5f-8736-41af-9961-f9f71ab2b41a · inbound

TOFU: A Task of Fictitious Unlearning for LLMs cites this paper.

TOFU: A Task of Fictitious Unlearning for LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:07:39.246388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T11:07:39.215164Z digest=sha256:524a87d678fb71337f6b5f4d656a86c0c49b74c77fff800afaef9c07b3cbe6d9

Observation 072c6e2a-0719-4282-b7aa-32e4992998ac · inbound

Whispers in the Machine: Confidentiality in Agentic Systems cites this paper.

Whispers in the Machine: Confidentiality in Agentic Systems Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:03:53.912178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T03:59:03.972043Z digest=sha256:f8630ef9c2d34124a0c9e0751c9b88b802c4f73431f61462e8db741e64796c9c

Observation 7f8da37b-d73a-46ac-acc8-04760c41a286 · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:28:02.900691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:4bf441344e27df29c249c871ae05249d610f0208b7073d0173c133bae173ad4d

Observation 4df00ab4-ea88-47ed-a0bc-0df4844ba04d · inbound

Defending Against Indirect Prompt Injection Attacks With Spotlighting cites this paper.

Defending Against Indirect Prompt Injection Attacks With Spotlighting Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:28:55.503509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T22:28:55.370749Z digest=sha256:c85ed9aee241b7dddf511cb3070d3d418b6583b0925492cd900e37c4b01e9893

Observation 7149086c-43aa-441f-a801-7da49abfc2b1 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:08:05.564496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:f042f515e8aa762207f11c2665c5aab8ddb17b2c702a718ecfdf598a6ea2c4d6

Observation 8350874e-ad20-47ca-ac27-48037e9179b5 · inbound

Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack cites this paper.

Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:12:41.520873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T02:12:41.431156Z digest=sha256:4088df0ab8010b9d18e18e80f8fd06cf8e1c1adba57ade88d33e6ab92da348e6

Observation 6370596a-fdd9-4567-b76b-1ba428927731 · inbound

LLM Agents can Autonomously Exploit One-day Vulnerabilities cites this paper.

LLM Agents can Autonomously Exploit One-day Vulnerabilities Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:18:27.680738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T04:18:27.597704Z digest=sha256:d6c5d40f048b513e103f6e381eb206912e43c7f519404b298044cf176a626b3f

Observation 1b2eada1-26ce-4905-a0c5-23380fb12b90 · inbound

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions cites this paper.

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:59:30.811638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T10:59:30.728091Z digest=sha256:bbaba8ad1e56c08c14a6d7e3d6dee6d7a8b951692aa50823ae980ccf0accfb24

Observation f70dc3f4-852e-49b1-8029-357de2a01387 · inbound

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment cites this paper.

Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:38:39.800763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T00:38:36.992597Z digest=sha256:b971efeea41ada36492c9c3350cfead57ad459cdb59cb8d321ac86d34965deb8

Observation 8990ca77-8e46-4295-973c-6fa1cb724aab · inbound

Refusal in Language Models Is Mediated by a Single Direction cites this paper.

Refusal in Language Models Is Mediated by a Single Direction Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 208

Resolution
verified exact
local_arxiv, observed 2026-05-13T10:47:56.181588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-13T10:47:55.934081Z digest=sha256:f7e155b4431441d10342c84324c400cbf03fcca079d74bee3a435232fbfe16b0

Observation d9c4eff1-78e4-4d0c-af2c-c8aeac3843c4 · inbound

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents cites this paper.

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T06:35:13.539981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:35:13.331872Z digest=sha256:36b70a06bfb1c53dfda87fcd2ebc1b0438bcf13ccab48ca0277b57b4557a5e7a

Observation 78e3adc0-7411-4d4b-96d4-4c86ceeeba18 · inbound

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs cites this paper.

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-17T16:25:14.923012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T16:25:14.744887Z digest=sha256:018eedb5906e89d5f989c48d6219bb53895f4e0ff46516d2268d6b7d5389e986

Observation 9ba9306c-28a8-4616-8a2c-4fd7c64240d0 · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 125

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T02:20:44.537423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:dbd469e93c80fed63dd90881a8ba9ec8bc09fb2fe8ea512fb8ca466660e03e55

Observation 0e127a1a-06d7-4776-95b0-0087710317ac · inbound

Guidance for twisted particle filter: a continuous-time perspective cites this paper.

Guidance for twisted particle filter: a continuous-time perspective Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-23T21:08:26.084514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T21:06:52.470243Z digest=sha256:d0a2b61c1604e3aff11790e1d365eb951803403bf2d1021590d9f8ec2f20aa34

Observation df5667e5-d446-46e4-a88d-0a5a584f919b · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 138

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T12:04:10.681068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:810bba57a20452c9f068b504c39ec86bd8e801399fbbb19a1af3c87098aa6f4c

Observation 979b9f49-bfce-45ca-8377-e8ad995d6700 · inbound

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey cites this paper.

Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 184

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T20:58:25.850959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T20:58:16.237327Z digest=sha256:dbe8bada6fddd752507ada7f8b1ee09c0d992fb0afe52d9a9e35348c9a7a34da

Observation 5a53c659-fa35-4dea-8ff0-6632e7ba4b1e · inbound

Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems cites this paper.

Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 98

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T19:32:19.876030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T19:32:19.405615Z digest=sha256:06a0fbedb4bc3a416082534e3eb88fcc2249c26ba7f989f6b71f3ce29b109757

Observation 6b640dcf-4154-4191-95c9-4adfb30dd792 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T01:35:51.188032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:1c65909942687c90f60f25da5297c26d018ddcd7ad6dc12fb2ff34c44ef7b2f9

Observation 61bec995-990b-4a62-b57d-70cfcaa12970 · inbound

Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models cites this paper.

Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:03:21.394452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T19:01:43.922848Z digest=sha256:1dc83362d438c823ad409d26af704ea701e58eef1eb320ddb23fc7b5f2d69da9

Observation ce3bce45-f4b3-47b5-884f-7d52f467af29 · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:50:13.931900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:af7aa04f78edf3fc68e4646039c69532e604766c5f17e783c320eed86e308a9e

Observation 3438c2bd-479c-41be-aa06-dd13c4777f2a · inbound

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference cites this paper.

Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T17:46:47.014121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-20T17:46:46.845424Z digest=sha256:575ab4a0c1ae31c38a706e328d50cfa01f95404b9b78128cfe8f92b77f094ca6

Observation b7cc975b-4979-4e18-82c7-854a94d69f03 · inbound

Adversarial Hubness in Multi-Modal Retrieval cites this paper.

Adversarial Hubness in Multi-Modal Retrieval Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:42:39.776660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:39:36.039613Z digest=sha256:6a6b9f9a7ff5f8e74c24971e7952836f68b2322e6d2f1c190ac168729ac2d981

Observation ccaf8640-23cc-4605-9a3a-18bca850bbc6 · inbound

Improving LLM Unlearning Robustness via Random Perturbations cites this paper.

Improving LLM Unlearning Robustness via Random Perturbations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:42:33.408884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T04:41:47.910423Z digest=sha256:c9d2a618d20c0e5429de44bd943ae84c4a995cd46efad41f49b79ecd9bbfe8f1

Observation cbb29ad8-3885-4d94-aeda-1032d8eb6078 · inbound

Peering Behind the Shield: Guardrail Identification in Large Language Models cites this paper.

Peering Behind the Shield: Guardrail Identification in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T03:45:21.382565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T03:45:14.234545Z digest=sha256:0abbf04ccd6f4be566f9f092b994bf09736394ed2935e1e86465122b5d2f6231

Observation 637f526e-2929-40eb-97d3-90a660c36cc2 · inbound

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety cites this paper.

Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:42:33.773359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T04:39:04.591722Z digest=sha256:600050d7b2b562d10974f3e64875c7072bb6e292eec85d6d2e759753df19529f

Observation 894a831f-546e-4e02-93cb-64991a8715a0 · inbound

Responsible Federated LLMs via Safety Filtering and Constitutional AI cites this paper.

Responsible Federated LLMs via Safety Filtering and Constitutional AI Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.163571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T02:20:07.800524Z digest=sha256:813a019898827f8c7fdbe217c841f01b212a00a335edfeca3f07f520a488e71a

Observation af241a98-56db-4f1d-94c4-6c9679ab0378 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:02:44.377608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:e8fcbea7e701179bafbee7d7cbd17c367697a1f4ccf80f2cad9f79b5d54a0169

Observation bc488420-641c-4c77-88b8-04b63f9be930 · inbound

Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction cites this paper.

Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:11:58.054185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T19:10:55.009810Z digest=sha256:78c15259ac5a4a300fd4d2c45bb4e8ed107861aa417cc1d6c778043e9f5159bd

Observation 24a4adc1-99d1-4f51-8574-5635935e58ac · inbound

Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs cites this paper.

Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:41:41.452731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:40:58.506345Z digest=sha256:4b06c1c08b2325f42a9faeefe64f57747d1201f7c228a2abc58448f15a17c439

Observation f02a76a8-8b5c-4f40-89c5-fab83ff4e166 · inbound

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling cites this paper.

Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
malformed identifier
local_arxiv, observed 2026-05-22T14:21:39.607873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T14:21:18.355360Z digest=sha256:368730a4ccf92393c716a05523965c153ae5e739aafb0f918c248aef52c8bbfc

Observation b76efd75-d0d9-46ea-ad2a-d362a8bce8f4 · inbound

Secure LLM Fine-Tuning via Safety-Aware Probing cites this paper.

Secure LLM Fine-Tuning via Safety-Aware Probing Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:11:35.797307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T13:07:09.402763Z digest=sha256:290c88b3950eb8ae7600e418ae7b0a3fabd820e7b61fa3f9e9b9e716eddef8b8

Observation b7c82321-6650-484a-962c-968250b0b367 · inbound

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs cites this paper.

Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 46

Resolution
malformed identifier
local_arxiv, observed 2026-05-22T13:31:36.614411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T13:26:47.663124Z digest=sha256:c92632d19a516a04499ba2b50b90455e10591b1727f5f85c3f72d84566154e3d

Observation 8f9b1d36-ed45-4d04-bae2-971bb6b06b23 · inbound

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction cites this paper.

ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:37:16.011602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:34:09.428653Z digest=sha256:a89dee39f3a7575638a5042d344a949e3ef04087d0114b67254f13e609f38184

Observation 9376c1d2-219c-4e2e-b956-503d17ea8e7d · inbound

Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG cites this paper.

Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:34.183042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T08:10:31.473565Z digest=sha256:a319833a8e9cd274095299b51b7720bcdea74c603bef2fe14f854d67c554302d

Observation 511005a8-5dd7-4be7-8719-e9692af9ad39 · inbound

Benchmarking Misuse Mitigation Against Covert Adversaries cites this paper.

Benchmarking Misuse Mitigation Against Covert Adversaries Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T10:32:14.672975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T10:29:05.104520Z digest=sha256:7e69c41be53dab921b17dad4dc050cd8ae573f043c9e99e24dce4361f1df15ed

Observation a5c26d1f-29f0-4915-ae9d-f3845d336592 · inbound

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations cites this paper.

Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:57:16.516818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T10:54:55.262180Z digest=sha256:e6a765ae560cefa165ded23cca9607599698978dfbd349b5747517cf66eb9556

Observation 80df81ac-fe22-4c21-b861-2c9cf89ea567 · inbound

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems cites this paper.

Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T09:22:15.935221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:20:12.827871Z digest=sha256:9e4c23a038fe88e9d1a4a321019eccc68a089bac9f5e2bf960687fc458f82ecc

Observation a1907dc4-5f4e-4c17-ae41-e158b954d2f3 · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:42:13.917686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:640dc603997b6db5c4fd9520dbe2504aac87199ce2cb9e7851247954778f0436

Observation da4f7547-1716-48ad-a6b2-97c30866c39b · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T08:42:12.809508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:e6547e6937bb71dfcd822ae54711fb5e17f31e3d7302d0b1a238db2540696bae

Observation a8a01a32-cf18-4ac8-aa00-4f59739f1c6d · inbound

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI cites this paper.

Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-22T12:21:31.168880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T12:17:59.458633Z digest=sha256:8453b269c88c819b66ef714e0f0100b8c8781e8113a69d819cb2863f468371ee

Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · inbound

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations cites this paper.

CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:47.094379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:18:47.094379Z digest=sha256:ab3699b499ca8ca7e336798eea94edfddd08f87b7201fbe380ab8dcf773a1cdf

Observation 509b5c76-8874-43d0-a16c-ce5ef743af98 · inbound

Data Compressibility Quantifies LLM Memorization cites this paper.

Data Compressibility Quantifies LLM Memorization Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-19T06:02:07.677164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T06:00:12.027708Z digest=sha256:42a73879ee1212a14338a61b4b27b3beb1df17642489d67f76527de434957fe9

Observation 05f8fcb2-eb3f-441e-99f1-8a526aa9d2fe · inbound

The bitter lesson of misuse detection cites this paper.

The bitter lesson of misuse detection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.283148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.283148Z digest=sha256:5cb65ec297b262eb7e748d68b2c6f6adf1184830b9b70247152c2c4ac13dfbc5

Observation 81543be5-060c-4908-a8f5-81c9cc76f904 · inbound

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms cites this paper.

Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:12:22.950253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:12:22.950253Z digest=sha256:87f632fa40d4a952481882a6c8ebd7bb98c54a4ef6a4f3a75054c3f021ce6f2c

Observation f682a274-4437-454e-9094-6f3cd42927a9 · inbound

A Mathematical Theory of Discursive Networks cites this paper.

A Mathematical Theory of Discursive Networks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:18.041331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:08:18.041331Z digest=sha256:fc1f4761462312b9bd93039fddd1667fd23b18672c60238470172b957eb862b2

Observation e3038363-73c7-4391-9fea-9f2d75a9285f · inbound

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection cites this paper.

Mitigating Watermark Forgery in Generative Models via Randomized Key Selection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T05:27:05.474981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T05:25:41.898297Z digest=sha256:0ffa80e705c4adf7f5f1558edfd8914a272c07f65fcf65bdd104b9136e9a3aad

Observation aad169b6-04a0-46ca-8be5-38d353ad6935 · inbound

Defending Against Prompt Injection With a Few DefensiveTokens cites this paper.

Defending Against Prompt Injection With a Few DefensiveTokens Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:32:44.285307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:32:44.285307Z digest=sha256:cc1b9ac38f2b75529e51bd578b7b8a8212bf529a9ee581147409887e294c3070

Observation a23e8d93-7f7b-4d19-8c45-61b2b66b355d · inbound

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation cites this paper.

Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T19:26:13.774858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:26:13.774858Z digest=sha256:cf84b4fc7e86282e4935a9725f72ac31a3de7a454dd446ba8586b8d64dfff77b

Observation 55d16c82-124c-46f0-820c-40016709b7f5 · inbound

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems cites this paper.

SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:46.576684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:28:46.576684Z digest=sha256:ea0df2eff5c0f600af2e75707a2b6904c94a4bbba73f0c1009a789aad882014e

Observation dda1b033-61d2-4105-9240-05780f8dee01 · inbound

Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks cites this paper.

Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:58.756338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:58.756338Z digest=sha256:052074e762addb3fbc5b9ab261273c034da8b94db0cbc408bcfc0447b36d4786

Observation 4777db18-432d-4d4b-b08a-62bbc6266ced · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.669576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.669576Z digest=sha256:84b21b06b41932dfdcf95517e847c7b872624da74dda4622fccdd80f6d60f41e

Observation 182fc9f9-0577-4c67-9158-41413fef0fbf · inbound

Scaling laws for activation steering with Llama 2 models and refusal mechanisms cites this paper.

Scaling laws for activation steering with Llama 2 models and refusal mechanisms Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:04.749036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:04.749036Z digest=sha256:acae142f5b947010b91dd3c2f483bf238c92466ca577dfcff282efe47b57c6a0

Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · inbound

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation cites this paper.

Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:24.884191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:09:24.884191Z digest=sha256:18c7e1a26c1c9ebd044b591e9b74c432b41c6aa87114724a424b5804e6068ab4

Observation c71ef73d-e629-4829-bd77-79b456a154dc · inbound

DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning cites this paper.

DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:30.788941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:30.788941Z digest=sha256:ca5092d6a64c468c664192de62f5c908130ed9b01abe3cc0a9dc7d5c68395cec

Observation 6b688f2f-1418-4268-8b2c-054e0d455c8b · inbound

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers cites this paper.

Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:28:54.067424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:28:54.067424Z digest=sha256:e73a23ad0cc3bb5324d85c1f33ce3c1ebb4d04d65637b227cd747b8a7b5b568e

Observation 0b34dbed-1ad6-4c8e-a379-7265076398a4 · inbound

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques cites this paper.

Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 165

Resolution
unresolved
no resolver link, observed 2026-08-06T16:24:30.380334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:24:30.380334Z digest=sha256:9b901202c39beb182451b818b78f30f6e12c4c6d4d3aebb6bf3e123341f882bb

Observation fa31761e-3949-45f7-ae22-ae12307ef775 · inbound

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models cites this paper.

Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:25:54.883404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:25:54.883404Z digest=sha256:83b488777cb6a40925156cb75940676964c22d565b752fe1acc9311e3f80bece

Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.626472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.626472Z digest=sha256:a230168ef072816241fd5279a6bdb345f4bf778a70e09f968b39d81b2efaecd3

Observation 72567c5b-b766-408c-84bc-fbb1cad14ede · inbound

Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree cites this paper.

Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:52:56.837539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:52:56.837539Z digest=sha256:2fcd8bdbaf841629b340aea9e569edc343aad413d5297795a49bbd97db5cd9ab

Observation 05a11772-9b8d-4347-87ad-2ea278f2011e · inbound

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning cites this paper.

AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:01.441397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:48:01.441397Z digest=sha256:88c3ae5b10112fbcf0488a8f07635dc68fb86139ee945072c052a6e577d21c76

Observation e5fd10a6-b584-415e-a924-c1a1834b7ae5 · inbound

PromptArmor: Simple yet Effective Prompt Injection Defenses cites this paper.

PromptArmor: Simple yet Effective Prompt Injection Defenses Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:01.109613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:42:01.109613Z digest=sha256:dc51efbc5ec53fcc1c281d6a7f1ba2aff5d7615993fa8911a13c2be91d94dcb1

Observation 46896952-598c-46e0-835f-975c032324c5 · inbound

Scaling Decentralized Learning with FLock cites this paper.

Scaling Decentralized Learning with FLock Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:19.586657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:19.586657Z digest=sha256:cb1a2f0b5180b5824970dfb2dd3b00e14087b2d1db053809b9c348d6ccbeb2ed

Observation 4109fe61-1555-4a87-afec-b2619bbd29d4 · inbound

Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems cites this paper.

Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:32:52.962454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:32:52.962454Z digest=sha256:9c2d778199cb71f46a9cb44a83eb54c127e968d4df12a62311124704f3131a15

Observation 2fc7540c-8fa6-4f35-9582-0abde34b4e65 · inbound

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs cites this paper.

When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:29.257814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:29.257814Z digest=sha256:5b7f428f518aa8fc1c8093a2d322b477d206f941e0cf2f3d50ec069680fa0cec

Observation 0d7e05da-9a35-4c94-9b5d-ee9effd20baf · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.710356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.710356Z digest=sha256:60e1129b0243040b44e2881eb4d2d25569d2e877cbc1bfd18d09f8da37e18cf5

Observation 0ef33d82-3afb-4942-8a18-8c3bbd3e0cd9 · inbound

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models cites this paper.

An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.952177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.952177Z digest=sha256:1456118536720b0483e675f2d19383520abef953d751dd8b68a99206773d1f16

Observation 67231121-8a7e-4839-b5ad-4050d547ecd5 · inbound

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models cites this paper.

From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:20.584079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:20.584079Z digest=sha256:edb08020f4157bf74ce2b5a90f8cb4f30d9e2e414c9c18f3c026f869837b8049

Observation 44dd8d47-a542-4d88-b2d7-e04cd9399c11 · inbound

Understanding the Supply Chain and Risks of Large Language Model Applications cites this paper.

Understanding the Supply Chain and Risks of Large Language Model Applications Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.918332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.918332Z digest=sha256:9a9935b7406e9f71257a96f2942c5f2fb0772846d489d4389c8e3fc1e1b1ea5c

Observation 39d40f62-47b1-4f12-a3ce-39c1cd4c782c · inbound

Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection cites this paper.

Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:08.218547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:42:08.218547Z digest=sha256:7bd43265d09730bb7dd12c3ead5dde6cb3aa4a2ce790f6f5e01c389109b657b3

Observation 258945a3-61a4-4d26-ad7a-a680973cd1d1 · inbound

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? cites this paper.

MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:34.222274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:34.222274Z digest=sha256:ab08ad57bfb407e547f224af4c737a57287fcf02ba5b2dbc5212586c1bf3aa1d

Observation 7824e3b7-baaa-4e5c-b1d7-1e519ffd28b0 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.203975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.203975Z digest=sha256:746307ca2569e041f95a1c1cf503e81b1091496732d1418f8201b0aff56a9fd0

Observation 0689a87f-add5-4750-9a33-7b636a5c79eb · inbound

A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction cites this paper.

A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 271

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:40.584133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:40.584133Z digest=sha256:8d23f428a9a42a009ffcd64cf0541e15ee9d8f950c0eb85ad805bf21bf89ded8

Observation 2c3fa86f-39a3-4e81-aef5-217f6cfced5e · inbound

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities cites this paper.

Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:10:29.544536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:10:29.544536Z digest=sha256:0c57b37ec25bf152b5a43e94f248db5f4ad3ef5c5f3b1dfd751e92b6c6ccf011

Observation ee54552d-dc47-4001-9de5-3984dec23f87 · inbound

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law cites this paper.

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:04:21.377383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:04:21.377383Z digest=sha256:fc978ba19e998f51c4435dd2622a25cd403e8176f1d99a99f10da11f812eaeba

Observation 9b19ed04-201d-49af-ac45-199cae965ab4 · inbound

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking cites this paper.

PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T03:37:01.111600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T03:36:24.013477Z digest=sha256:da5552ace710156794300a37a25c9fcbdfa18db9e99038995507959aeff8605d

Observation c4fdc14f-445e-4781-9f32-f9361afa7df3 · inbound

Training language models to be warm and empathetic makes them less reliable and more sycophantic cites this paper.

Training language models to be warm and empathetic makes them less reliable and more sycophantic Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:37.932831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:37.932831Z digest=sha256:de5a18e94c4e0344dac741e28f1b47217510daa37de7e33f92bbfc0e9080aa8f

Observation 6ab8a4e0-999b-4b93-a1ca-820f02528e21 · inbound

Strategic Deflection: Defending LLMs from Logit Manipulation cites this paper.

Strategic Deflection: Defending LLMs from Logit Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T12:05:44.471626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:05:44.471626Z digest=sha256:c4a5c9794ede313c0f8bbebe4879edbc21ad64ec4bd5b156227338cfce6a33f6

Observation c7e21a6a-5c9e-4963-b3b2-45149780c19b · inbound

ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction cites this paper.

ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T10:13:24.003904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:13:24.003904Z digest=sha256:397cdda5c696f9396c85b9bd006ddcc22b8b9ccd347977f041daddafa6297026

Observation 63eef6ee-9127-41be-a63b-b15d1a0cda18 · inbound

Adaptive Content Restriction for Large Language Models via Suffix Optimization cites this paper.

Adaptive Content Restriction for Large Language Models via Suffix Optimization Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:58.417889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:58.417889Z digest=sha256:3899ccf1eba4a04561cca4ba26c116d72e7fa409fe6b0454b525733cd70462d8

Observation d008524b-2600-43f7-aabf-d8db455af81e · inbound

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles cites this paper.

PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T05:43:40.073275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:43:40.073275Z digest=sha256:8236ee5607b2f2dde1c9c4cdddd509daa71f95743b51dca97890cf07b1ea1ff1

Observation 554aa273-f8f3-4e03-88a6-cfc38ee93423 · inbound

A comprehensive taxonomy of hallucinations in Large Language Models cites this paper.

A comprehensive taxonomy of hallucinations in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 108

Resolution
unresolved
no resolver link, observed 2026-08-06T05:29:20.984433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:29:20.984433Z digest=sha256:c59be7917b8f69e9935654ee3b673cb165bf2e4e0789d3e7a20fca571d19774f

Observation 04d16a38-0f21-415b-a5ff-0d1c0fc63bcc · inbound

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning cites this paper.

Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:47:22.349778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:47:22.349778Z digest=sha256:15fed2b4312a012ac1eb63c31996ac51a30f23484565bad6ea9aafde91db8bd2

Observation bc60518b-deb8-4fc3-88a7-a602e4919a1c · inbound

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments cites this paper.

ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T01:02:54.791365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T01:02:07.088724Z digest=sha256:83295e17c80637de8381d19851fc96130efd2066143546dfb2bde24dda689e53

Observation fe9819de-170b-47b6-b817-f4b5472bb3f3 · inbound

Automatic LLM Red Teaming cites this paper.

Automatic LLM Red Teaming Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:04:25.509141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:04:25.509141Z digest=sha256:549aebbb2dbaad0f7cd5dd8c4284423bb57dd7bfa60a8511f2b7a49f5b7405f5

Observation c4660857-0f0c-4709-959c-2603ffaed132 · inbound

Quantifying Conversation Drift in MCP via Latent Polytope cites this paper.

Quantifying Conversation Drift in MCP via Latent Polytope Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:51:28.192282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:51:28.192282Z digest=sha256:c2f78ca911aad5dda7fe44a467b111c82d4cd4aca5b43f54e8348dab0aeba9de

Observation f9154070-1a34-4e73-923e-6ee24ea89336 · inbound

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection cites this paper.

A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:24:23.367757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:24:23.367757Z digest=sha256:0893127d01be1e5ec8d1cd65732eb43f53804e50ae478848b33c5d3064599d14

Observation e05cc146-80ba-4697-b710-292789aa3a39 · inbound

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training cites this paper.

From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:34.616461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:34.616461Z digest=sha256:973d8f5b92b4da8e5f6c471ae5650683015932d901dab83fd76ef79015beb963

Observation a38a564c-74ba-4d95-803e-320d6bbb4945 · inbound

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs cites this paper.

NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:08:06.140982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:08:06.140982Z digest=sha256:08dda01550c923db69ded84f614a9d619aeee932f50bf2a5ae8c0d015f98e0c5

Observation 9ae6274f-493e-4d41-9408-b6a083ff8ea8 · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.913269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.913269Z digest=sha256:0db8c7f3ea89ce16267e75a5286fdcbf5f8b15103904901dfb5200b2126d9803

Observation 6a2bac1f-2a01-4ccd-bdec-c3257360aa44 · inbound

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal cites this paper.

ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T23:31:54.549906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T23:27:52.438709Z digest=sha256:9d66a93bbf1ed19ab5fa33bcdcf81e162d0c7e68424a74d9979b2e128f35a6f5