Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T07:42:09.112946Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 100 inbound Pith citation observations for arXiv:2307.15043.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-24T07:42:09.112946Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:26:13.774858Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
29 of 29 outbound references displayed
External citation measurements
187
pith, observed 2026-08-05T02:28:24.338817Z
Observation ca4e928a-9d59-41b2-b973-a170c0b100b1 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Generating Natural Language Adversarial Examples
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b2f5e9bd-290b-470c-9202-5eb960eee6c9 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a9e9bee-8dd3-4138-a99f-f2f458f35de7 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Evasion attacks against machine learning at test time
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 70b52c00-5d1e-4019-8c6b-dbede855f858 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Are aligned neural networks adversarially aligned?
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a777885-6e35-43da-8749-bb209fb2e591 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models QLoRA: Efficient Finetuning of Quantized LLMs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97d41f36-9e04-4305-898e-8d65624e5c6f · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models HotFlip: White-Box Adversarial Examples for Text Classification
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d5dc8f3f-5cae-44a5-9567-4dfa8e5f5e6e · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Improving alignment of dialogue agents via targeted human judgements
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f7530c69-8ed4-4384-bf38-f51aeeed4585 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Explaining and Harnessing Adversarial Examples
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d85c7d76-6cfc-4435-86f8-85b94843a510 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Gradient-based Adversarial Attacks against Text Transformers
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e297ad03-7c85-42a5-bc48-d7e4bd71d09e · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial Examples for Evaluating Reading Comprehension Systems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2057dccc-700a-4cf2-9f2f-cc98e76fe724 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Automatically Auditing Large Language Models via Discrete Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 321cb681-049f-486e-96dc-72cd871218ab · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Scalable agent alignment via reward modeling: a research direction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f77806a-b0a6-430c-a32d-59bb0a883152 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models The Power of Scale for Parameter-Efficient Prompt Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1ecd6f6f-ef44-4f02-8e17-ff3599895285 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Sok: Certified robustness for deep neural networks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b979029f-6eec-4452-a10f-2985c4c8fd37 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Exploring Targeted Universal Adversarial Perturbations to End-to-end ASR Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d76ba06f-d1c6-48a0-b92b-75a91a292791 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Black Box Adversarial Prompting for Foundation Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac6388fc-d75b-4f5b-a80b-472482296aca · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Universal Adversarial Perturbations for Speech Recognition Systems
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d83fd42c-7ee9-4e4a-9ba9-b8cea034257a · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c0389a04-54c2-4631-9fdd-c8aa7ef0fd9c · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1045f988-1058-450c-9c07-c2c184a9cd74 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3831fdf3-9c73-4317-93af-53f49c229acb · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 023b5a70-651b-4c14-b264-82fadcd3dcd6 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models The Space of Transferable Adversarial Examples
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 51751e4c-a6b8-468f-8be4-42ff42147cec · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Universal Adversarial Triggers for Attacking and Analyzing NLP
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 299597fc-8011-4ba3-900f-bd62627e7964 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3bd9b335-1537-4360-a4d6-67d4c015a45d · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Jailbroken: How Does LLM Safety Training Fail?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e62e4a33-f511-4137-b300-e7a0247455d2 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and Discovery
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3efae1f4-54d7-4687-aafa-7b082e8c0467 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Fundamental Limitations of Alignment in Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cebe60e5-147d-4845-8fa1-b8dd4a3d0d07 · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b2e6116-13c7-4016-8e76-f738ee64014d · outbound
Universal and Transferable Adversarial Attacks on Aligned Language Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 689f133a-5324-4b83-8b9e-c4eea52b2441 · inbound
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ec25416-37d4-4a6f-826e-2f9e0880554d · inbound
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation acaf8eee-4c15-4042-b03f-d453f71624a3 · inbound
Baseline Defenses for Adversarial Attacks Against Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cfc55ff5-8613-4f87-b8a0-da5af497d12f · inbound
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8bc99e31-5022-4f5d-8b1b-659d0ef8fd2e · inbound
Low-Resource Languages Jailbreak GPT-4 Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9f633d5a-b017-40db-b951-b5cc550752bf · inbound
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb3c49d1-813d-41b1-94d9-da2e6c6bff3e · inbound
Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9079ed20-f698-4c86-9b07-489d1b1d3578 · inbound
Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19f7b153-cf69-4d84-a46e-93245620482a · inbound
AI Alignment: A Comprehensive Survey Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 58a9aa42-db13-4505-a043-5befa3274470 · inbound
Scalable Extraction of Training Data from (Production) Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8e84bd5f-8736-41af-9961-f9f71ab2b41a · inbound
TOFU: A Task of Fictitious Unlearning for LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 072c6e2a-0719-4282-b7aa-32e4992998ac · inbound
Whispers in the Machine: Confidentiality in Agentic Systems Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7f8da37b-d73a-46ac-acc8-04760c41a286 · inbound
A StrongREJECT for Empty Jailbreaks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4df00ab4-ea88-47ed-a0bc-0df4844ba04d · inbound
Defending Against Indirect Prompt Injection Attacks With Spotlighting Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7149086c-43aa-441f-a801-7da49abfc2b1 · inbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8350874e-ad20-47ca-ac27-48037e9179b5 · inbound
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6370596a-fdd9-4567-b76b-1ba428927731 · inbound
LLM Agents can Autonomously Exploit One-day Vulnerabilities Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1b2eada1-26ce-4905-a0c5-23380fb12b90 · inbound
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f70dc3f4-852e-49b1-8029-357de2a01387 · inbound
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8990ca77-8e46-4295-973c-6fa1cb724aab · inbound
Refusal in Language Models Is Mediated by a Single Direction Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 208
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9c4eff1-78e4-4d0c-af2c-c8aeac3843c4 · inbound
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 78e3adc0-7411-4d4b-96d4-4c86ceeeba18 · inbound
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9ba9306c-28a8-4616-8a2c-4fd7c64240d0 · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0e127a1a-06d7-4776-95b0-0087710317ac · inbound
Guidance for twisted particle filter: a continuous-time perspective Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df5667e5-d446-46e4-a88d-0a5a584f919b · inbound
Training Language Models to Self-Correct via Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 979b9f49-bfce-45ca-8377-e8ad995d6700 · inbound
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a53c659-fa35-4dea-8ff0-6632e7ba4b1e · inbound
Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6b640dcf-4154-4191-95c9-4adfb30dd792 · inbound
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61bec995-990b-4a62-b57d-70cfcaa12970 · inbound
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ce3bce45-f4b3-47b5-884f-7d52f467af29 · inbound
VoiceBench: Benchmarking LLM-Based Voice Assistants Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3438c2bd-479c-41be-aa06-dd13c4777f2a · inbound
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7cc975b-4979-4e18-82c7-854a94d69f03 · inbound
Adversarial Hubness in Multi-Modal Retrieval Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ccaf8640-23cc-4605-9a3a-18bca850bbc6 · inbound
Improving LLM Unlearning Robustness via Random Perturbations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbb29ad8-3885-4d94-aeda-1032d8eb6078 · inbound
Peering Behind the Shield: Guardrail Identification in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 637f526e-2929-40eb-97d3-90a660c36cc2 · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 894a831f-546e-4e02-93cb-64991a8715a0 · inbound
Responsible Federated LLMs via Safety Filtering and Constitutional AI Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation af241a98-56db-4f1d-94c4-6c9679ab0378 · inbound
Towards an AI co-scientist Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc488420-641c-4c77-88b8-04b63f9be930 · inbound
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 24a4adc1-99d1-4f51-8574-5635935e58ac · inbound
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f02a76a8-8b5c-4f40-89c5-fab83ff4e166 · inbound
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Output Prefilling Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b76efd75-d0d9-46ea-ad2a-d362a8bce8f4 · inbound
Secure LLM Fine-Tuning via Safety-Aware Probing Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7c82321-6650-484a-962c-968250b0b367 · inbound
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f9b1d36-ed45-4d04-bae2-971bb6b06b23 · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9376c1d2-219c-4e2e-b956-503d17ea8e7d · inbound
Through the Stealth Lens: Attention-Aware Defenses Against Poisoning in RAG Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 511005a8-5dd7-4be7-8719-e9692af9ad39 · inbound
Benchmarking Misuse Mitigation Against Covert Adversaries Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a5c26d1f-29f0-4915-ae9d-f3845d336592 · inbound
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 80df81ac-fe22-4c21-b861-2c9cf89ea567 · inbound
Textual Bayes: Quantifying Prompt Uncertainty in LLM-Based Systems Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a1907dc4-5f4e-4c17-ae41-e158b954d2f3 · inbound
Exploring the Secondary Risks of Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation da4f7547-1716-48ad-a6b2-97c30866c39b · inbound
Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a8a01a32-cf18-4ac8-aa00-4f59739f1c6d · inbound
Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 822a0e55-2997-4ae1-8622-df97a4992f59 · inbound
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 509b5c76-8874-43d0-a16c-ce5ef743af98 · inbound
Data Compressibility Quantifies LLM Memorization Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05f8fcb2-eb3f-441e-99f1-8a526aa9d2fe · inbound
The bitter lesson of misuse detection Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81543be5-060c-4908-a8f5-81c9cc76f904 · inbound
Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f682a274-4437-454e-9094-6f3cd42927a9 · inbound
A Mathematical Theory of Discursive Networks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3038363-73c7-4391-9fea-9f2d75a9285f · inbound
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aad169b6-04a0-46ca-8be5-38d353ad6935 · inbound
Defending Against Prompt Injection With a Few DefensiveTokens Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23e8d93-7f7b-4d19-8c45-61b2b66b355d · inbound
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55d16c82-124c-46f0-820c-40016709b7f5 · inbound
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dda1b033-61d2-4105-9240-05780f8dee01 · inbound
Explicit Vulnerability Generation with LLMs: An Investigation Beyond Adversarial Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4777db18-432d-4d4b-b08a-62bbc6266ced · inbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 182fc9f9-0577-4c67-9158-41413fef0fbf · inbound
Scaling laws for activation steering with Llama 2 models and refusal mechanisms Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1612cb63-f7b3-4493-b43e-c44ffd16bb53 · inbound
Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c71ef73d-e629-4829-bd77-79b456a154dc · inbound
DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b688f2f-1418-4268-8b2c-054e0d455c8b · inbound
Paper Summary Attack: Jailbreaking LLMs through LLM Safety Papers Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b34dbed-1ad6-4c8e-a379-7265076398a4 · inbound
Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 165
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa31761e-3949-45f7-ae22-ae12307ef775 · inbound
Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba77999-6c87-41b9-bf16-cc2d7c3a81ad · inbound
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72567c5b-b766-408c-84bc-fbb1cad14ede · inbound
Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05a11772-9b8d-4347-87ad-2ea278f2011e · inbound
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5fd10a6-b584-415e-a924-c1a1834b7ae5 · inbound
PromptArmor: Simple yet Effective Prompt Injection Defenses Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46896952-598c-46e0-835f-975c032324c5 · inbound
Scaling Decentralized Learning with FLock Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4109fe61-1555-4a87-afec-b2619bbd29d4 · inbound
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc7540c-8fa6-4f35-9582-0abde34b4e65 · inbound
When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d7e05da-9a35-4c94-9b5d-ee9effd20baf · inbound
Agent Identity Evals: Measuring Agentic Identity Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef33d82-3afb-4942-8a18-8c3bbd3e0cd9 · inbound
An Uncertainty-Driven Adaptive Self-Alignment Framework for Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67231121-8a7e-4839-b5ad-4050d547ecd5 · inbound
From Seed to Harvest: Augmenting Human Creativity with AI for Red-teaming Text-to-Image Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44dd8d47-a542-4d88-b2d7-e04cd9399c11 · inbound
Understanding the Supply Chain and Risks of Large Language Model Applications Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d40f62-47b1-4f12-a3ce-39c1cd4c782c · inbound
Safeguarding RAG Pipelines with GMTP: A Gradient-based Masked Token Probability Method for Poisoned Document Detection Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258945a3-61a4-4d26-ad7a-a680973cd1d1 · inbound
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7824e3b7-baaa-4e5c-b1d7-1e519ffd28b0 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0689a87f-add5-4750-9a33-7b636a5c79eb · inbound
A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 271
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c3fa86f-39a3-4e81-aef5-217f6cfced5e · inbound
Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee54552d-dc47-4001-9de5-3984dec23f87 · inbound
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b19ed04-201d-49af-ac45-199cae965ab4 · inbound
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4fdc14f-445e-4781-9f32-f9361afa7df3 · inbound
Training language models to be warm and empathetic makes them less reliable and more sycophantic Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ab8a4e0-999b-4b93-a1ca-820f02528e21 · inbound
Strategic Deflection: Defending LLMs from Logit Manipulation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7e21a6a-5c9e-4963-b3b2-45149780c19b · inbound
ProbGuard: Proactive Runtime Monitoring for LLM Agent Safety via Probabilistic Prediction Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63eef6ee-9127-41be-a63b-b15d1a0cda18 · inbound
Adaptive Content Restriction for Large Language Models via Suffix Optimization Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d008524b-2600-43f7-aabf-d8db455af81e · inbound
PUZZLED: Jailbreaking LLMs through Word-Based Puzzles Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554aa273-f8f3-4e03-88a6-cfc38ee93423 · inbound
A comprehensive taxonomy of hallucinations in Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d16a38-0f21-415b-a5ff-0d1c0fc63bcc · inbound
Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc60518b-deb8-4fc3-88a7-a602e4919a1c · inbound
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe9819de-170b-47b6-b817-f4b5472bb3f3 · inbound
Automatic LLM Red Teaming Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4660857-0f0c-4709-959c-2603ffaed132 · inbound
Quantifying Conversation Drift in MCP via Latent Polytope Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9154070-1a34-4e73-923e-6ee24ea89336 · inbound
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05cc146-80ba-4697-b710-292789aa3a39 · inbound
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a38a564c-74ba-4d95-803e-320d6bbb4945 · inbound
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae6274f-493e-4d41-9408-b6a083ff8ea8 · inbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2bac1f-2a01-4ccd-bdec-c3257360aa44 · inbound
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.