Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:31:42.328780Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 100 of 182 outbound references and 0 inbound Pith citation observations for arXiv:2508.10404.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T20:31:42.328780Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 182 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b3d67068-1cc3-4e6f-8be0-7b572fade415 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe71ad6-57ad-4266-8612-7fd0530e0210 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627938d6-e79e-4b04-9299-f5e9e42df1d5 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Applying sparse autoencoders to unlearn knowledge in language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d5b5438-d810-47fa-81f3-99b1a64fb6ed · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d84fa0-46e6-4bbc-826a-a011a1d5b69e · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Steering Language Model Refusal with Sparse Autoencoders
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2f3f4b-2430-4bbd-8a35-ef1c2fac1daa · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scaling monosemanticity: Extracting interpretable features from large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f76d637-2727-4dd4-bf88-41c468e1229b · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Understanding the decisions of large models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6265bbc8-9c57-46f2-b42d-6d9d0eb65852 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e984a50b-2dfc-4d54-ab0a-fa2c943e4b45 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation SPML: A DSL for Defending Language Models Against Prompt Attacks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f91128-20e3-4d51-bd0c-f1da26be06c8 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae6274f-493e-4d41-9408-b6a083ff8ea8 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9886d1-7f7c-423b-8d0b-5e05f6d5512c · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Adversarial Examples Are Not Bugs, They Are Features
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d688554-cd12-48cc-a878-7d4c7e750fad · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc2f18e-501a-4833-a2ea-6835864e9231 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Ddsa: A defense against adversarial attacks using deep denoising sparse autoencoder
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca38d9a1-9a9d-45f8-bcfc-7b880faf321e · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a646562-beb4-455b-9b45-1a18b333f3f8 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f6c8a8-3d64-4da7-8f0d-c9205117bdcc · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scaling and evaluating sparse autoencoders
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b910cfd2-824c-4ba5-8039-15b0bb4b80b4 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Route Sparse Autoencoder to Interpret Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71db0da8-18ab-4522-966b-281e91d32b94 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbafd58-6fab-4d99-bc2e-9860eb474bd6 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparsegan: Sparse generative adversarial network for text generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25aa254-2579-4ce9-9ad3-283b9ebb2b99 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Real-time segmentation of on-line handwritten arabic script
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a7542f-c27b-409c-94ed-94f66272b653 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fast classification of handwritten on-line arabic characters
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf25f153-f647-4869-9713-2dd73cdb6e1b · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c41a788d-39bb-4780-83d5-16f11f3d3ec6 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Boosting Jailbreak Attack with Momentum
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2da513-690e-4b56-99c0-748ef52c5472 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd52154-d50a-4659-b063-cc2bbd32eb46 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Attacking Large Language Models with Projected Gradient Descent
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bbb770-7390-4731-8445-b13309036222 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c46d0d5a-1d65-40fa-b1f0-f571846445f8 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Hijacking Large Language Models via Adversarial In-Context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8da5bc0-1678-4e7d-a684-7202620ad8c0 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc76b078-c515-4956-b91d-b93bed59d299 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2ddbd0-6ab7-48b3-8a03-042e64e1955d · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Automatic and Universal Prompt Injection Attacks against Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1489f2-1e86-47c7-baba-c5cc027fd0c3 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Generation of Adversarial Examples Against Safety-aligned LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2861fd44-e50f-4d4a-8c03-1ae8259f8c1b · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b10907-dc4b-451b-b535-e0841e28f91e · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Don't Say No: Jailbreaking LLM by Suppressing Refusal
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36244384-4873-47b0-b7d7-bad4ca7beacb · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7a5a669-88f3-429a-8668-8c7d187bf6ea · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Make Them Spill the Beans! Coercive Knowledge Extraction from (Production) LLMs
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512e92fe-7ea0-4a83-b1d9-7f13210766df · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 395708b6-947e-43d8-b952-e562b4ba357a · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fast Adversarial Attacks on Language Models In One GPU Minute
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5593cdcc-082e-4ea9-a8d5-680a15548693 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7c2c932-e692-4384-ad68-e40230d919db · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba196737-b187-4781-964d-195ff8a829bd · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d9129f-7bb4-4215-ae7b-53ab8a841f8e · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2fffb31-5954-4016-ac38-f43310662a1b · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Removing RLHF Protections in GPT-4 via Fine-Tuning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec2fb5a-a4f3-44fc-835b-60ebb1957c71 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Learning diverse attacks on large language models for robust red-teaming and safety tuning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d72aad72-62ff-46d8-b916-81bc4d105fe5 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Exploiting Novel GPT-4 APIs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff7773a-6323-4b3c-bce2-1b0df8d3a687 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Steering Language Models With Activation Engineering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180e7a9a-bfa9-40af-a697-a47145e9befa · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25011324-4e75-4273-8cdd-0c3901a57082 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725280ed-c100-4917-8f79-e5c55895d556 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98dbcbb2-387b-49ad-99aa-53fe56fca48d · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Prompts have evil twins
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd7a346-7dde-45fc-bf34-b6b32029aa32 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation PAL: Proxy-Guided Black-Box Attack on Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 662af245-6dca-4328-98e9-2e7229c2d120 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50128bf5-2104-432c-8ff7-0d38647da5fd · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Uncovering Safety Risks of Large Language Models through Concept Activation Vector
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f845f7-2d66-4cbd-ab40-9e84ce9c8173 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation $\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd092592-9b9a-45b2-b9c7-b628184385d1 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Low-Resource Languages Jailbreak GPT-4
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cee7763d-f474-4e9d-a874-767b13499608 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Multilingual Jailbreak Challenges in Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700de88b-473c-478c-af8f-f9dd1003bb73 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94077c99-5a47-4e5b-b033-1f4b3e82c09f · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Cross-Language Investigation into Jailbreak Attacks in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33cb0b08-c7f5-4146-80e5-df085cd44bc1 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Does Refusal Training in LLMs Generalize to the Past Tense?
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8a82589-fba1-4a2d-a4f7-8ccec9762110 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5338dce9-2623-4bc2-bb50-8d8089c4cec1 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb88860-73dd-47a7-a440-dae0fdc7137d · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Distract Large Language Models for Automatic Jailbreak Attack
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988661c8-3b8b-4dae-b34b-8ba76988e131 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d301f50f-4848-491b-b0b0-448c66818f48 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9352be-3e94-4c12-a714-e05670708b7b · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6572016-2ca6-486f-b7cd-6ddbf51ce786 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9832347-69c1-485d-b07d-92509c29f29d · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbadbbac-3e64-4778-90c7-60019d30923c · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddc25817-9bbc-4f3c-9f6c-b6894256d338 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cd6c39-faad-47d4-9419-a2595bb633ec · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a546f1b-868d-443b-9131-adc1338483e6 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 994455b5-e1c7-40d3-a964-8916c933abce · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking proprietary large language models using word substitution cipher
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3394ec4-0e10-4687-bc61-bf5b581bd413 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Endless Jailbreaks with Bijection Learning
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78b3b375-eb49-4fd2-8927-49c564fdca30 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa6d8ef1-779b-4e8f-956b-f105db6f570c · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21da2a50-cc45-46b1-bf0a-ac8480d8d34d · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b717d4d-1dc1-44c4-9960-8658c2b0d429 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ecd295a-6107-4aa3-9682-247e5b5f1aa8 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ee3035-118a-407b-8085-916f7b3b6269 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14348ef6-23eb-4fc7-95c7-32497fdfcd76 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4996865d-87dd-432e-bb2b-4d92ae4f1f62 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec4bc9c-2de1-4dc3-b8e1-8a9345ed3384 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eccfcfc-c96a-4ec0-9160-2dc7dcd13a09 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e46ad7-42c9-47ad-9012-227974bb1b5c · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation DeepInception: Hypnotize Large Language Model to Be Jailbreaker
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5062fa6e-774e-4d32-9460-93e018b24fd0 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 503994c0-873c-4d01-bc23-7d6f96f3ef5c · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b527fb2d-3981-4b5b-8a4a-fea67a2d4492 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Harris, and Marcel Carlsson
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9b37606-de90-4af1-a6e7-3428a4bc4d79 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a35bf87d-fa26-4e8b-bd01-a5cef0f76913 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9741435e-9d62-4bda-8b93-9ddb7eaf1513 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e071d78-c297-4a80-8084-43d8cac74c4f · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 203f472a-cddb-40ee-9ce1-98549ebf1a3e · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Masterkey: Automated jailbreaking of large language model chatbots
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a625e3c-4afa-4f3b-b1e3-eb81901693df · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a686db4-7f0f-4504-8970-4b21fbfc98b9 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b336e46-60e1-46ee-b303-d1a7e4301c73 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d6fe14f-9446-49ad-91f7-75c0948f0528 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation All in how you ask for it: Simple black-box method for jailbreak attacks
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3f44fa2-acd8-4f49-8551-f7f32084e4ed · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7830ba-e151-4cca-829e-887234174a10 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdf52f4-855e-4da7-b519-afef192658e4 · outbound
Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Query-Efficient Black-Box Red Teaming via Bayesian Optimization
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.