Pith. sign in

Paper Citation Record · LEDGER

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation

As of 22 August 2026, this Paper Citation Record lists 100 of 182 outbound references and 0 inbound Pith citation observations for arXiv:2508.10404.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10404 v1

Coverage vector

measured 100 of 182 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:31:42.328780Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 182 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b3d67068-1cc3-4e6f-8be0-7b572fade415 · outbound

This paper cites Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Feature-Level Insights into Artificial Text Detection with Sparse Autoencoders

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.199445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.199445Z digest=sha256:8de732190684216019a53c8f557bc786ac8b824a8edba36c5878ccbd1a984c71

Observation 7fe71ad6-57ad-4266-8612-7fd0530e0210 · outbound

This paper cites I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.253466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.253466Z digest=sha256:b974dc3b9a151f0517807402a9f241f0c013dd6743ccb5b2d6a8e30ae9f8e3d8

Observation 627938d6-e79e-4b04-9299-f5e9e42df1d5 · outbound

This paper cites Applying sparse autoencoders to unlearn knowledge in language models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Applying sparse autoencoders to unlearn knowledge in language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.321750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.321750Z digest=sha256:3831b7596f1458cfff3ec70456d6934f57b7ec88ac80bd506eb484a1e4727fa5

Observation 5d5b5438-d810-47fa-81f3-99b1a64fb6ed · outbound

This paper cites Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.388754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.388754Z digest=sha256:508cea416722410e5c046c224f55012e6847909bb3504b19084e6467de1fd659

Observation 77d84fa0-46e6-4bbc-826a-a011a1d5b69e · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Steering Language Model Refusal with Sparse Autoencoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.454509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.454509Z digest=sha256:3942c913864724d23f2bd5d02ce8d99159ff8c1c28fc5b3d82cf74110129abf0

Observation ba2f3f4b-2430-4bbd-8a35-ef1c2fac1daa · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from large language models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scaling monosemanticity: Extracting interpretable features from large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.552312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.552312Z digest=sha256:bd05493953021eb83abbde0eee8249d08bb2af97d496adefbd646a10d89b447e

Observation 7f76d637-2727-4dd4-bf88-41c468e1229b · outbound

This paper cites Understanding the decisions of large models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Understanding the decisions of large models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.616113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.616113Z digest=sha256:babf4e66733ebe9702cebca11f098e16dff99705aa126f159a61b39377e94cf9

Observation 6265bbc8-9c57-46f2-b42d-6d9d0eb65852 · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.703450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.703450Z digest=sha256:bed3f04375680528b5975795a27323e0e17b900cce809683d2d988b81d340495

Observation e984a50b-2dfc-4d54-ab0a-fa2c943e4b45 · outbound

This paper cites SPML: A DSL for Defending Language Models Against Prompt Attacks.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation SPML: A DSL for Defending Language Models Against Prompt Attacks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.765489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.765489Z digest=sha256:b0fb03e4e6b39f14b7258f8e73fdca3681384e39c5faba0bbd311409452a0925

Observation f0f91128-20e3-4d51-bd0c-f1da26be06c8 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.837394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.837394Z digest=sha256:81cba1ea668eb2cda6da1dc2f85982a5e78e51adea3bb29cce5cc8217e9017c2

Observation 9ae6274f-493e-4d41-9408-b6a083ff8ea8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.913269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.913269Z digest=sha256:d81f1e6fa97f4e8efa68399b12e78f5b8ae782d753a9035580d904d15cc6e76d

Observation 9c9886d1-7f7c-423b-8d0b-5e05f6d5512c · outbound

This paper cites Adversarial Examples Are Not Bugs, They Are Features.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Adversarial Examples Are Not Bugs, They Are Features

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.006974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.006974Z digest=sha256:66606469ef770253d2589e63a652f56ba026ffd38ad2af213a2db6ab6cd5a4ea

Observation 6d688554-cd12-48cc-a878-7d4c7e750fad · outbound

This paper cites an unresolved cited work.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.096752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.096752Z digest=sha256:c4f20a82063d8911335634b36422e6e5c4ea9a6a526cfe462faa3163c7aa0965

Observation 1bc2f18e-501a-4833-a2ea-6835864e9231 · outbound

This paper cites Ddsa: A defense against adversarial attacks using deep denoising sparse autoencoder.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Ddsa: A defense against adversarial attacks using deep denoising sparse autoencoder

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.154418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.154418Z digest=sha256:6ec79799b0dc9eeb44247646e13e892a9ac469d6e84b4b10e48a012c06a0f34a

Observation ca38d9a1-9a9d-45f8-bcfc-7b880faf321e · outbound

This paper cites Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.239962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.239962Z digest=sha256:de52901c4dda6b8863d4533ae5d1b3e7b81753a474b089342bcbb619096d0a93

Observation 0a646562-beb4-455b-9b45-1a18b333f3f8 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.324026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.324026Z digest=sha256:a73add9e741a63900bb3baa295ef9337d25a5161c6e96557c1e54a440150c2ac

Observation c7f6c8a8-3d64-4da7-8f0d-c9205117bdcc · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scaling and evaluating sparse autoencoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.382109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.382109Z digest=sha256:b72bb5bf56c7693eceb20a4a965faf929bfe46de6f195f87ef2745e928a8c7d0

Observation b910cfd2-824c-4ba5-8039-15b0bb4b80b4 · outbound

This paper cites Route Sparse Autoencoder to Interpret Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Route Sparse Autoencoder to Interpret Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.471158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.471158Z digest=sha256:23d3e34e6ba450346dde4283140aacef949eb14520236f5ac2f69bb934af21ff

Observation 71db0da8-18ab-4522-966b-281e91d32b94 · outbound

This paper cites Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.540179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.540179Z digest=sha256:1888f9e17ec97e5351d4465e77387a3fafe2e25968f12964e4a91c45180be7f7

Observation 1dbafd58-6fab-4d99-bc2e-9860eb474bd6 · outbound

This paper cites Sparsegan: Sparse generative adversarial network for text generation.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Sparsegan: Sparse generative adversarial network for text generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.604718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.604718Z digest=sha256:f55265f9448337a9993a6b785a3c81fdd4a607f5d618f7e5206ca3710fac4dab

Observation e25aa254-2579-4ce9-9ad3-283b9ebb2b99 · outbound

This paper cites Real-time segmentation of on-line handwritten arabic script.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Real-time segmentation of on-line handwritten arabic script

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.667105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.667105Z digest=sha256:0cf4cc11a967e0c0db12c358bde9303e5c0a013837c0b9b0b878ebf6f4a283a0

Observation 39a7542f-c27b-409c-94ed-94f66272b653 · outbound

This paper cites Fast classification of handwritten on-line arabic characters.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fast classification of handwritten on-line arabic characters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.758980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.758980Z digest=sha256:b0a148dca5ecb42716a17ef4effa7aa0e533285f9a937cd38f95fabbc22ac104

Observation cf25f153-f647-4869-9713-2dd73cdb6e1b · outbound

This paper cites Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Estimate and Replace: A Novel Approach to Integrating Deep Neural Networks with Existing Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.841903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.841903Z digest=sha256:f428787e1c3fc9df2d45d66c5fb68fd4c6925d8ebdc4e081f244313c71e8acdf

Observation c41a788d-39bb-4780-83d5-16f11f3d3ec6 · outbound

This paper cites Boosting Jailbreak Attack with Momentum.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Boosting Jailbreak Attack with Momentum

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.922403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.922403Z digest=sha256:d4f94380f27abc73ad84a721a3bbb998333f6374fa6616021088856b7766397a

Observation 5a2da513-690e-4b56-99c0-748ef52c5472 · outbound

This paper cites Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:32.991210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:32.991210Z digest=sha256:be0a3cc2d67ecfdf549aa0dbe194a9d7e34bda5677982ab9ad5a13f0b473d58c

Observation abd52154-d50a-4659-b063-cc2bbd32eb46 · outbound

This paper cites Attacking Large Language Models with Projected Gradient Descent.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Attacking Large Language Models with Projected Gradient Descent

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.071055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.071055Z digest=sha256:4a56c4e85119352d618af7100ede07328b5386fa251899c67518fa4a014914bd

Observation 98bbb770-7390-4731-8445-b13309036222 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.151173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.151173Z digest=sha256:b9ac23291fd388a788c50e5055e23a8e5ee34722da8ce80d38972888146578f1

Observation c46d0d5a-1d65-40fa-b1f0-f571846445f8 · outbound

This paper cites Hijacking Large Language Models via Adversarial In-Context Learning.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Hijacking Large Language Models via Adversarial In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.221051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.221051Z digest=sha256:ef99bc36c36be27dd82553fab204f0370e5f2f43a79739dbd63749c8a11eafdb

Observation c8da5bc0-1678-4e7d-a684-7202620ad8c0 · outbound

This paper cites PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.306997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.306997Z digest=sha256:1164d2df27b685b5eda707d4cfb2dcfaa8e8606ba095e343a47229afb684feae

Observation bc76b078-c515-4956-b91d-b93bed59d299 · outbound

This paper cites ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.378159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.378159Z digest=sha256:37c9c893b172b61049ef930e571f56712855b2301d235d594609c5744281aacc

Observation eb2ddbd0-6ab7-48b3-8a03-042e64e1955d · outbound

This paper cites Automatic and Universal Prompt Injection Attacks against Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Automatic and Universal Prompt Injection Attacks against Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.469643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.469643Z digest=sha256:6e9e79d8b3465d7b6e96f98f4a23170f9d552b331c8d6d25c7152a1d02f151ec

Observation db1489f2-1e86-47c7-baba-c5cc027fd0c3 · outbound

This paper cites Improved Generation of Adversarial Examples Against Safety-aligned LLMs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Generation of Adversarial Examples Against Safety-aligned LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.554453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.554453Z digest=sha256:71070c11c57284cb7f6b612574fe6200a77745bb9a773b6d4e1d022de36b5d46

Observation cd0431ae-c85a-46f5-878f-6dcce83602e6 · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Weak-to-Strong Jailbreaking on Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.631918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.631918Z digest=sha256:4ef08509a6d6651319f0e35a73dc8ac917c6c354bf181ed672214689c4ba5434

Observation 2861fd44-e50f-4d4a-8c03-1ae8259f8c1b · outbound

This paper cites Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven Jailbreak

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.736025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.736025Z digest=sha256:eb06a3f9687f4c53d4472937e4bf594818a5667f2366a549da93ea1ebdbca216

Observation 17b10907-dc4b-451b-b535-e0841e28f91e · outbound

This paper cites Don't Say No: Jailbreaking LLM by Suppressing Refusal.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Don't Say No: Jailbreaking LLM by Suppressing Refusal

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.756873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.756873Z digest=sha256:99895395e7098530489bfa1440f854fd0e070a84a76e477cd243740f1e3309db

Observation 36244384-4873-47b0-b7d7-bad4ca7beacb · outbound

This paper cites Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.814112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.814112Z digest=sha256:465049a701ff3209ba5c3442336be34096e0938dd82ede7f0edf1244b95c6142

Observation c7a5a669-88f3-429a-8668-8c7d187bf6ea · outbound

This paper cites Make Them Spill the Beans! Coercive Knowledge Extraction from (Production) LLMs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Make Them Spill the Beans! Coercive Knowledge Extraction from (Production) LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.859800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.859800Z digest=sha256:1ea107dfca75e9ae2e4761abfb2d94016fb905245905bbdfd84f027188aadab5

Observation 512e92fe-7ea0-4a83-b1d9-7f13210766df · outbound

This paper cites COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:33.972622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:33.972622Z digest=sha256:3f5c8eb71ee977ecfbf92103d989e1d589c25139073ad19b0da67cb84336a9eb

Observation 395708b6-947e-43d8-b952-e562b4ba357a · outbound

This paper cites Fast Adversarial Attacks on Language Models In One GPU Minute.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fast Adversarial Attacks on Language Models In One GPU Minute

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.060079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.060079Z digest=sha256:2a8a690852f7bc522c74de68705c9512d01f2431e8d5e56928e92f11f3c0d6b6

Observation 5593cdcc-082e-4ea9-a8d5-680a15548693 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.192244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.192244Z digest=sha256:65e62569128339467d0d913a5f2deda57d148b66cde2a96f92e4edcce0bc83bd

Observation e7c2c932-e692-4384-ad68-e40230d919db · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:34.960696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:34.960696Z digest=sha256:40dd33f58e01d6636841d85f1a4d4827e0ee0894642055d00f09f68aa8cf6489

Observation ba196737-b187-4781-964d-195ff8a829bd · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.139445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.139445Z digest=sha256:1d1c874046159a7ddcf871f1877bee834721a90fc9ecafbda3c95c0d4abfac2a

Observation 42d9129f-7bb4-4215-ae7b-53ab8a841f8e · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.299965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.299965Z digest=sha256:d7b6470f1daf2327409d4106ada9d4379c270055aa8e8e98900d188712049820

Observation e2fffb31-5954-4016-ac38-f43310662a1b · outbound

This paper cites Removing RLHF Protections in GPT-4 via Fine-Tuning.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Removing RLHF Protections in GPT-4 via Fine-Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.407857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.407857Z digest=sha256:dac4b86a62109c40e55f42bb265bcd777457d8e2961959d8e1eaaa039e77daf4

Observation eec2fb5a-a4f3-44fc-835b-60ebb1957c71 · outbound

This paper cites Learning diverse attacks on large language models for robust red-teaming and safety tuning.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Learning diverse attacks on large language models for robust red-teaming and safety tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.553261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.553261Z digest=sha256:b0379ec7b50464cd12c67f3d48e984b877250d44d5ad48870bc37158aa5e2c2a

Observation d72aad72-62ff-46d8-b916-81bc4d105fe5 · outbound

This paper cites Exploiting Novel GPT-4 APIs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Exploiting Novel GPT-4 APIs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.707987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.707987Z digest=sha256:1d0d4d72f9eab16e03fd9348a54407fdf289042b489a53f36453c9be9f67758a

Observation 5ff7773a-6323-4b3c-bce2-1b0df8d3a687 · outbound

This paper cites Steering Language Models With Activation Engineering.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Steering Language Models With Activation Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:36.921320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:36.921320Z digest=sha256:ddfb3847eed6b6c86224c9ae8da4a062c8bbb9b8dd51f6b7fb8198bf3ff5ee4c

Observation 180e7a9a-bfa9-40af-a697-a47145e9befa · outbound

This paper cites Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.044270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.044270Z digest=sha256:ddf4b46314912ec76a63ba3f9f9f7726678818d36d110fd3546c2df07a5b344e

Observation 25011324-4e75-4273-8cdd-0c3901a57082 · outbound

This paper cites Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.155433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.155433Z digest=sha256:3ea0281d2061eec2c670cff3e909466bc39cf9618cd0f9e01726b76f789bcf15

Observation 725280ed-c100-4917-8f79-e5c55895d556 · outbound

This paper cites LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.285221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.285221Z digest=sha256:18937e4a46293c720ea3cd590e9bc21a2b1cc3ae9c0f7c8b3756fc925cb897a7

Observation 98dbcbb2-387b-49ad-99aa-53fe56fca48d · outbound

This paper cites Prompts have evil twins.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Prompts have evil twins

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.394563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.394563Z digest=sha256:2275108f08364100515ffd53ee877172c7e3817ca1895ced05de1de486201a60

Observation 8cd7a346-7dde-45fc-bf34-b6b32029aa32 · outbound

This paper cites PAL: Proxy-Guided Black-Box Attack on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation PAL: Proxy-Guided Black-Box Attack on Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.579293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.579293Z digest=sha256:84b25deb725dcc797a80e9aac67899acafece019123c803e9d501ab73ba5263e

Observation 662af245-6dca-4328-98e9-2e7229c2d120 · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.738668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.738668Z digest=sha256:11387378973ec7a3f03dd2e5ac1c023e7f7805b7467e5665c2d0822b3f902bd0

Observation 50128bf5-2104-432c-8ff7-0d38647da5fd · outbound

This paper cites Uncovering Safety Risks of Large Language Models through Concept Activation Vector.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Uncovering Safety Risks of Large Language Models through Concept Activation Vector

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.875871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.875871Z digest=sha256:5a299108a24f32a11ca6475a5d1ddea23bdf07dcc0df136b212bfce3559a783a

Observation b2f845f7-2d66-4cbd-ab40-9e84ce9c8173 · outbound

This paper cites $\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation $\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:37.935832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:37.935832Z digest=sha256:7132be08679c183fb7d6192021a4283dec670675eb98ac5357a454e21fa8ed31

Observation fd092592-9b9a-45b2-b9c7-b628184385d1 · outbound

This paper cites Low-Resource Languages Jailbreak GPT-4.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Low-Resource Languages Jailbreak GPT-4

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.034933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.034933Z digest=sha256:9df4a7e42767490dc1fcb10f4edb2cb4a2a6246f827d82b276b6cd504aacd3dd

Observation cee7763d-f474-4e9d-a874-767b13499608 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Multilingual Jailbreak Challenges in Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.142157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.142157Z digest=sha256:8ec3f9ffd2b1850034669dce1271751410d7d8aaef7b107bfdf7ced4974b1330

Observation 700de88b-473c-478c-af8f-f9dd1003bb73 · outbound

This paper cites Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Cognitive Overload: Jailbreaking Large Language Models with Overloaded Logical Thinking

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.266606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.266606Z digest=sha256:8b4f662ab87ee0f385996ddc43024d663b7491a24c335ed90c46f22240be9b0a

Observation 94077c99-5a47-4e5b-b033-1f4b3e82c09f · outbound

This paper cites A Cross-Language Investigation into Jailbreak Attacks in Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Cross-Language Investigation into Jailbreak Attacks in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.472071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.472071Z digest=sha256:13adbcac5d654d9c04874e63cc5f3236e05f41dded14a225e78ad653020cdbfc

Observation 33cb0b08-c7f5-4146-80e5-df085cd44bc1 · outbound

This paper cites Does Refusal Training in LLMs Generalize to the Past Tense?.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Does Refusal Training in LLMs Generalize to the Past Tense?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.588823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.588823Z digest=sha256:7d5c7b213b37647de915e5bcbad98a2a3d6eaf32282de2bcd6350cc054706ddf

Observation c8a82589-fba1-4a2d-a4f7-8ccec9762110 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.691439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.691439Z digest=sha256:ad1ca9c089778aa2426995ec607f220c84f2469a57dd2806dc019282effa216f

Observation 5338dce9-2623-4bc2-bb50-8d8089c4cec1 · outbound

This paper cites Tree of Attacks: Jailbreaking Black-Box LLMs Automatically.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.748228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.748228Z digest=sha256:47ffdb1bec8e59c5c2e1cac76700547cc214c42f13741a95502ffc6a84c5973a

Observation 0cb88860-73dd-47a7-a440-dae0fdc7137d · outbound

This paper cites Distract Large Language Models for Automatic Jailbreak Attack.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Distract Large Language Models for Automatic Jailbreak Attack

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.816400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.816400Z digest=sha256:e5a333b9c25eb85bb60ef6d2daa2ab14c120a38c2f97017711fade722e9b57b8

Observation 988661c8-3b8b-4dae-b34b-8ba76988e131 · outbound

This paper cites GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:38.916981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:38.916981Z digest=sha256:03da5e9d0dc8993a2c75c69a27090aa2bdd775ddd47131537e2053a6e40ac3ea

Observation d301f50f-4848-491b-b0b0-448c66818f48 · outbound

This paper cites Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.012270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.012270Z digest=sha256:94a0dd9715f3effed00505f48ab69c144ec41437546a2028fd02c642d4739f36

Observation db9352be-3e94-4c12-a714-e05670708b7b · outbound

This paper cites Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.106420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.106420Z digest=sha256:13fbbd715d41d196b3a79eb862b458fee8cfd3b328f83e5af10d96aeed6d723d

Observation b6572016-2ca6-486f-b7cd-6ddbf51ce786 · outbound

This paper cites Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Foot In The Door: Understanding Large Language Model Jailbreaking via Cognitive Psychology

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.185931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.185931Z digest=sha256:024643e9b8c7a93de61fde9a81df6b728b2c4e77736d4acd361e9dcf6238168b

Observation f9832347-69c1-485d-b07d-92509c29f29d · outbound

This paper cites Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.270858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.270858Z digest=sha256:ceb0527ae081d262bb823955ceb1f084a6bcedceaebdfb6d3e1c7fc80e2fc616

Observation cbadbbac-3e64-4778-90c7-60019d30923c · outbound

This paper cites Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.383251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.383251Z digest=sha256:00fdf3d02991dd45b3286c1e3ee0bfeff5afdf95dde17d8c5bd09f395f6f25d6

Observation ddc25817-9bbc-4f3c-9f6c-b6894256d338 · outbound

This paper cites Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.494024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.494024Z digest=sha256:73814a7ddcd6cd19147e6fe9eb4c72f370b8b08421c317b4c95f44fd50b55630

Observation a1cd6c39-faad-47d4-9419-a2595bb633ec · outbound

This paper cites How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.578072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.578072Z digest=sha256:45593fea494e39b399fa4f07462dc0d375593ee0dd8c483bd90f777d004dec0f

Observation 7a546f1b-868d-443b-9131-adc1338483e6 · outbound

This paper cites GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.676177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.676177Z digest=sha256:fe9443916b15dbf9c0b65a0acd1ffac736fc3a3de840699fe7623b832e88218b

Observation 994455b5-e1c7-40d3-a964-8916c933abce · outbound

This paper cites Jailbreaking proprietary large language models using word substitution cipher.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking proprietary large language models using word substitution cipher

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.753248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.753248Z digest=sha256:63a10a1a23316f14a367f42f00625b50180632fbe6f627fac597ab3256278db9

Observation d3394ec4-0e10-4687-bc61-bf5b581bd413 · outbound

This paper cites Endless Jailbreaks with Bijection Learning.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Endless Jailbreaks with Bijection Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.837061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.837061Z digest=sha256:67588087df5b18e96d1b49747f9a93840ce6e873a2914855c2e7459c6718324f

Observation 78b3b375-eb49-4fd2-8927-49c564fdca30 · outbound

This paper cites ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:39.978896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:39.978896Z digest=sha256:19d782f436c68bbcfdd639a6230fd1b451e55abeb273d2d4361074a20ace2bab

Observation fa6d8ef1-779b-4e8f-956b-f105db6f570c · outbound

This paper cites Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Mind the Inconspicuous: Revealing the Hidden Weakness in Aligned LLMs' Refusal Boundaries

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.072419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.072419Z digest=sha256:818be2dfabf2f991b0de1a0e2e286b3c03d36f66e16e7e69db2e500809b5adb3

Observation 21da2a50-cc45-46b1-bf0a-ac8480d8d34d · outbound

This paper cites Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.174290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.174290Z digest=sha256:ab2e079b66383c9e02a8ff6e8abf6e08a7b24da5d6f98bbcf7c7ba1a2f1dcc86

Observation 1b717d4d-1dc1-44c4-9960-8658c2b0d429 · outbound

This paper cites AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.261208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.261208Z digest=sha256:0a96854d5b4d3a1a5a8f1933f4bfbca1540a0aeb25c586d16ed4f337f23cab3d

Observation 3ecd295a-6107-4aa3-9682-247e5b5f1aa8 · outbound

This paper cites Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit Clues

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.373214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.373214Z digest=sha256:bda323f1dbde5855b2a57e092b7b0a300f6b8a75cd94641e2750c140ede4933d

Observation 79ee3035-118a-407b-8085-916f7b3b6269 · outbound

This paper cites DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.480555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.480555Z digest=sha256:62bb41e8caa740b19879aa9f8e6c8b0d62c482dd083289b90087900f48fff53c

Observation 14348ef6-23eb-4fc7-95c7-32497fdfcd76 · outbound

This paper cites Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and Reconstruction

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.561131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.561131Z digest=sha256:2f559a0aee9b47e606a9e30ffcaebc3b6cfbb0d60f6cc55816c6b3786ddefa7e

Observation 4996865d-87dd-432e-bb2b-4d92ae4f1f62 · outbound

This paper cites CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.641624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.641624Z digest=sha256:9b914941c4b6abb0e85ec32cd534e07c7bc4fd2b1affb5b29b06683ae6508e49

Observation 7ec4bc9c-2de1-4dc3-b8e1-8a9345ed3384 · outbound

This paper cites CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.726264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.726264Z digest=sha256:0625011c6d97f5999d7a5cd05ad9132988b9718655fae3b7ff3e9863797441e1

Observation 1eccfcfc-c96a-4ec0-9160-2dc7dcd13a09 · outbound

This paper cites A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.813666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.813666Z digest=sha256:8ede842aa10668ac4912736d8754f818c3f2bf68fa51d45b1092408655988020

Observation 75e46ad7-42c9-47ad-9012-227974bb1b5c · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:40.900099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:40.900099Z digest=sha256:6137508e8f9ff2ee326328c645511db83ead546d966bc95a4682db6dca56b3f9

Observation 5062fa6e-774e-4d32-9460-93e018b24fd0 · outbound

This paper cites LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.015324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.015324Z digest=sha256:0f64120cae34582bf56bc4f5cd18ab5aaa5af816f2585cfa21012ddd0504f752

Observation 503994c0-873c-4d01-bc23-7d6f96f3ef5c · outbound

This paper cites Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.107700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.107700Z digest=sha256:e9bc2a2e1268ed9b50efb666df0300c866199270c30c3b4de411f668c20660c2

Observation b527fb2d-3981-4b5b-8a4a-fea67a2d4492 · outbound

This paper cites Harris, and Marcel Carlsson.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Harris, and Marcel Carlsson

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.204266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.204266Z digest=sha256:fcff2ef9375ff45f33910a1e767a0e91dfe84390ba14d93ed05c008a93c954a5

Observation f9b37606-de90-4af1-a6e7-3428a4bc4d79 · outbound

This paper cites Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Hidden You Malicious Goal Into Benign Narratives: Jailbreak Large Language Models through Logic Chain Injection

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.267214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.267214Z digest=sha256:ff98a8fad7e22cc0e0374185fb1424abde5fd3b86764665f7d2e7fd337f8c7d2

Observation a35bf87d-fa26-4e8b-bd01-a5cef0f76913 · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.373394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.373394Z digest=sha256:d9cb97a72010ddd2eba2081e5adf5d6e76fdf8000a15bff425d8f2bb6303dcc0

Observation 9741435e-9d62-4bda-8b93-9ddb7eaf1513 · outbound

This paper cites Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.449950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.449950Z digest=sha256:816a4381e22064c0ff54d5ae646d3d678b3c09eb088b0c09595a1617b8854936

Observation 3e071d78-c297-4a80-8084-43d8cac74c4f · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.559678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.559678Z digest=sha256:eaa4488743863532be36a735599687eeadcb02a5970ef154a0fb57063ae24c2f

Observation 203f472a-cddb-40ee-9ce1-98549ebf1a3e · outbound

This paper cites Masterkey: Automated jailbreaking of large language model chatbots.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Masterkey: Automated jailbreaking of large language model chatbots

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.666187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.666187Z digest=sha256:ab4f76598ad43a969aae99c5fb0311552b7165bf60a20d387bdcf20604b69df7

Observation 1a625e3c-4afa-4f3b-b1e3-eb81901693df · outbound

This paper cites Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.745657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.745657Z digest=sha256:0dc9b25da24df9e8636a4cd024b830507423c72b8e90f7a4e4234b66cae01642

Observation 3a686db4-7f0f-4504-8970-4b21fbfc98b9 · outbound

This paper cites Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Guard: Role-playing to generate natural-language jailbreakings to test guideline adherence of large language models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.840089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.840089Z digest=sha256:34970b7f47ebcbcd54d353795d1ffcbb5dedb9df74fadb30f5cc062fce32aacb

Observation 7b336e46-60e1-46ee-b303-d1a7e4301c73 · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:41.927775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:41.927775Z digest=sha256:1d0e67994c752dd7a819f396325de91c1f5839a7a42d9121c2be49e41e5bda49

Observation 0d6fe14f-9446-49ad-91f7-75c0948f0528 · outbound

This paper cites All in how you ask for it: Simple black-box method for jailbreak attacks.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation All in how you ask for it: Simple black-box method for jailbreak attacks

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.032201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.032201Z digest=sha256:899e0746e7935d48e3d4c8357c8bd0fe8ff18a4b5347d60bd1b871f33255c8fb

Observation a3f44fa2-acd8-4f49-8551-f7f32084e4ed · outbound

This paper cites Open Sesame! Universal Black Box Jailbreaking of Large Language Models.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.141679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.141679Z digest=sha256:e563856be672d66f7c49fb990b9e9ca14be3ee6a589cd7835820e08ff8b8ac02

Observation ee7830ba-e151-4cca-829e-887234174a10 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.237749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.237749Z digest=sha256:441cc4f2f8e0af3189ba4dfd82194d32b28d8417a7d55672065232d948f13f73

Observation 1fdf52f4-855e-4da7-b519-afef192658e4 · outbound

This paper cites Query-Efficient Black-Box Red Teaming via Bayesian Optimization.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Query-Efficient Black-Box Red Teaming via Bayesian Optimization

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.328780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.328780Z digest=sha256:d51b3e34f9e6638ac7ee087e3cd68009ef9428542cfd96a10b4a0c1763297509

Pith citing papers

No inbound Pith citation observations are available.