Pith. sign in

Paper Citation Record · LEDGER

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.07402.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07402 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:40:20.752585Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:03.319920Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy9
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ac0cda6-c77c-48c6-99e7-985e17627304 · outbound

This paper cites Universal language model fine-tuning for text classifica- tion, 2018.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal language model fine-tuning for text classifica- tion, 2018

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.505766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.514067Z digest=sha256:139aed87925b8aad4dfbafb8939c407bda2d6bf1672ec10efa17c0614ed9fbf4

Observation 7a8c6123-ac07-45b4-aaef-ea22fed4a683 · outbound

This paper cites Training language models to follow instructions with human feedback.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.519392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.519392Z digest=sha256:2e888343a11c51154f375732af236ed2fb7ba3152dbf91310820a9c877f01c30

Observation d7e06a19-25a4-49ea-9cf4-46033b91e68d · outbound

This paper cites Scaling instruction-finetuned language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Scaling instruction-finetuned language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.525164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.525164Z digest=sha256:25312dd187e606f6121248715b46c6b538883fe50ae7c5b54ca438afa4d2cbe3

Observation f687c9cc-53af-4811-acf1-5026fc9bf588 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Fine-Tuning Language Models from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.529978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.529978Z digest=sha256:f19e366d2ad50699b81d95bc9e970131d2b044df76957b903f3a94a117006d37

Observation 6938c8e7-dd8a-473b-a01a-e7b0717c189e · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Direct preference optimization: Your language model is secretly a reward model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.535236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.535236Z digest=sha256:b612abc043ca6693cab0428b5dff493cf12eed2207fa07f14064da6c654f1905

Observation be8403d9-f7ee-4916-a14e-cf1608a7fa44 · outbound

This paper cites Jailbreakchat.com, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreakchat.com, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.439636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.539989Z digest=sha256:7f1ec23806a876589de321e05d8bf1069a82d0003506b745fbfd807550f54f1a

Observation f6ff2025-0476-4e55-9c57-d49f5a91cb13 · outbound

This paper cites Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbroken: How does llm safety training fail? Advances in Neural Information Processing Systems, 36:80079–80110, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.545428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.545428Z digest=sha256:a0aa8359a829f37a4816707ae679ef9ddc22666ed36fcf2049a0bfaf19b2246b

Observation 2f1a9f34-036d-4335-8a32-9f3ab8954e01 · outbound

This paper cites Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36:61478– 61500, 2023.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36:61478– 61500, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.415500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.550874Z digest=sha256:05c966249ed44d2a21c3f6dafa389dae58a1575817ac68125dc273dc99af3632

Observation bd3b9ee6-56ab-4793-ba5a-dc13d7fd6fcf · outbound

This paper cites do anything now.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures do anything now

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.555354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.555354Z digest=sha256:91b0db8ea386eddc7c30521ea5f98b2dd01b9fc378f558c0867a83dabdd1dbdc

Observation a1753045-f0d9-442a-aecf-64ed9abd11ff · outbound

This paper cites Don’t listen to me: understanding and exploring jailbreak prompts of large language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don’t listen to me: understanding and exploring jailbreak prompts of large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.390445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.560611Z digest=sha256:0d649bf27111805e1fa2f95a4eeb6808f7320610b33a1a290b4770cabe4e889b

Observation ac5bca8e-c73d-4f9e-90ca-fa8bed0495f8 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.565977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.565977Z digest=sha256:d0a63181d04a6cb6d84d74e94668748f29474429838ef8f48b1d2bd9183d262a

Observation 36c6d1dc-e66c-4e72-868b-e30305efdd7a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.570649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.570649Z digest=sha256:ca391d566c60a46325ee71571bf87cf288b20f07afbce00532dfb452e324ca74

Observation 7d463ea2-dc58-4f1d-b807-d3b19f063b03 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.575426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.575426Z digest=sha256:1c7969248b94e1f5ad8ee2d239f7411756aad30d47fe8dc326499852e119b5f2

Observation ff06be60-b1ab-4ac7-b4c7-f3d678f2af6b · outbound

This paper cites MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.580018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.580018Z digest=sha256:0aaa7a4b073e92da964291ced9e669da3e1b9d6d7bf7a61cdb8e9e8be64ef478

Observation afa7ba33-3fe7-41de-877d-2461d41a0ce3 · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.585362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.585362Z digest=sha256:579f1d19c311a6d26766c3ef3922c9aca494b7a314b5a7f8d362ed9a05b23fb4

Observation d202259b-976c-48e2-8129-af12e45802bc · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.590342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.590342Z digest=sha256:9b469f5cc34e4c1e323bea33489e5de1d4952311ce2713ae5cdca2656a7014a1

Observation da3d46d1-a7b7-4829-a1a1-a80d0775e3ca · outbound

This paper cites AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.595942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.595942Z digest=sha256:68a4f724f112c1e2f854b1aff5348bcabe6cfa55956d92d1c4ee4a34da5eb14e

Observation 8cd526e5-5d8d-4e82-91ab-2b9b4db91aa4 · outbound

This paper cites Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.601664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.601664Z digest=sha256:930bbc335ad5393b93408cdfbe5b799fc73ee09f156e631e441064374db9d67e

Observation 595452c4-6357-4760-8f52-30d432d8a57b · outbound

This paper cites Weak-to-Strong Jailbreaking on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Weak-to-Strong Jailbreaking on Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.606600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.606600Z digest=sha256:af24c79450522727c6969d2387b2c31ca7f9ada9bced66ff5a70ea3189d30567

Observation 81904555-64f5-41c7-a0fe-bfecc3e67c03 · outbound

This paper cites AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.611175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.611175Z digest=sha256:0e26a73022294625abe0f7ac6531c37c033042698f915d2ecb1b85b1db56a87f

Observation 72aa392a-d53a-4eb3-8989-88f9158796ad · outbound

This paper cites Improved Techniques for Optimization-Based Jailbreaking on Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Improved Techniques for Optimization-Based Jailbreaking on Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.616360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.616360Z digest=sha256:e351b7ed0f535fd68a81669bc6279fb7714ce6996598cea812382530cb084c5f

Observation 72ba783a-5b5a-4037-a109-1035b69cdfe8 · outbound

This paper cites Don't Say No: Jailbreaking LLM by Suppressing Refusal.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Don't Say No: Jailbreaking LLM by Suppressing Refusal

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.621417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.621417Z digest=sha256:7e82599c42c8579ae270cf8d5dba5d9ba4689a99cde0d70bafe80a25ede541fe

Observation 95cbf8a9-6f0a-4ea8-b67d-d55cec96b58f · outbound

This paper cites AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.626685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.626685Z digest=sha256:ca98b40500492a3862784e4f6dce27498221a65cb3f44014bf7e16649e29d1cf

Observation 37d45666-6cd8-4238-a472-ce1bc8adba7e · outbound

This paper cites AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.632046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.632046Z digest=sha256:fa5c7f6833f771a74d9a52bb9345b7050e07e7135f25f303db51a30ad4af17ae

Observation 7d542fe5-5475-47de-b4ff-013c8b89ec90 · outbound

This paper cites Jailbreak Attacks and Defenses Against Large Language Models: A Survey.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreak Attacks and Defenses Against Large Language Models: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.636809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.636809Z digest=sha256:37545478555aa9d67136a8adec808f178b11cbb94d37353efb37e3ee08371df7

Observation 9bcc24c3-5661-475e-b745-7e0a568132f8 · outbound

This paper cites Llm jailbreak attack versus defense techniques–a comprehensive study.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Llm jailbreak attack versus defense techniques–a comprehensive study

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.365440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.642041Z digest=sha256:59659b42724e64d889f5f254ce90abc8af7f6efa1a368744da058f2c3ffd5a4e

Observation ec8ff256-1793-4bab-b989-42bd5d49e3b0 · outbound

This paper cites LLM-Safety Evaluations Lack Robustness.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures LLM-Safety Evaluations Lack Robustness

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.646900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.646900Z digest=sha256:1a200b0263648e040904298648021373cd4ffac3f51ff22d2ad27ff1b1452881

Observation b2eaf344-2622-4f66-b217-4850c0eb5904 · outbound

This paper cites JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.651700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.651700Z digest=sha256:e9b445a03e03fb065a7f07436ef6c3d14149bbe069bdf9d3322ce3cfca364869

Observation f20d886e-3a37-4af6-864a-5e093bcc8e5e · outbound

This paper cites Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.349668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.657041Z digest=sha256:139401e02034a48753f20c102dbd7f8c518d17478c7f0a0ceec301288abe9904

Observation ebeab4e4-5f8f-4310-a1ba-9de7ff330d25 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.661675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.661675Z digest=sha256:17c179581c111ac727e3b2e862cc248d891a3a1dae636164531efb43b3844dba

Observation 51c51fb7-d940-44ea-83c4-8445e1beb86d · outbound

This paper cites JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.666525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.666525Z digest=sha256:2b8234888eae047bc25e38cc493a0e587b82ed74938cbc7b2f8857b060f0f0ae

Observation 3f30fa75-a85a-42b8-b79d-0bb40852e600 · outbound

This paper cites Clas 2024: The competition for llm and agent safety.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Clas 2024: The competition for llm and agent safety

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.333247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.671409Z digest=sha256:5a7a1f547c5d689ec9232bcfcd547db23078507eff42856eab01a5041927ecce

Observation 7db82a51-965f-4deb-8037-686fccefef3d · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures A StrongREJECT for Empty Jailbreaks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.676687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.676687Z digest=sha256:c4261523036fa52960243df77d9b136aa0bda07c73471bc3874a62db5ca67d9e

Observation 606f5adb-80f0-41a5-8703-0e69f61b6ff6 · outbound

This paper cites The art of saying no: Contextual noncompliance in language models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The art of saying no: Contextual noncompliance in language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.682110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.682110Z digest=sha256:49a973d92e6567f904a0437bf2c2cc8685a80a381d425673f40d6587cb4db372

Observation 931ffb9d-d955-4d82-b723-9b4705eebfc9 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.307575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.686767Z digest=sha256:5b2964cbce3321e151950ba1da52e11f3b4688011d0ddd7b30d3fe7c8765dbef

Observation a0dc3ecf-570a-4d1e-ae01-7785e1e52f36 · outbound

This paper cites Jailbreaking large language models with symbolic mathematics, 2024.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Jailbreaking large language models with symbolic mathematics, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:40:21.291327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.691795Z digest=sha256:5e78723ff540b70896b37ac94d324a9d2a3a7fc91ee3bf28095d98478bd9d575

Observation 5b995f12-018a-4ea8-83a6-6fbb92a7a99d · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.696608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.696608Z digest=sha256:5168370550187c3862429242d3a3a19f2e54d7aeb320acf1b7ad77b3086ff7ea

Observation 34268b1e-ffbf-4159-9ee4-6aeecca7a4bf · outbound

This paper cites Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.701570Z digest=sha256:d58a36b45d732474b29a38783e6275241d90edb46bf55ff562125bac42e37333

Observation 37041393-4ed4-4b92-a263-8c859f7bd858 · outbound

This paper cites Rethinking How to Evaluate Language Model Jailbreak.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Rethinking How to Evaluate Language Model Jailbreak

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:40:20.926237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:40:20.706822Z digest=sha256:b5f389f2a65f8aba4ee519de8c459a7fd0b0a79f8042abcf109a974450653edb

Observation 2a7751e3-04d6-4088-b57d-1229ae8704ae · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.711744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.711744Z digest=sha256:cbe54cf80b2a0425ecde81eff0cccb38db4527fa23e8b516cddf1d791aaed201

Observation a1070a30-5b1e-4ab6-992e-ec2f7ad6ff83 · outbound

This paper cites The Jailbreak Tax: How Useful are Your Jailbreak Outputs?.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures The Jailbreak Tax: How Useful are Your Jailbreak Outputs?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.716732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.716732Z digest=sha256:88a587497c617948f7ff768223116ee71b76df86ff5db1abf59a759a3b4f9ddf

Observation 26835b23-1641-4939-ab42-09bdb134a4a5 · outbound

This paper cites AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.722684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.722684Z digest=sha256:19e68503d4913e7ded876c2dc40e18ca194edb971e878c7e778aaeb001479122

Observation 2912aaeb-2854-4b65-8fa6-47bfc05b8e3b · outbound

This paper cites Multimodal Situational Safety.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multimodal Situational Safety

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.727309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.727309Z digest=sha256:a34eb60264472cd7a7d868819d360b92be5af4522aa3a93e689ac04fc8d52206

Observation b139af91-eb49-41c9-8e53-aad5732b7f87 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.732373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.732373Z digest=sha256:f6e5828d443e6858552e6984bcc18e6abac92a7396a820007d994dd0d5ee2fb6

Observation 19b5f8c9-896d-4480-b287-6ebff5ed5fad · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.737235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.737235Z digest=sha256:c0738a2765b0e2532a57e3b14a65f81bc33ab1cce5f3401bb668945c6bd31d0a

Observation d8e565ca-a31a-48a8-acc5-b88382c9daf0 · outbound

This paper cites Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.742100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.742100Z digest=sha256:9e385abdf2715fc4f283d32ddc5933ee7a883cc0697389905c172ec53ceda014

Observation f9ee1f79-47cb-4735-aa43-bfa24bca4413 · outbound

This paper cites Multilingual Jailbreak Challenges in Large Language Models.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Multilingual Jailbreak Challenges in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.747279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.747279Z digest=sha256:c124703fef812b6facc83ccd0596e55a23e76ab5dfb4de29afa6503f5a256a6b

Observation c549df12-5d4b-49ff-9913-808afde46e72 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:20.752585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:40:20.752585Z digest=sha256:0aaa5dd76cfad451954d4918e3b3108428bcfbb6b4bef15c1845cad14408745b

Pith citing papers

Observation 394b7429-0807-402b-917e-b5967f811225 · inbound

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions cites this paper.

How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Beyond Jailbreaks: Revealing Stealthier and Broader LLM Security Risks Stemming from Alignment Failures

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:03.319920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:03.319920Z digest=sha256:adfb1807f62b3d48d8c299f0c925214e997b2f28eacfb356b03095c14e713189