Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:08:05.386345Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 72 inbound Pith citation observations for arXiv:2404.01318.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T06:08:05.386345Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T11:28:36.924563Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T06:15:00.866473Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f52f41d0-e120-43c0-9c46-915c932a32de · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Get my drift? Catching LLM Task Drift with Activation Deltas
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb1a5952-78e4-43b9-9b13-f69f8f2895d1 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Llama 3 model card
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee1558a9-17a5-4cf4-b065-2372b912e852 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Croissant: A metadata format for ml-ready datasets
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 94d4132a-2e54-4831-b5a4-1a3ff306350c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbreak chat
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d81435c3-e78d-408d-8c71-d1a8023331eb · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Detecting Language Model Attacks with Perplexity
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9b4695f0-3376-423e-ae6d-b5f91e084b40 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bf1517d-d7fa-42a2-885c-0c52dff62180 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Refusal in llms is mediated by a single direction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d1a78591-99a7-4ff0-b1ad-6a855965ff3b · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Are aligned neural networks adversarially aligned? Advances in Neural Information Processing Systems, 36
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 417ae944-6027-4753-b94c-3d629d48b954 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Non-determinism in gpt-4 is caused by sparse moe
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da0555d9-a6eb-406c-a06b-f113cc3ac3ff · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7beb1d43-7e0f-4423-9e8f-cbfbe07a1595 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Robustbench: a standardized adversarial robustness benchmark
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bde5b79-fc98-4f7d-bd61-96df390e823c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Multilingual Jailbreak Challenges in Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dcd6d2d5-77a7-4685-a090-86134a0fa128 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Attacking Large Language Models with Projected Gradient Descent
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40a44745-888d-4f5e-8e58-53bfc3307b1c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Gemini v1.5 report
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3d40c8bc-92e7-45af-9a9b-13b4450f0166 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Query-Based Adversarial Prompt Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c80d48f4-6846-45c8-be08-ffe06c6a00fc · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Measuring massive multitask language understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 451e27d6-fd76-4838-a4d0-ac044ccecea1 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 14eda907-262d-4a5e-9978-ccab91d7815c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b032a506-52d4-41cf-8c2f-027640a9e8d0 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7788ce21-f353-4643-aa33-3dfeb7d0fc56 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 658d527c-31f9-449c-a5a1-f82f5704d7a6 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Defending Large Language Models against Jailbreak Attacks via Semantic Smoothing
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31ddbe07-947e-412c-b49b-1f4779586d60 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Mixtral of Experts
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a44d343e-a3e5-4c9b-ab65-b7f13694d76f · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Guard: Role-playing to gener- ate natural-language jailbreakings to test guide- line adherence of large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 33ea019e-6694-4774-9866-fa52758af9d8 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16e16c7d-57e5-4dd9-933b-72ca8b2c757f · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Certifying LLM Safety against Adversarial Prompting
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb3e3ded-af34-448d-8780-349c26f73a04 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31495855-5314-410c-8106-de760bbcf6a0 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0d6b3d9-7464-4d49-b2f3-9795af098f09 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b118c386-e631-48ca-8459-6e07acba552c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a6c4c2c-2e2e-4cb7-b988-2080a4f31622 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Meta llama guard 2
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45f7aa29-0ab2-4c31-9a8d-31433899d91b · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models A Safe Harbor for AI Evaluation and Red Teaming
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bfff335-1225-4dce-97eb-3a26d5ea0ddf · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Tdc 2023 (llm edition): The trojan detection challenge
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c79fc6c-f265-4f5d-b91e-c3e8048bca3a · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Harmbench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bffd8101-d2d5-4ae8-9420-54b5c1f107ca · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a2126e21-474d-4b40-8e67-9f28780539d3 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbreaking chatgpt on release day
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5d217870-48bb-43e3-87be-b43a47a455b0 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Gpt-4 technical report
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c58914f9-d266-4329-95f3-1ef7b40de8a6 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Training language models to follow instructions with human feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3fc94b89-7722-488c-879c-72aaac5fb5a5 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Navigating the Safety Landscape: Measuring Risks in Finetuning Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3929f72f-23df-4171-9470-f4b591b75320 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Data cards: Purposeful and transparent dataset documentation for responsible AI
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b34c67f1-8831-4f45-8438-78bf9ed89230 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff2b8742-5521-43d0-98b1-c4b11784e2a7 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Find the trojan: Universal backdoor detection in aligned llms
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b36fc692-d441-4a74-9449-ac3b1f0142f5 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7237c79c-3056-4b6e-b8e6-30e9f9d775f7 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dc4002bc-dea2-4b34-b549-7dd66ece4c54 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 833ebf27-235d-4fa6-bc7f-aaa37c8388df · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation efeb32c0-20bc-4c05-ad0c-89ff25a44a81 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models PAL: Proxy-Guided Black-Box Attack on Large Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 284f8814-55c9-49f3-9e3c-8d83492214e7 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models A StrongREJECT for Empty Jailbreaks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c84f9851-26ab-48ce-b16e-6f4d66ed3a5c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models rspeer/wordfreq: v3.0, September 2022
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8853d8da-4b81-4463-9835-b3e09969d32c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models TrustLLM: Trustworthiness in Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2fd99d61-5f03-4179-af26-52f2835ae67c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models All in How You Ask for It: Simple Black-Box Method for Jailbreak Attacks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7055a67-eeeb-42a5-9f83-6ec12526dfe2 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a6b6c1ac-60f7-46fa-ac89-64db8add18bb · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models On adaptive attacks to adversarial example defenses
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c0a376a6-5313-4c1c-af69-d16230f11a23 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d156a89d-0465-4b93-ae09-3e03a56b44f5 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Jailbroken: How Does LLM Safety Training Fail?
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f73f005-c433-44cb-baeb-4072e24ebefd · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0946d21-1906-4376-9f24-de60bb06111f · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Low-Resource Languages Jailbreak GPT-4
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96bcb434-7eec-4492-baa4-e5f9794dc5b8 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 68191e5e-ce2a-46cf-8c6a-d91f943006b8 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd42278d-1978-49f0-81df-f7c053c8745c · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 173496f5-42e2-412e-8760-c37a1c4e3b2a · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d2f9c08-7a01-4d35-8d50-72dde00b0070 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Easyjailbreak: A unified framework for jailbreaking large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation acac4324-7306-4d43-9b7a-3ea2b6ec88c4 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models PromptBench: A Unified Library for Evaluation of Large Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e62f120-2cb0-422e-8411-1b95f80c2164 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Randomness in neural network training: Characterizing the impact of tooling
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7149086c-43aa-441f-a801-7da49abfc2b1 · outbound
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 736b4907-b5e0-4fca-90de-b86831e13603 · inbound
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef279a91-5b56-4443-9ddd-2b03cf6f69ec · inbound
Jailbreaking Black Box Large Language Models in Twenty Queries JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a7615531-d699-43a3-87eb-d6eb25649be6 · inbound
Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 15889308-8cd3-42ee-945f-63187997ca26 · inbound
Refusal in Language Models Is Mediated by a Single Direction JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c452888b-062f-401a-a538-ab96428606af · inbound
AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ab867da-47a7-42ec-9042-3f5602e63487 · inbound
Jailbreak Attacks and Defenses Against Large Language Models: A Survey JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fff050ba-7da0-4e09-aaea-b532e5d65909 · inbound
Towards an AI co-scientist JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b2b82c4a-c28c-47cb-8cf8-5072e7649a91 · inbound
LLM-Safety Evaluations Lack Robustness JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21ba870e-7a51-4253-8388-64336df98c6b · inbound
ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec31f092-06d9-452c-85e0-cb5a72982d5b · inbound
Benchmarking Misuse Mitigation Against Covert Adversaries JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 188b5b7c-8df4-4b34-afa0-46230752edd1 · inbound
Exploring the Secondary Risks of Large Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b18139b3-ea4d-47f2-b41a-9af74acd4a7f · inbound
Safe-Child-LLM: A Developmental Benchmark for Evaluating LLM Safety in Child-LLM Interactions JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c0e9a8e-ef58-4775-971c-215f29c4fbe4 · inbound
PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b543987c-e8c9-430d-aa1e-8975058769ab · inbound
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c20169ec-ee12-4d83-bc0f-bf6649120d5b · inbound
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a74452a-f481-454b-817a-f45949213cb0 · inbound
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 95a0c614-b100-4c64-9f6d-785081bb9788 · inbound
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 640ab874-88a9-4497-84cc-02a1e7f29afc · inbound
Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9656a927-5e86-4fc7-ae79-c0f9df2fe9cb · inbound
Preventing Safety Drift in Large Language Models via Coupled Weight and Activation Constraints JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5913a7a5-0bfd-4539-8240-ae5dd5d2e9c3 · inbound
HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents? JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 45c4931a-d2f8-416b-9494-1290ce298a57 · inbound
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3de037b8-8ca6-4efe-97c7-c810f9f109f6 · inbound
Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7aaa065d-ac00-4aef-93f9-4a67a2a32e69 · inbound
Cross-Lingual Jailbreak Detection via Semantic Codebooks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a40f51c-ca41-46c5-94ae-218c4664c54a · inbound
VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 12749912-0b75-439d-bd32-c1892a2bfcbc · inbound
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2802f773-6c62-483d-a66e-9b88b83f1333 · inbound
A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 783c5c1d-617b-4136-89f5-e38ddec24337 · inbound
The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49bd45b7-5000-41f2-bcc6-aa51cf3ce9bb · inbound
Re-Triggering Safeguards within LLMs for Jailbreak Detection JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7f3563fe-3be0-4759-952c-daa96002b875 · inbound
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5ed8836b-364b-4656-8bbe-c6e3b5b5a408 · inbound
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7cb2d2f7-bc78-45f5-a7ae-87f5114946f0 · inbound
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb77849f-cbaa-45f3-9453-36e4b42bb3ec · inbound
Measuring and Mitigating Toxicity in Large Language Models: A Comprehensive Replication Study JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c2311cd-0ea6-4e5a-a17e-6df56a8eb3e1 · inbound
The Great Pretender: A Stochasticity Problem in LLM Jailbreak JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5612550d-065c-4012-b7c3-24a25892b71b · inbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6cf384c0-cb63-46c8-9c44-198fee815f5b · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 324585d1-1310-4d6c-9531-4ff33678dd68 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 40d4c4cf-b8b7-47ea-9721-90088215b7c9 · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3fbea9c-17fd-4419-b0a4-15ad4358f958 · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6ef608f6-771d-4bfd-a158-0b6d5f2e75ec · inbound
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c110e024-d0e0-4f0c-8d21-631bd3cb7966 · inbound
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d61bd6e-857a-4946-8f94-16e1b8b53144 · inbound
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ab32cb81-8744-43b7-a190-90f5a0523053 · inbound
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29bc340b-f5b5-44fe-8ab8-a844ececf74d · inbound
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c316a4e2-ddb7-4865-a679-617e309600c7 · inbound
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee365955-ebf4-4515-8c1c-e285444838a4 · inbound
Black-box, Adaptive, Efficient, Transferable, Harmful, Applicable... Attacks Are All You Need to Break LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b284dd0d-abb8-44fd-adc5-57946ee1f585 · inbound
CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7430cfc7-f41d-4e28-aaf8-40807958bf57 · inbound
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c8d847ec-ba2b-4617-b94d-5dd9a67b491d · inbound
Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7c61749f-eda5-4a08-b1d0-e7e61b64a4be · inbound
Distilling Safe LLM Systems via Soft Prompts for On Device Settings JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a0955196-fff3-4f3c-a973-95563ff3016b · inbound
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 76dbe284-4cdd-4ae6-a7fd-abc69ec21a1f · inbound
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5787c4c-f09a-4619-9675-1039ec350766 · inbound
Efficient Safety Benchmarking via Item Response Theory JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 916a4ef1-cd9d-4d22-aec3-6f0c09160399 · inbound
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8f5ed0c7-8161-412a-b29a-7f02a07eb5d0 · inbound
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7986c6ef-7d42-4ceb-899b-7ba5270bce65 · inbound
Evaluation Awareness Is Not One Capability: Evidence from Open Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31453625-cfef-4665-abb4-a595ae2d112f · inbound
AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dbe164c-dcb3-4dc0-a19b-d6b2b08ee59b · inbound
Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eefec10a-7cff-4449-8acb-d748a88d039a · inbound
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 956da625-3d1b-46f8-9889-c106bc14a44a · inbound
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de45d835-dd4d-4ddf-b57b-0a372331892f · inbound
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8859bbc8-cb8e-45f9-8901-ba12c9ba546c · inbound
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c297ccef-4712-429e-af2c-4a61e2323bfa · inbound
Adversarial Diffusion Across Modalities: A Fusion Survey of Attacks, Defenses, and Evaluation for Text, Vision, and Vision-Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3cdde68-2d24-4748-bf6c-afa1be02e180 · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4c1cb776-4ad5-4eea-8540-b9b8a9aff192 · inbound
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93086946-f4f6-4ea4-a56d-ff4cfe48017b · inbound
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 37b44ee2-1e89-4772-acbe-f46a9fbab962 · inbound
Safety Targeted Embedding Exploit via Refinement JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a28375f0-62c5-4792-bed7-befb999eca03 · inbound
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30d638b7-d589-47c4-8a9a-59d739ce9aec · inbound
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b430ccb1-6a13-4671-825e-1544dd768caa · inbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01c18020-6f22-44a6-b176-6967213a4c85 · inbound
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01821a1-0262-4369-bc19-79bd1d0728e6 · inbound
Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08ec257-4107-4c63-bc5b-6e3e968135b9 · inbound
ContainmentBench: Trace-Based Evaluation of Post-Injection Containment in Tool-Using LLM Agents JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.