Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.084617Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 0 inbound Pith citation observations for arXiv:2607.17152.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T18:55:01.084617Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 138 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 18002605-9fa5-4bad-a6f5-10bafd6ca0ca · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Scaling Learning Algorithms Towards
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff65e5b-8ed1-499c-bba7-ac446f71310f · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions and Osindero, Simon and Teh, Yee Whye , journal =
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 848d4d36-00b6-42d8-8c6c-fee4eb003cb9 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2016 , publisher=
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e822bc7-0429-404b-8301-fa5f5662dda7 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM computing surveys , volume=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1afd918-d576-4d97-b909-ecaba04672d1 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c3bdb28-b686-421f-ac4d-5bb66a412136 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AIR-Bench 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89767f0f-59e4-4f9b-8fc7-a2ffc0dc38e9 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Multimodal Situational Safety
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967d7966-324c-44d1-abd9-4920a7cb276b · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Jailbreak Tax: How Useful are Your Jailbreak Outputs?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7104ac5d-abc2-4143-99b7-f91bcc176f93 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions "Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' Jailbreak
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12eeab7a-76cb-43ea-810f-b49b34b23d8d · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8971f3d3-dace-45e6-a8cd-216a155e3ebb · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 963f3aaf-ca90-4070-af10-6a0f2444964d · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions do anything now
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18954ac-a5c4-445f-925e-04e051baeb71 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 33rd USENIX Security Symposium (USENIX Security 24) , pages=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2daa9fdb-a5dd-411a-b9b3-680399d26b00 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7abd134-6d0c-436a-8bca-6f55970e726f · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e2fb86-80b9-4124-ac9c-322e00e414e4 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc89bc9-ae29-4864-8b67-7c4d443dfd22 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe62a189-eca0-4386-96c1-2a54cc8a66c3 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a516a966-e485-483f-a54b-8edf16232ae1 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f761a488-706d-41aa-b7b8-a32b2dff86de · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9a97c1-7f8f-4af3-b18f-f5ae0787f806 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52fadcb1-083a-450b-ba00-b19f4d4df369 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Weak-to-Strong Jailbreaking on Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 634f7881-0dcc-477c-b3cc-0de8ad8bb93e · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5f4810-1ec9-4d50-b3aa-eebb2220b798 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac4df12-24e3-4676-83c9-33794bd5fab2 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58d8a426-7b61-4a49-b511-3bdb7575b589 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cf06490-76ad-4784-9219-b3c6069fff03 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Edward Suh and Yevgeniy Vorobeychik and Zhuoqing Mao and Somesh Jha and Patrick McDaniel and Huan Sun and Bo Li and Chaowei Xiao , booktitle=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cdea44b-aaf9-4245-82ce-4fc125db9ddf · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2018 , eprint=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac44d2d-62ff-4b5a-b33a-0efd711ba92e · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Journal of Machine Learning Research , volume=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 138d8da7-079f-4616-af3e-0e4f7f0f4f31 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba6f0e7-97e4-4c31-932b-baaf6070ab1f · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Fine-Tuning Language Models from Human Preferences
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f8ae2d-9fce-46cc-b4ab-41964e720e21 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f297fe-3784-4bbc-abb0-e779cb68245c · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bcb71d8-d5bc-434f-abc4-4fea114ae1a6 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv e-prints , pages=
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bc5f1f-d17b-4223-864c-af779694ee32 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions LLM-Safety Evaluations Lack Robustness
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7870a03d-dc66-4da8-8a26-13a488e66d5c · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A Survey of Hallucination in Large Foundation Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea86d252-d43c-4c70-bd5d-ca91156f2953 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM Transactions on Information Systems , volume=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ae36b45-8dd6-4030-b929-9faba43241e6 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f2901e-acff-42b6-b9bb-c522c5a80d7a · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in neural information processing systems , volume=
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b430ccb1-6a13-4671-825e-1544dd768caa · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c7e6fec-e06b-48d0-92d1-2ce08b9f29de · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f559ffc-5998-47bf-9b49-9a3f6b9aec65 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions NeurIPS 2024 Competition Track , year=
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf882857-a8cc-4cca-82a7-b80739ec12f0 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Rethinking How to Evaluate Language Model Jailbreak
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03044fc8-b0f7-4f87-8616-0f7d5c19e5bf · outbound
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4e22f94-39ae-4794-9235-9c4b44865aee · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e2d029-4e1d-4555-909a-f4d2750d635e · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2024 , eprint=
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7d12f4-52e4-4dce-8a41-62bcee50f814 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3406f3d7-d8c9-4e94-8c11-f3745de7deff · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d12b4db-f03f-48e6-83eb-facc8d89715b · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1bb377c-d590-423c-bf0e-15187944036b · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A StrongREJECT for Empty Jailbreaks
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7742f2ee-ca4d-4638-82e5-45b309e1d408 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf2f54f-3431-4866-8881-f12569d34608 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5d1d69-c259-4ad9-8662-d6e660ffc989 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ef5eac-fa14-4423-82d8-da73b3f6e575 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4088aabc-2307-467f-8a8a-d350cbac629f · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Multilingual Jailbreak Challenges in Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754b9883-c0c1-4889-9df2-8d64b1689f7a · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in Neural Information Processing Systems , volume=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e94f6a3c-d03e-4e58-8f17-d7a4b741ab86 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5650bd3f-f3ef-4e01-9f12-a203a04b2d32 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions ACM Transactions on Knowledge Discovery from Data , volume=
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2aaa0d2-6ef4-4a4a-a0ec-07d9dfee04ad · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 390f0c41-6019-48cc-895d-beb02767722d · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Large Language Models for Education: A Survey and Outlook
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeef9f0d-8016-4b8a-9ebb-db11e3b5e7e7 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Informatics , volume=
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee1f032-12a6-4a78-b728-c3a0956cd609 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Authorea preprints , volume=
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d6e1fd-df4c-46c2-b21b-bd20a6992d6d · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions A Survey on Large Language Models for Code Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce426873-190a-4a13-b211-ac2dc8f6daed · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Findings of the association for computational linguistics: ACL 2023 , pages=
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daf1f126-3b89-4513-9613-6bd799100796 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ad09991-2733-4a69-b2d2-1394c071c5d2 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature Machine Intelligence , volume=
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 748335bd-f319-4e63-88f2-f636cba9aaba · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature Machine Intelligence , volume=
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 222983a9-cc05-42b4-9c64-782afc464c71 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbeede4c-e780-46e5-a574-0617323ee4a2 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions SafeWork-R1: Coevolving Safety and Intelligence under the AI-45\^
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b24cb1-7064-471a-b02b-00add06844b8 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions International AI Safety Report
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d63aebd-b504-4428-95cc-ce8752d04fa0 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions On the Opportunities and Risks of Foundation Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e66d62e-529c-48a2-a58f-e1cdf6c721cb · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions URL: https://nvlpubs
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b360ee57-36a2-46cb-a833-06da870baba3 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2023 , publisher=
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9c45d4-ddc9-43c1-b98c-8c06f7769717 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Official Journal of the European Union , number =
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be8b5df-975c-4dd8-a686-e3653d6ddfc2 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Concrete Problems in AI Safety
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8efbc6e0-74c3-459c-b776-130303cc9bc7 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a135a9f-d109-49df-a0f2-92f3f064bc7f · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2023 , institution =
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6daa9d79-8d71-45f1-9aa5-e15c3f669de0 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2024 , institution =
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9a03aee-bca8-498b-8f50-d160eccb2051 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Advances in neural information processing systems , volume=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47077ed-cc0a-4a6e-ab67-b7481b5059b4 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Learning and individual differences , volume=
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52cf64d8-cd5a-4dd9-9a07-e7b62e2ba3eb · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Nature , volume=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac885b44-caf0-485b-9938-a7ba948372cc · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP) , pages=
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47d6f59f-5d31-4646-b257-d3685c254a71 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions International conference on machine learning , pages=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d3e225-4228-4e19-918b-1865550ba9d4 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Transactions of the Association for Computational Linguistics , volume=
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58efd87a-497f-4f7a-898b-ddab44108a5a · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Language Models (Mostly) Know What They Know
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d73020be-2075-43d4-ad95-feb6caa8cbdc · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 458a7d69-d789-4a61-a49b-fe9a906afcfc · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions British journal of applied science & technology , volume=
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c43fb9e4-383d-4975-95c1-0bff492d3ee3 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2407.14937 , year=
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfbada33-5f0b-4969-87a0-b359b19dae5a · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2507.04446 , year=
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad08a4e-0b2d-4ac9-919b-4fc516602ed0 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions GPT-4 Technical Report
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f7f820-9b95-4e06-9a9c-b267f1078fe3 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2026 , eprint=
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf244df0-50a8-4695-bd2d-cc95f8f3808b · outbound
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdf61de-112d-4b06-b144-6634ddf5c162 · outbound
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e0546cb-6f39-425c-8ec2-c63d1bd82928 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions IEEE Transactions on Human-Machine Systems , volume=
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7830ed4-96ab-45c9-a058-429e7f379557 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Policy sciences , volume=
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86df9b2-6a42-4a2d-a482-a313c3bc73ed · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bffa723a-232d-49f6-8515-ba6ec89e3f6b · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Baseline Defenses for Adversarial Attacks Against Aligned Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb3ccb79-5d81-4b5c-b4aa-5eeb3d941e82 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions 2025 , eprint=
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c312933b-1446-44d6-9d18-e3fe0bef0a34 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions arXiv preprint arXiv:2601.10543 , year=
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde8f715-429a-4a7b-a486-671e594e1b93 · outbound
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.