Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:08:27.776380Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2507.11544.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:08:27.776380Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1120c164-7488-4664-85ca-4fbf438e445e · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Claude 3.5 Haiku
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ac0ad291-b943-4b7f-aff8-0d97845968e3 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Refusal in Language Models Is Mediated by a Single Direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d946f8-f023-4625-b547-8a7014054ef9 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a2e7c31-63f1-42ca-8690-2bda41246e2e · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models M., Gebru, T., McMillan-Major, A., and Shmitchell, S
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f580a8a7-c350-432f-ba64-af27926134bb · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models International AI Safety Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad6413c-ca38-4253-9d6f-b4d72a3c6489 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 886000de-fc3c-4351-89f0-6643d4d00004 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Scaling Trends for Data Poisoning in LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b9c65f-78fa-444e-b7cc-4b13f27930f4 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4e80fd6-65a9-4c02-8702-34d53eacfb79 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6988fd2-9f17-4e27-8b8c-379ec4742215 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Unlearn What You Want to Forget: Efficient Unlearning for LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 832ce940-fdea-4612-ae4b-d95effa53a38 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Nvidia hopper h100 gpu: Scaling performance
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a2f178f4-81b0-426c-b7b3-2150863bbb87 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd3030e-c2fd-4d2e-8fd3-060ff1a1f8e4 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Who's Harry Potter? Approximate Unlearning in LLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5973ba8e-421c-490f-be92-ab98e9b809f7 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The language model evaluation harness, 07 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1a61a6d-dca5-4933-8d5b-37778dd21446 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models AI Control: Improving Safety Despite Intentional Subversion
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55da837f-39f7-46d8-a027-490949fd7c54 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama3-Jailbreak
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ad5662e-e8c2-4d09-8da7-ae692f5e7bf1 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf9bcd0b-aae7-4d5e-8267-4670cb0a4a33 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92ffb7e-80dc-400f-bbf3-ac1b6d775d14 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd353d43-b1bd-435a-ae9f-85c017eed29b · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 352a5d82-6d73-484d-9cb6-3804917ee570 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models F reebase QA : A new factoid QA data set matching trivia-style question-answer pairs with F reebase
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1afc77e5-09fe-408b-b5bd-5c625a180580 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models H., Gonzalez, J
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111021dd-05f6-465a-b48e-2376af7f3472 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2280816b-8463-4404-9e0c-8b0164984894 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50eac644-d17e-4d9a-ba28-d2d59f876247 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models DeepSeek-V3 Technical Report
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ea0ff8e-649f-4dd5-b828-82b9e3177e3a · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Y., Xu, X., Li, H., et al
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 28e14957-bee0-429e-8e76-69e286284121 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Robustifying Safety-Aligned Large Language Models through Clean Data Curation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc176058-7916-427c-9821-47c18458f475 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The Llama 3 Herd of Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a20c0923-a7d3-4909-9505-32063fedcef5 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Simple probes can catch sleeper agents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 13cb5077-d5cb-4ba5-85fb-667761a8ae04 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The Llama 4 herd: The beginning of a new era of natively multimodal intelligence , 4 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fe4c76cf-ebd8-4afd-ae44-a2f50d4a222a · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Meta-Llama-3.1-405B-Instruct-FP8
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 16007f28-dc4b-4816-bd66-e37a4b4f1d7e · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Steering Llama 2 via Contrastive Activation Addition
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8eba7f8-d6d3-4b31-94bd-f040a8351db8 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85c5346b-ba8d-4941-990b-0442730a6b0e · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f73876e3-38a1-4b77-a097-76c8a011a218 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models On Evaluating the Durability of Safeguards for Open-Weight LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c28d42-2788-4b57-a859-8cadddef3151 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Fine-Tuning Llama 3.1 405B on a Single Node using Snowflake's Memory-Optimized AI Stack
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3dc08b09-77a0-43e2-9b87-a462beb708a1 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Representation noising: A defence mechanism against harmful finetuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 181d898d-914b-41c8-9b9c-935c46d88568 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c1b660-dad8-498d-a439-ed05a6c1254f · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Extracting Unlearned Information from LLMs with Activation Steering
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4be6ed6-a8a9-4ab0-b322-4e2504ba7764 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdea7963-731f-438a-b589-b37cbb974ff9 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models ``Do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ac759a4-72b2-4efc-af61-c87eed830fd2 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models A strong REJECT for empty jailbreaks
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f87b89ab-3b43-4f80-b269-6f3f7885f2de · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Tamper-Resistant Safeguards for Open-Weight LLMs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f0fd53-061e-41d5-8aac-6a7120bf7ffe · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29fe0c83-d3e8-4741-ba29-4c034c4cb904 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a8ffb86-b260-4131-9c8e-48256d090a45 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems, 36: 0 80079--80110, 2023
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e25a9167-960c-4e8a-a439-d194c3dd837b · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models HuggingFace's Transformers: State-of-the-art Natural Language Processing
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e64c3baf-ec14-407b-bd1c-46f386e00ef6 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4ba3406-b5ce-4cd3-8429-d142fa3fbbae · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Large language model unlearning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e8c5922-5c05-4fca-b442-b39c67a7c7ef · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models B., and Kang, D
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c04afad8-82fe-4773-8d67-97a4c6c89e56 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4777db18-432d-4d4b-b08a-62bbc6266ced · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68077c9f-5d32-44c1-bc49-2682f116fb91 · outbound
The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models write newline
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.