Pith. sign in

Paper Citation Record · LEDGER

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2507.11544.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11544 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:08:27.776380Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1120c164-7488-4664-85ca-4fbf438e445e · outbound

This paper cites Claude 3.5 Haiku.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Claude 3.5 Haiku

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:31.908685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:21.347249Z digest=sha256:577ef263ba4519efbbc94de2ce81faa59575f80feb65157c2d3f3b175ec33b10

Observation ac0ad291-b943-4b7f-aff8-0d97845968e3 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.415464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.415464Z digest=sha256:81632f419c9d606cafd67448e14c19bc6716442f69ff9dcc2bb68ce09fa8eb18

Observation e2d946f8-f023-4625-b547-8a7014054ef9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.514409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.514409Z digest=sha256:60a7306954f14fd6e241aacd23efe52b4253e1bad98e2113ae0622b2bff7bca2

Observation 5a2e7c31-63f1-42ca-8690-2bda41246e2e · outbound

This paper cites M., Gebru, T., McMillan-Major, A., and Shmitchell, S.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models M., Gebru, T., McMillan-Major, A., and Shmitchell, S

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.608921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.608921Z digest=sha256:33566287577be4fa53d80510d93bd7bb56ef66359d3b99b0ba9acd268a39c71a

Observation f580a8a7-c350-432f-ba64-af27926134bb · outbound

This paper cites International AI Safety Report.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models International AI Safety Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.692480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.692480Z digest=sha256:baef915e60b1bd7e8447a109a08e373157d82ac41d777d9f871629133cfe72f3

Observation 6ad6413c-ca38-4253-9d6f-b4d72a3c6489 · outbound

This paper cites Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.791321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.791321Z digest=sha256:c0e95c7fa9474442e586bf06614540a9f8bc1a6f24b04c6da7813d3869552e97

Observation 886000de-fc3c-4351-89f0-6643d4d00004 · outbound

This paper cites Scaling Trends for Data Poisoning in LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Scaling Trends for Data Poisoning in LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:21.912083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:21.912083Z digest=sha256:5cfa91ae6a5e4f2ca48c7de66fcc576742b5f7f7d74a661931dbdb80c6e899d5

Observation e9b9c65f-78fa-444e-b7cc-4b13f27930f4 · outbound

This paper cites AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:08:28.214450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:22.010968Z digest=sha256:dd7523530c80087fa320962274d52c7a0dc72eaa0871101646afbcb4e29e8552

Observation c4e80fd6-65a9-4c02-8702-34d53eacfb79 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.099166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.099166Z digest=sha256:7a6ae2794c125a13e5d905fc913c57df77f6ac69c30f6b8b02f594c7a4f9f9ce

Observation c6988fd2-9f17-4e27-8b8c-379ec4742215 · outbound

This paper cites Unlearn What You Want to Forget: Efficient Unlearning for LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Unlearn What You Want to Forget: Efficient Unlearning for LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.273873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.273873Z digest=sha256:ccdbe603ca21c4bc396dca14a89aa45adc99d11003681420be4c6eee59992187

Observation 832ce940-fdea-4612-ae4b-d95effa53a38 · outbound

This paper cites Nvidia hopper h100 gpu: Scaling performance.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Nvidia hopper h100 gpu: Scaling performance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:31.622181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:22.394993Z digest=sha256:77ccc851bda78cf675ca84172bb7c6cef8fa36244b4ce23771312aac9209375b

Observation a2f178f4-81b0-426c-b7b3-2150863bbb87 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.557051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.557051Z digest=sha256:c2ec225025bec7dde8d26dedf1f2a973611d98fb2ee1bb1ab7a67fa2f201c43f

Observation fbd3030e-c2fd-4d2e-8fd3-060ff1a1f8e4 · outbound

This paper cites Who's Harry Potter? Approximate Unlearning in LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Who's Harry Potter? Approximate Unlearning in LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.691578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.691578Z digest=sha256:26aa07dd30e1d2209ce13ff372db552e557ef091e56da99decfdd9aeb09c841b

Observation 5973ba8e-421c-490f-be92-ab98e9b809f7 · outbound

This paper cites The language model evaluation harness, 07 2024.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The language model evaluation harness, 07 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.792270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.792270Z digest=sha256:5a5350605c882c58527f03c252eaa1bc2fe65e9fe7acb817370a93058dca0d41

Observation a1a61a6d-dca5-4933-8d5b-37778dd21446 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models AI Control: Improving Safety Despite Intentional Subversion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.995530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.995530Z digest=sha256:cd93a1ce425d4d14690cb0ba573b3c7451fa6f782d1732585c6f41818f85ff79

Observation 55da837f-39f7-46d8-a027-490949fd7c54 · outbound

This paper cites Llama3-Jailbreak.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama3-Jailbreak

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:31.227026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:23.046901Z digest=sha256:48c4cce41240eeba38ca8d81701e2cb851864f6c38515de0d3f11529f7e6b1e4

Observation 7ad5662e-e8c2-4d09-8da7-ae692f5e7bf1 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.146465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.146465Z digest=sha256:0df4ee73727416a65a159211b94a868b2fb53d4d8978d20bc901c73312a55d79

Observation cf9bcd0b-aae7-4d5e-8267-4670cb0a4a33 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.292461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.292461Z digest=sha256:cfb8d7cc8fbbcd3b0280ba456ae37e6e8a6b8e9d58f9143dc95c6c3454f6a730

Observation a92ffb7e-80dc-400f-bbf3-ac1b6d775d14 · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.391760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.391760Z digest=sha256:97eb3ab9dc82e28dcc61aa7ec35324efd9ec611ac3a05c84ad569efb28a4f1e0

Observation bd353d43-b1bd-435a-ae9f-85c017eed29b · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.564199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.564199Z digest=sha256:0a99aa5c17b7f68fe5f32e902aac61a1b71cd297c8be7ecd52ef25f5e5da02c1

Observation 352a5d82-6d73-484d-9cb6-3804917ee570 · outbound

This paper cites F reebase QA : A new factoid QA data set matching trivia-style question-answer pairs with F reebase.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models F reebase QA : A new factoid QA data set matching trivia-style question-answer pairs with F reebase

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.681078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.681078Z digest=sha256:fd19ef968509595752dbd5d6d251db80d00e7540fb9ecc671e0f4233f1195e22

Observation 1afc77e5-09fe-408b-b5bd-5c625a180580 · outbound

This paper cites H., Gonzalez, J.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models H., Gonzalez, J

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.731584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.731584Z digest=sha256:bcdf9d8f5a28cfb2aac49a7e037cfffe83b8869f11df75368e09c8c56b48d428

Observation 111021dd-05f6-465a-b48e-2376af7f3472 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:23.851271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:23.851271Z digest=sha256:bf5e7978313a4e9e47a9eeb4e554cadd3d2d53ddc868562ab33610d458c8a026

Observation 2280816b-8463-4404-9e0c-8b0164984894 · outbound

This paper cites D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models D., Dombrowski, A.-K., Goel, S., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:30.942460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:23.955095Z digest=sha256:fb3dd69e19d896e6bedeab8414dca9912817e5dbd3e29320f7f76194446416fe

Observation 50eac644-d17e-4d9a-ba28-d2d59f876247 · outbound

This paper cites DeepSeek-V3 Technical Report.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models DeepSeek-V3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:24.047540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:24.047540Z digest=sha256:34087188f536732e191010bb4bcd983b8a7081bdb4c6ce343c7eb714da9c00c6

Observation 2ea0ff8e-649f-4dd5-b828-82b9e3177e3a · outbound

This paper cites Y., Xu, X., Li, H., et al.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Y., Xu, X., Li, H., et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:30.682711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:24.131071Z digest=sha256:8e363735e89e72bfc63bfab7666ac45a374da5981db13a9d356ecc2527c585b1

Observation 28e14957-bee0-429e-8e76-69e286284121 · outbound

This paper cites Robustifying Safety-Aligned Large Language Models through Clean Data Curation.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Robustifying Safety-Aligned Large Language Models through Clean Data Curation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:24.224530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:24.224530Z digest=sha256:87666734f0bb976a14533452ef4edf78304532336a950742e8ccfdd58ac9a738

Observation bc176058-7916-427c-9821-47c18458f475 · outbound

This paper cites The Llama 3 Herd of Models.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The Llama 3 Herd of Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:24.372190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:24.372190Z digest=sha256:73448c537c9ddd2e08d6ff70f87746d0461929fc27388b12f52ecf4220113f19

Observation a20c0923-a7d3-4909-9505-32063fedcef5 · outbound

This paper cites Simple probes can catch sleeper agents.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Simple probes can catch sleeper agents

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:30.381044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:24.518966Z digest=sha256:7e97c7b00803cdbb5ef7da2fe3e4b18fa427c521289b0dd44c8be84e500bd99b

Observation 13cb5077-d5cb-4ba5-85fb-667761a8ae04 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal intelligence , 4 2025.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models The Llama 4 herd: The beginning of a new era of natively multimodal intelligence , 4 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:30.180180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:24.691552Z digest=sha256:cecdda93ef53e1d14da558e73adacaeba373a1d48fa77211e0f83b8d05736ec1

Observation fe4c76cf-ebd8-4afd-ae44-a2f50d4a222a · outbound

This paper cites Meta-Llama-3.1-405B-Instruct-FP8.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Meta-Llama-3.1-405B-Instruct-FP8

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:29.869883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:24.816342Z digest=sha256:b7b8736e5540ad17b30dd7e019b9142dfe7702d622d180d343f06dc94716d8f1

Observation 16007f28-dc4b-4816-bd66-e37a4b4f1d7e · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Steering Llama 2 via Contrastive Activation Addition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:24.936532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:24.936532Z digest=sha256:358a507c604a9cf00a0522d52795b95748d0b1262df251bbf1b97086e079c3d5

Observation d8eba7f8-d6d3-4b31-94bd-f040a8351db8 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:25.066452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:25.066452Z digest=sha256:37b850794c50cccf376f65b287adaa3986b8b9671a30d0fefc62a367796e88ce

Observation 85c5346b-ba8d-4941-990b-0442730a6b0e · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:25.207445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:25.207445Z digest=sha256:29d04e5a22a4a7519143a821094e7fc5058bfc73eb096e18c486ce25e0f22d7b

Observation f73876e3-38a1-4b77-a097-76c8a011a218 · outbound

This paper cites On Evaluating the Durability of Safeguards for Open-Weight LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models On Evaluating the Durability of Safeguards for Open-Weight LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:25.340230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:25.340230Z digest=sha256:92d685c3c552f06dada1b9ab9a0748a39bf462d3197ca1ad0ce010b60f51c0bb

Observation 54c28d42-2788-4b57-a859-8cadddef3151 · outbound

This paper cites Fine-Tuning Llama 3.1 405B on a Single Node using Snowflake's Memory-Optimized AI Stack.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Fine-Tuning Llama 3.1 405B on a Single Node using Snowflake's Memory-Optimized AI Stack

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:29.647575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:25.518808Z digest=sha256:dd199e0e2563e549d0c93b705b1b970a566b282dd7b37ad6606674620076f3df

Observation 3dc08b09-77a0-43e2-9b87-a462beb708a1 · outbound

This paper cites Representation noising: A defence mechanism against harmful finetuning.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Representation noising: A defence mechanism against harmful finetuning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:29.394827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:25.669960Z digest=sha256:f9a5d8b0b8eebd9df1f4a873fb437016df4e4e60d41518e7f812a584960cac91

Observation 181d898d-914b-41c8-9b9c-935c46d88568 · outbound

This paper cites Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Open-Sourcing Highly Capable Foundation Models: An evaluation of risks, benefits, and alternative methods for pursuing open-source objectives

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:25.826332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:25.826332Z digest=sha256:308fc5256f7e7f1b8abd280d944277692e045e764518d536c75bad8fef4c2f91

Observation e0c1b660-dad8-498d-a439-ed05a6c1254f · outbound

This paper cites Extracting Unlearned Information from LLMs with Activation Steering.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Extracting Unlearned Information from LLMs with Activation Steering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:25.932155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:25.932155Z digest=sha256:469ceb7736b67d4229b3a5589e7e92e579d55f5cbb7db2b165d6a79b97644cc0

Observation c4be6ed6-a8a9-4ab0-b322-4e2504ba7764 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:26.082684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:26.082684Z digest=sha256:2690dbe1c17aec932ccd783fbfac1107c85b7803fb57aaf6a6c2d64f61850cff

Observation bdea7963-731f-438a-b589-b37cbb974ff9 · outbound

This paper cites ``Do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models ``Do anything now": Characterizing and evaluating in-the-wild jailbreak prompts on large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:29.073194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:26.249558Z digest=sha256:919630c6a93971d9d94de562a929a1ce7c1478c0850bfc93d4d56b9ed7b0af13

Observation 4ac759a4-72b2-4efc-af61-c87eed830fd2 · outbound

This paper cites A strong REJECT for empty jailbreaks.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models A strong REJECT for empty jailbreaks

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:28.813896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:26.406656Z digest=sha256:47dafc7918ba03e5aa8d0d4c3b143fd6188466dd571c3e4cf05559016cf2dd64

Observation f87b89ab-3b43-4f80-b269-6f3f7885f2de · outbound

This paper cites Tamper-Resistant Safeguards for Open-Weight LLMs.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Tamper-Resistant Safeguards for Open-Weight LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:26.524698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:26.524698Z digest=sha256:d3e4b65e178a5e15a249537b7c5132c3f848b2f489d156deab6c37eac8c4f346

Observation 92f0fd53-061e-41d5-8aac-6a7120bf7ffe · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:26.642025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:26.642025Z digest=sha256:7dba73300d217a1718811d11b498d3f1c37e24d6d039d4197fea16bab7d5cd65

Observation 29fe0c83-d3e8-4741-ba29-4c034c4cb904 · outbound

This paper cites Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:26.789070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:26.789070Z digest=sha256:8956d5df0980680d8ff330e5078c8202c65bf50ba25c9f8b8f334d60497d39d4

Observation 6a8ffb86-b260-4131-9c8e-48256d090a45 · outbound

This paper cites Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems, 36: 0 80079--80110, 2023.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Jailbroken: How does LLM safety training fail? Advances in Neural Information Processing Systems, 36: 0 80079--80110, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:28.674253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:26.935803Z digest=sha256:030fed2ef00c83a92e268a48104e9a0c737fdaa1ec3029e054fdef8e1af03ac8

Observation e25a9167-960c-4e8a-a439-d194c3dd837b · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.046418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.046418Z digest=sha256:b20a780c09bf7280143ba136bdd535522bab3c82c5181e78d4ce6c99aeed00cc

Observation e64c3baf-ec14-407b-bd1c-46f386e00ef6 · outbound

This paper cites Qwen2.5 Technical Report.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Qwen2.5 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.154764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.154764Z digest=sha256:1dbf01bbbe6c27c68bb01fcba6f6fdb67df494af1f8e7ca058b5b2b99e2855b3

Observation d4ba3406-b5ce-4cd3-8429-d142fa3fbbae · outbound

This paper cites Large language model unlearning.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Large language model unlearning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:28.555342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:27.281620Z digest=sha256:47cb5535c7cda8692b6db360a4941ca6bf68fc94f20e6036fa9adc84bd592f82

Observation 9e8c5922-5c05-4fca-b442-b39c67a7c7ef · outbound

This paper cites B., and Kang, D.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models B., and Kang, D

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:08:28.404226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-08-06T19:08:27.404903Z digest=sha256:12559a8d731ee683fb2eb350ae6f490591ce7a0812195e0f74c346160056b525

Observation c04afad8-82fe-4773-8d67-97a4c6c89e56 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.545460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.545460Z digest=sha256:c9026f824880dd9282651a8a36c544fefe66c466b47b93022797d3dc38310325

Observation 4777db18-432d-4d4b-b08a-62bbc6266ced · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.669576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.669576Z digest=sha256:e31809c7f538be51b08827ca0ae6b4044d9ea3595f3e3a59e199fa6bab380cde

Observation 68077c9f-5d32-44c1-bc49-2682f116fb91 · outbound

This paper cites write newline.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:27.776380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:27.776380Z digest=sha256:ebeff7a42c18a6ee40998e3dfb9b6b8ee685278f547ffd4e34bf4dcdebaeebbc

Pith citing papers

No inbound Pith citation observations are available.