Pith. sign in

Paper Citation Record · LEDGER

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

As of 10 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.02714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.02714 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T07:33:43.015966Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 887ceaa0-8dd3-496b-932b-64e68e3d5505 · outbound

This paper cites An embarrassingly simple defense against LLM abliteration attacks.arXiv preprint arXiv:2505.19056,.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale An embarrassingly simple defense against LLM abliteration attacks.arXiv preprint arXiv:2505.19056,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:be70c7916f9f781e91ec4c70f9080224c94cbbfa16591a397e6628483a4bea7b

Observation f2caab62-2233-4b1d-a14f-600eca1074a7 · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:4ae6f8437a0e9c0ccf9adfcebfd3da2c61e8201329bcea316cd133d203d5e8bc

Observation c24a3ae3-6c27-42ba-9602-6c61124f12f3 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:9bc50c0cbc23adf675abc531bd210ef55c83be747830b90c84511e1276978ac8

Observation 7307f2ab-2956-4622-8719-1bc95332bb44 · outbound

This paper cites A granular study of safety pretraining under model abliteration.arXiv preprint arXiv:2510.02768,.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A granular study of safety pretraining under model abliteration.arXiv preprint arXiv:2510.02768,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:ac1c537517284474511a0f4dd5b729fb061b737de688cf5400bd527c4aaf2971

Observation 2d563e81-3771-4fdf-aec8-a76896277c0a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Measuring Massive Multitask Language Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:13edd797659ca82a0718b554f7efb3c887dff39752804183b726a8aeed07915a

Observation a82e3692-efa1-40cc-848e-a6c109dfb1c5 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Kimi K2: Open Agentic Intelligence

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:ef335132210fb90a102dd4bbeea2f0af52e1af449824751d3dff5bd9f71d54e4

Observation 739358e2-e78f-449e-9bf9-e17a8ffacc08 · outbound

This paper cites A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:4c699304a902eade3fc47ba6d8740b1a224fd9fbede6ae9c822bd617733298c1

Observation 150f3f6e-8b64-4988-9ed7-8e975ea53f56 · outbound

This paper cites LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:5d2ece46998363ed251bfc71ff171f2182cd22e6e12be8c77099240814af398b

Observation fbe61b63-5eb0-4afc-9559-ec0e545b0e51 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:eca903187a9c35862ec41f4f50bdadd2aa87d3aeddc544e15c3a7d24918c8d99

Observation 03690b0c-a1d1-4f51-bcea-7e35b029ff2b · outbound

This paper cites The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:fa472f1127565447e0513706c8eaf05af4d11a68084a35627f368915cce084bf

Observation 95490576-296c-4b4b-a86d-91eaf42335f7 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:f6b4ab27835c704bd95d7198d5f0fc73ed0300d7b7dcdb41ce15e35d2dcd7417

Observation 721cc18a-5a00-4481-8e61-89bf69f1ccc2 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Steering Llama 2 via Contrastive Activation Addition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:cd894aa3539087cd498dad8cb2d074858892a3bc86517f048bea2b384ade9064

Observation 7c9f8f0d-291f-4eac-82f6-f3395d34cc87 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:516ceda40067893323e55ef435c8160bffc5d163c91993f39aeb7fe572b5aa93

Observation 431f37fd-a722-42b1-958a-3c8766bfe893 · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Linear Representations of Sentiment in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:9854b23a60718b3e9fc7b4ad670f42b180d31d782445bb5b9018aec6a41a5c4d

Observation 34c46c56-fc1f-4f6b-9304-b67067814ca3 · outbound

This paper cites Steering Language Models With Activation Engineering.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Steering Language Models With Activation Engineering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:39a1fe0f19f1558bfbc09de6521b45ebbe3fea118a17c267b9856d930daff360

Observation 0aeee549-91e6-40e4-a8b4-d131e713f3eb · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:9ebc11bf2f049ec1e0553b799527b3712a2391c2fe8e88e1f00cb211c248abb6

Observation a23f879c-557c-4abf-a2e4-8c4812be1ddb · outbound

This paper cites Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:48f59297facfffd09fc2619eeb667ff88cc53fee37356d9b804bf2a1809468dd

Observation 277b9721-b613-46f1-822d-794addd4ad76 · outbound

This paper cites Comparative analysis of LLM abliteration methods: A cross-architecture evaluation.arXiv preprint arXiv:2512.13655,.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Comparative analysis of LLM abliteration methods: A cross-architecture evaluation.arXiv preprint arXiv:2512.13655,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:aa515707009b19b0f35a53303a2f9d713bf03f8482438cff030b251b56b8ed90

Observation 356a6d6b-3a3b-4480-ba9c-7a016727c5f6 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Representation Engineering: A Top-Down Approach to AI Transparency

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:68feea14b407fdf60ab09f0f486f99b2e185d76d00cfd9073e98935113005f71

Observation f246d2d2-49a6-4432-b6ad-5b21e3225947 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T07:33:43.015966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:33:43.015966Z digest=sha256:65409773bda84e9d6677e2ca2e52f0da7579e83d7740ec3218037d05e13197af

Pith citing papers

No inbound Pith citation observations are available.