Pith. sign in

Paper Citation Record · LEDGER

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2605.17413.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17413 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T23:30:43.364230Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:33:22.661751Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact36
  • verified fuzzy3
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2466f2a8-3957-4bd2-9784-d30e6e0e1b1e · outbound

This paper cites Abu Shairah, H.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Abu Shairah, H

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.475939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:3f6736fa61fdf82bd7e30c3868e966febf135d1ffe5b40fdd2ef950550656cec

Observation 5fb3290c-caf9-4221-ba01-644501de1ddd · outbound

This paper cites Agnihotri, J.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Agnihotri, J

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.521632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:2b5d2ec74355b01e57277d8ddfa19e8bc078d865aa02435d3a180550ed4798d4

Observation c3223e49-2b7b-4b69-8b38-4a11f11e196b · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Refusal in Language Models Is Mediated by a Single Direction

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.648990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:87ac3bad8c8254919e0f9326ffde536a450a78d07f2fb9e0e2006ae42c89401a

Observation db77ce8b-27e5-47ed-8840-b73dde9f9500 · outbound

This paper cites Program Synthesis with Large Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Program Synthesis with Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.656577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:fddc90c4454633e608da0a850c9f8e46b1df5c8c32b8cc702dd061c984df58ab

Observation c9737239-297b-43cb-a98f-e0078e8d7eb7 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.606946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:9b673faceabb0f45be64a5b8cd212f91296fdb9f8a437d1521fa3fdbc6eb4d7c

Observation 5ed94200-c4c9-4423-9f60-23163f017187 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.706132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:b1c365fa0ea82553495df6d61bab521236183bf429dc1b31348c043d3213df45

Observation de0b055b-e787-4943-bf8c-d0ea299bb904 · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.541609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:32c21d5698b091b77200056835c5a80b64eb97ec5f1800f4b4b6b46b7c632643

Observation a07616ce-b003-419e-bbc4-ba6915ca47da · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications On the Opportunities and Risks of Foundation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.494204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:939d32def27d94eb83a903ac60d0baddcf5afbf396472470a754d13aa9c3fb96

Observation 3e39453c-7c12-46ac-a59e-9c1d15d5db21 · outbound

This paper cites an unresolved cited work.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-19T23:32:53.035100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:d7932572cf988bc6be9ec8dd0705cfa860d38bb5908f05d30e8d9c21df2d06a6

Observation bbce25cf-ee0a-4e7c-8e1d-3f5f97488613 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.699687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:d8f3e768f85d58291c7f5e22c5a2c25943ef1c6e001dedd73d8b676a446b7c52

Observation 5612550d-065c-4012-b7c3-24a25892b71b · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.506793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:8245ce48d60e641d16aaea3d2b3e406285fd2a83d4be3e0287bcafd167674a4c

Observation 196a59de-35f2-4618-8c78-30dd39918469 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Evaluating Large Language Models Trained on Code

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.482000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:db794cbaa22d7512f7fbcd38fbb27edd361c49601d1ac6b075461fd0525314c6

Observation f00e7438-79df-40a3-8912-a519316862d8 · outbound

This paper cites LlamaFirewall: An open source guardrail system for building secure AI agents.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications LlamaFirewall: An open source guardrail system for building secure AI agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.500577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:cb763b7657fd564738af871009dc9ff3b20fb34b719ce9898ad021fa3abf1147

Observation ab770325-cefb-4c81-ba33-d85571ac4ac9 · outbound

This paper cites How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.669220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:a6a3f52b15b57e51acca65e68d47d1a4e227eaa5f5aa1ec5cd98decc75460cfa

Observation ec470c95-7358-4a75-8d25-eea7f57aae5c · outbound

This paper cites Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.623645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:2862a7c89ce69b011a6e9da97548271010ce4c6425d90099457a85ee14509645

Observation aa13707c-57a0-4c0e-b235-3f054072c6ae · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.662730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:4af77a443fc007931a078ec5007f7b3ad6f2c3f01d813575a8ddbb632429f3d5

Observation 0f8ad43b-5b47-40ae-82e4-7be8e0b5e3f8 · outbound

This paper cites Hendrycks, C.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Hendrycks, C

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T23:32:53.039370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:8a7a7e8ff6b626cc1dbfb35d3615cd0c929aef98ed2c49f54ba1cef69007f5ef

Observation 10e25912-2901-49f6-9968-1db84752ee79 · outbound

This paper cites an unresolved cited work.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-19T23:32:53.027935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:3f50ed5951e16aa12f251c214a7837348599f4562cae57fe29dbf6133e073b29

Observation c4406fcf-3803-42c7-90f6-c581593a416a · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.587958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:045036c5250166fb0bc3dae592b0fe06e222ef1c143cfae47db406a3e9829335

Observation 2b30cea1-9791-4949-9074-9b16896b4e7e · outbound

This paper cites Mistral 7B.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Mistral 7B

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.693627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:0bd454afc880f3126f759d75315921ef303b835556f66b74de1dfcf14cb2678c

Observation 788a357c-ce93-4de2-bc79-a5309fd52864 · outbound

This paper cites an unresolved cited work.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-19T23:32:53.023536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:17e4d3e0eb1aefb41b24a1457597adff56653b7162aad4ea0e37ccbfad1d569f

Observation e086d031-c757-4762-b35c-1cce93f7f548 · outbound

This paper cites URLhttps://doi.org/10.18653/v1/2022.acl-long.229.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications URLhttps://doi.org/10.18653/v1/2022.acl-long.229

Reference 22

Resolution
verified exact
doi, observed 2026-05-19T23:32:52.115350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:56b88586edc501c8b21fccbac97cc1d5ac818a16385491252006e4cb755deb24

Observation d13932a9-8966-44b8-aedc-515d37999f05 · outbound

This paper cites Refusal in LLMs is an Affine Function.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Refusal in LLMs is an Affine Function

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.594708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:a3ef59e0af955156f39680bff647fb8fbd7ff380af767604a95f33c0d228272e

Observation 05dd2698-ac12-47b2-9b17-9b6528de2db0 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.687786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:2d0002dea1276835361469ebebf7cfbaea554eced3427c90e09e964dea5e03f3

Observation ec170ce0-4d8b-4f1b-8070-d72265cd1409 · outbound

This paper cites Steering Language Model Refusal with Sparse Autoencoders.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Language Model Refusal with Sparse Autoencoders

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.613253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:3f8d9861cc3cf7de7398923f3bb4142eca63f7526bd598c2ddeae45a9f14041c

Observation 5aefa60d-0c1f-4483-924d-289f0b53cef5 · outbound

This paper cites Ouyang, J.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Ouyang, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T23:32:53.031948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:13f9b2b7cf5c598b96f12cf34e90f0d1244c843e2bda6f99f227c76c95f1d7f7

Observation 68375a11-2dae-4b5d-93fc-c272a288fa0d · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Llama 2 via Contrastive Activation Addition

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.575702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:6834ebe9fc4fbb55579b75f2e068b476bc78e0d068b90f3b1e37e603a84482e4

Observation 36dd57f7-f487-42f6-91ba-c90b4d7ae2e1 · outbound

This paper cites Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures , url=.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures , url=

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:32:52.096363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:6113c216468e3aee9c5b985cd1300c84a99564f7f323eeb6ba8d9cefc40033ee

Observation bd069f4b-20c1-4907-a459-b5fdfa511f0c · outbound

This paper cites Red Teaming Language Models with Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Red Teaming Language Models with Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.528288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:425271d14936814eda9ddc758d6b78ca5946fabe49b361272b606a447faab5d5

Observation bd5e148a-6f0e-485a-ba02-f9f20c721942 · outbound

This paper cites an unresolved cited work.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.582385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:eac5e21fd2a57e59341949bec7838a587b7d6a22e0167bb48070da822afc1604

Observation 0ca43b31-5dad-4cb7-b1ed-b6cfed1821fb · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.569631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:9297bf2e2ed981bc59a0057ccf760f4c15e55d79535ce4aa1cc1e0e71dda1574

Observation 50336ac5-2974-406f-ac2d-16be30884918 · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.534733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:db1bbb0c7684d392e0fde10e65e9da6e0e3d75cac3b76589db20d947a0eaba62

Observation cb14639c-b09f-4200-876d-e7027daa343b · outbound

This paper cites Qwen2.5 Technical Report.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Qwen2.5 Technical Report

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.488164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:4bdb238325e5d81d9274b0b661e55f0bc192d0fc64a9029c24541d9b2bf0ea49

Observation 2c6813e8-b676-45f0-ab46-e1eb6fe5266f · outbound

This paper cites NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.681567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:0aa826968631be3aea89b185705f3eac873c46603bbae88dc0f4116f057dd64e

Observation 3d64e567-144f-4188-8b61-b79ad3f04ca7 · outbound

This paper cites an unresolved cited work.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work

Reference 35

Resolution
verified exact
doi, observed 2026-05-19T23:32:52.122157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:cfdb5d513804e31904e9fe5a2af53dc5f46262e0cc0f700c4e74f9e5a9149025

Observation 07f7e6ea-6969-459b-98df-da0589472736 · outbound

This paper cites Srivastava, A.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Srivastava, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T23:32:53.019136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:22e69be0724c05c45f3a4c8cea229406ee19fbbbff77e74e00663b5b11acf8b2

Observation e33c2bcd-0063-4e5b-94fb-a28d5f2ee048 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.553992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:291bd1742d7a261874ef5cd741e1d264207017f663a9c2b6faa32d43aa01b5da

Observation 9b529663-0de8-4b57-9db9-fe1d9617ab69 · outbound

This paper cites Steering Language Models With Activation Engineering.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Language Models With Activation Engineering

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.547630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:35ad44792e6a4700ada0a378d662937471c8fd0f334ce50eaacba15802acf72e

Observation 4bd3a8df-887f-4d62-8d3c-56f212a73d5c · outbound

This paper cites Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:45.367569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:8b86425a90b733b17b01d4f94782d74c3c0d1e2b56575c90ff6c09be4c2f8f15

Observation d3fa68e8-f483-445a-8136-3424e1513475 · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Jailbroken: How Does LLM Safety Training Fail?

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.675659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:7ab0e634f68a2ce1ac4f2c8f3d64853416061728e6a6b5d58be8e1afdcaef517

Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · outbound

This paper cites Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.563721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:acb34333578bd719ebf96efdc58f53dcd7f8d89783bfe28dfbf8567f5272d140

Observation 8d49065c-9fa6-4caf-a8b1-93fde51bc7d1 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Representation Engineering: A Top-Down Approach to AI Transparency

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.468958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:e76e2a3772ee3ae6749709092c11f5a7e9dfee4446fea474ab6d27017f98f086

Observation 5c523340-275e-4231-8476-1c3c69f4d01e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-19T23:32:52.514719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:7cd62ee0c7ba572381c1884f234370d4cbcff5c413f331f08535d492c5d1d3fa

Pith citing papers

Observation e1a4d8d8-0d3d-4adc-a36d-73935169244a · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-01T02:40:29.482309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:40:29.482309Z digest=sha256:c2a09df9e9c16d9513a3968ef69a54e5a04b98982c18ee446956f2be291ac960

Observation bc1711a2-3954-4e0d-bf5d-70cbf9104ef4 · inbound

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response cites this paper.

Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:22.661751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:22.661751Z digest=sha256:9bf52d16f8d09b262f38b063f4a174608eb04aa58d35f08e12b5beb8d0e999b7