Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T23:30:43.364230Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2605.17413.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T23:30:43.364230Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:33:22.661751Z
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2466f2a8-3957-4bd2-9784-d30e6e0e1b1e · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Abu Shairah, H
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5fb3290c-caf9-4221-ba01-644501de1ddd · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Agnihotri, J
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c3223e49-2b7b-4b69-8b38-4a11f11e196b · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Refusal in Language Models Is Mediated by a Single Direction
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db77ce8b-27e5-47ed-8840-b73dde9f9500 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Program Synthesis with Large Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9737239-297b-43cb-a98f-e0078e8d7eb7 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5ed94200-c4c9-4423-9f60-23163f017187 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation de0b055b-e787-4943-bf8c-d0ea299bb904 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a07616ce-b003-419e-bbc4-ba6915ca47da · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications On the Opportunities and Risks of Foundation Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3e39453c-7c12-46ac-a59e-9c1d15d5db21 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bbce25cf-ee0a-4e7c-8e1d-3f5f97488613 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5612550d-065c-4012-b7c3-24a25892b71b · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 196a59de-35f2-4618-8c78-30dd39918469 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f00e7438-79df-40a3-8912-a519316862d8 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications LlamaFirewall: An open source guardrail system for building secure AI agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ab770325-cefb-4c81-ba33-d85571ac4ac9 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec470c95-7358-4a75-8d25-eea7f57aae5c · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa13707c-57a0-4c0e-b235-3f054072c6ae · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0f8ad43b-5b47-40ae-82e4-7be8e0b5e3f8 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Hendrycks, C
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 10e25912-2901-49f6-9968-1db84752ee79 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4406fcf-3803-42c7-90f6-c581593a416a · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2b30cea1-9791-4949-9074-9b16896b4e7e · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Mistral 7B
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 788a357c-ce93-4de2-bc79-a5309fd52864 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e086d031-c757-4762-b35c-1cce93f7f548 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications URLhttps://doi.org/10.18653/v1/2022.acl-long.229
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d13932a9-8966-44b8-aedc-515d37999f05 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Refusal in LLMs is an Affine Function
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 05dd2698-ac12-47b2-9b17-9b6528de2db0 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ec170ce0-4d8b-4f1b-8070-d72265cd1409 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Language Model Refusal with Sparse Autoencoders
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5aefa60d-0c1f-4483-924d-289f0b53cef5 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Ouyang, J
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68375a11-2dae-4b5d-93fc-c272a288fa0d · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Llama 2 via Contrastive Activation Addition
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 36dd57f7-f487-42f6-91ba-c90b4d7ae2e1 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures , url=
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd069f4b-20c1-4907-a459-b5fdfa511f0c · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Red Teaming Language Models with Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd5e148a-6f0e-485a-ba02-f9f20c721942 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ca43b31-5dad-4cb7-b1ed-b6cfed1821fb · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 50336ac5-2974-406f-ac2d-16be30884918 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb14639c-b09f-4200-876d-e7027daa343b · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Qwen2.5 Technical Report
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c6813e8-b676-45f0-ab46-e1eb6fe5266f · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d64e567-144f-4188-8b61-b79ad3f04ca7 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 07f7e6ea-6969-459b-98df-da0589472736 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Srivastava, A
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e33c2bcd-0063-4e5b-94fb-a28d5f2ee048 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b529663-0de8-4b57-9db9-fe1d9617ab69 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Language Models With Activation Engineering
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4bd3a8df-887f-4d62-8d3c-56f212a73d5c · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Few Tokens, Big Leverage: Preserving Safety Alignment by Constraining Safety Tokens during Fine-tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d3fa68e8-f483-445a-8136-3424e1513475 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Jailbroken: How Does LLM Safety Training Fail?
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 83f64bae-f5e2-4773-88c2-c67e3df1b3d2 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d49065c-9fa6-4caf-a8b1-93fde51bc7d1 · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Representation Engineering: A Top-Down Approach to AI Transparency
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5c523340-275e-4231-8476-1c3c69f4d01e · outbound
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1a4d8d8-0d3d-4adc-a36d-73935169244a · inbound
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1711a2-3954-4e0d-bf5d-70cbf9104ef4 · inbound
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.