Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:53.661855Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01414.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:53.661855Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1f600858-8d64-49c6-a2d6-d67d271eb025 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Phi-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e3ad4a-aa29-4c3c-8322-fc053dbcbf88 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Refusal in language models is mediated by a single direction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84f80de4-8a14-447b-9b6a-96163838e204 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen technical report, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4074e66f-a830-4527-aeae-d624851a7bd2 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen2.5-vl technical report, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce0f009-250b-4a6d-9397-5e55f263ec5f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c523e77-2bcf-4195-86f6-0f3d87f8611d · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Mechanistic Interpretability for AI Safety -- A Review
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2733ef-2b62-43f7-8dba-730a14994635 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Language models are Homer simpson! safety re-alignment of fine-tuned language models through task arithmetic
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e4ed37-5070-4455-b867-09c47f93ea5b · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 392568be-bb36-4c6b-a44b-9b8c8ead0abb · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Brown et al
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a82b5044-2036-4a88-b5c9-3d237b593f6e · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Towards understanding safety alignment: A mechanistic perspective from safety neurons
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 099a1440-6997-43b3-840a-86ce3e9ea498 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d799eba4-ccc5-4f82-8abb-6a5162de9a0f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training verifiers to solve math word problems, 2021
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9497005e-7db5-4842-bdde-59636bd6558f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dna: Uncovering universal latent forgery knowledge.arXiv preprint arXiv:2601.22515, 2026
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7329596c-6e32-404e-8807-106b89078f3f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Gemma: Open Models Based on Gemini Research and Technology
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afe3a815-d650-4634-8ae1-922834e6efcf · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks FigStep: Jailbreaking large vision-language models via typographic visual prompts
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3904cd3-be16-40e6-ab81-a1f911ef0b4f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks The llama 3 herd of models, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965dcc06-edae-409d-a29e-d0eabdd9e20e · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cac4d5-8985-4e00-974c-21b577d2d13d · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks PKU-SafeRLHF: Towards multi- level safety alignment for LLMs with human preference
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5a78fc4-7c1b-4f4b-8762-7559e5ee21e5 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9efe5ece-fc1d-4b04-a61f-183b40f76cd2 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 391afc04-b9e6-4eea-b279-73d6226f29f5 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95d81f17-dbd3-4b39-8784-51c43cd96727 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TruthfulQA: Measuring how models mimic human falsehoods
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5f373f9-07a4-4730-bcde-6323a666b9bb · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Visual instruction tuning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642773eb-c65f-414d-ba00-0ccf0ba78eef · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks MMBench: Is your multi-modal model an all- around player?, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517f68ac-7397-4396-b81b-ca04a9826e64 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Decoupled weight decay regulariza- tion
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e0cc30-eb5d-471d-84cc-5cd5421fb6a2 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks HarmBench: A standardized evaluation framework for automated red teaming and robust refusal
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83054a01-f0cd-4ff9-a024-7392f6604c96 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Pruning convolutional neural networks for resource efficient inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 777b2304-c637-43ff-89cf-f44664ae95d7 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 67ad9ef2-ede9-4589-83c2-ac872a61d6fc · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation noising: A defence mech- anism against harmful finetuning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3874d651-a443-4ef3-a1ff-56db695739b6 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b45e36db-5e0f-406e-a599-c1632c39a1f4 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0382d27c-9de5-4212-a464-d3152fe14255 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282, 2025
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ad0204-b8c7-497d-be98-9ce6e5cde755 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TraceRouter: Robust safety for large foundation models via path-level intervention
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c64ee951-f054-4f4e-8519-3ab80e966c19 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Orthoeraser: coupled-neuron orthogonal projection for concept erasure.arXiv preprint arXiv:2603.11493, 2026
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51dcc1f5-2ec3-4bba-abf7-f0e651ef2a17 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks A StrongREJECT for Empty Jailbreaks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7018b5b6-ffcc-434f-8f23-acb594ded70c · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dropout: A simple way to prevent neural networks from overfitting.Journal of Machine Learning Research, 15(56):1929–1958, 2014
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77401e9b-678a-40dc-9eaa-01507ef87998 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3844070b-028c-4aed-b328-806cd2808c4f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5881ff72-767d-43ea-ad4a-d7f368a58b0b · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks SafeNeuron: Neuron- level safety alignment for large language models.arXiv preprint arXiv:2602.12158, 2026
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d07861b-3234-4a57-b119-647180c55b74 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Taxonomy, opportunities, and challenges of representation engi- neering for large language models.arXiv preprint arXiv:2502.19649, 2025
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b650a3e8-a2df-4c1a-b276-41f8ed9d33ba · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2023
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f1293c3-bd93-43c6-a8d7-123bee059ef4 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NeuroStrike: Neuron-level attacks on aligned LLMs
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11b3413a-478c-4fd6-9f1d-e45d8ad267b3 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NLSR: Neuron-level safety realignment of large language 9 models against harmful fine-tuning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4639a0e-6f5b-4dbe-8f57-3ad4a739267c · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation bending for large language model safety
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 433c468d-dbc8-42b1-8e23-0e716e9f8166 · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Weston, and Xian Li
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc92f761-cf4f-40ed-91ff-dfb68f534e9c · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 552e4e83-f8bc-4208-a42c-770f591127eb · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Improving Alignment and Robustness with Circuit Breakers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c28e9822-e015-46e2-99c6-67dbb638237f · outbound
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.