Pith. sign in

Paper Citation Record · LEDGER

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2608.01414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01414 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:53.661855Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f600858-8d64-49c6-a2d6-d67d271eb025 · outbound

This paper cites Phi-4 Technical Report.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:47.641655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:47.641655Z digest=sha256:755f68ed80eca483f9434a6e597a53d8e29a9381f4279011c995cb11c0dcc864

Observation 90e3ad4a-aa29-4c3c-8322-fc053dbcbf88 · outbound

This paper cites Refusal in language models is mediated by a single direction.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Refusal in language models is mediated by a single direction

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:01.147174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:47.790026Z digest=sha256:f383af95b77310b4d5e1ba879092a9b6bcab14a4c79d9f8bc8d5cd01cfd590d8

Observation 84f80de4-8a14-447b-9b6a-96163838e204 · outbound

This paper cites Qwen technical report, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen technical report, 2023

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:47.909570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:47.909570Z digest=sha256:b8ad564a201bd906e3fa00790310f0d8359f23a4979c24773a30f05990c0b6b2

Observation 4074e66f-a830-4527-aeae-d624851a7bd2 · outbound

This paper cites Qwen2.5-vl technical report, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Qwen2.5-vl technical report, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.015481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.015481Z digest=sha256:f54c23eb80f4833314cd11377ec230c968ea7ff901c0201cb4b5d33df751b1de

Observation 1ce0f009-250b-4a6d-9397-5e55f263ec5f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.167155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.167155Z digest=sha256:e331fb22fb6f056674a5bea62eda2692bbdc96e74b02cad25d98d586bfb03369

Observation 3c523e77-2bcf-4195-86f6-0f3d87f8611d · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Mechanistic Interpretability for AI Safety -- A Review

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.335534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.335534Z digest=sha256:01cbfe1082e8c0af9a5969606b8b60b2b8ab8cfc25afe18b6fec88287d75ac3a

Observation 3c2733ef-2b62-43f7-8dba-730a14994635 · outbound

This paper cites Language models are Homer simpson! safety re-alignment of fine-tuned language models through task arithmetic.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Language models are Homer simpson! safety re-alignment of fine-tuned language models through task arithmetic

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.427125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.427125Z digest=sha256:8e34a42cd84be31471fb88ed9b545cfaf1361081d5ff766f36fdc283328c4d60

Observation a8e4ed37-5070-4455-b867-09c47f93ea5b · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:48.516596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:48.516596Z digest=sha256:4c74161d2c26c74be4722b5138cb14a2851d627102a5f24b68afca790341aabe

Observation 392568be-bb36-4c6b-a44b-9b8c8ead0abb · outbound

This paper cites Brown et al.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Brown et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.861651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:48.616841Z digest=sha256:5b25f624eb61afad319a71cb553765c9ac7c50a8dfb1bf4714ff1a65c009a50c

Observation a82b5044-2036-4a88-b5c9-3d237b593f6e · outbound

This paper cites Towards understanding safety alignment: A mechanistic perspective from safety neurons.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Towards understanding safety alignment: A mechanistic perspective from safety neurons

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.512350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:48.747271Z digest=sha256:617fd68d52d0eafca4cda78c4c863643a78e4255dc737be16cc2b1a3a6aabe26

Observation 099a1440-6997-43b3-840a-86ce3e9ea498 · outbound

This paper cites Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Think you have solved question answering? try ARC, the AI2 reasoning challenge, 2018

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:19:00.155206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:48.913473Z digest=sha256:42053c401cf85e18dc048c99dde7c814f42edfd7994cb41dd426f1786500b31e

Observation d799eba4-ccc5-4f82-8abb-6a5162de9a0f · outbound

This paper cites Training verifiers to solve math word problems, 2021.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training verifiers to solve math word problems, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.043158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.043158Z digest=sha256:82c2a5fa1f3a01aaba92d3a6ffea46fcc4605bca8074dc81660e14c355591ec5

Observation 9497005e-7db5-4842-bdde-59636bd6558f · outbound

This paper cites Dna: Uncovering universal latent forgery knowledge.arXiv preprint arXiv:2601.22515, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dna: Uncovering universal latent forgery knowledge.arXiv preprint arXiv:2601.22515, 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.155369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.155369Z digest=sha256:f4904013f496d1d7f0499994684891826ca0679ae78db2499ca1a993865ee8e0

Observation 7329596c-6e32-404e-8807-106b89078f3f · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.251487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.251487Z digest=sha256:12df6b1b7fc7f8b72d7cdb8fc415a30728f7b208c266d1f0f4da8bb20eefe9e2

Observation afe3a815-d650-4634-8ae1-922834e6efcf · outbound

This paper cites FigStep: Jailbreaking large vision-language models via typographic visual prompts.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks FigStep: Jailbreaking large vision-language models via typographic visual prompts

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.736632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:49.356894Z digest=sha256:35fe8fe967fabc550b3613435d21626b5c4c5d58d23c523fec967e0e2eba4963

Observation a3904cd3-be16-40e6-ab81-a1f911ef0b4f · outbound

This paper cites The llama 3 herd of models, 2024.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks The llama 3 herd of models, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.592647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.592647Z digest=sha256:6234bb53ac7add55534eb6a9214347f31f104fd82d09ecf8fc043cb3af6742af

Observation 965dcc06-edae-409d-a29e-d0eabdd9e20e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.709052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.709052Z digest=sha256:21e4b87a01c8d6f8bc8cca8e55f83c7b88aa89746cd61aa622f3dfd4ffae4a5b

Observation b3cac4d5-8985-4e00-974c-21b577d2d13d · outbound

This paper cites PKU-SafeRLHF: Towards multi- level safety alignment for LLMs with human preference.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks PKU-SafeRLHF: Towards multi- level safety alignment for LLMs with human preference

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:49.818291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:49.818291Z digest=sha256:3bf7d4fc4dc04b2a8f9cf2ef364259960f64d72c38196db62fba93467f4cb572

Observation a5a78fc4-7c1b-4f4b-8762-7559e5ee21e5 · outbound

This paper cites an unresolved cited work.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:59.462162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:49.948433Z digest=sha256:3b8adf396f665f62c944cbeb6e1303f7ed276b8194add8636a300e755ed5fde3

Observation 9efe5ece-fc1d-4b04-a61f-183b40f76cd2 · outbound

This paper cites Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.246967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:50.077899Z digest=sha256:7a0e4ca4968825b74d7089ee567b8803419b2e2a9e98729153cb6488a7711428

Observation 391afc04-b9e6-4eea-b279-73d6226f29f5 · outbound

This paper cites Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:59.008836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:50.254018Z digest=sha256:a2834da51d0f05bdf96f802b4f6d74ecc1e9f9a5ec75095c33a4a95641897754

Observation 95d81f17-dbd3-4b39-8784-51c43cd96727 · outbound

This paper cites TruthfulQA: Measuring how models mimic human falsehoods.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TruthfulQA: Measuring how models mimic human falsehoods

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.773225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:50.386078Z digest=sha256:d9ef9f8da61301ee8ae4adacbf500c4826d05941a387646fdb768719bdc52d95

Observation b5f373f9-07a4-4730-bcde-6323a666b9bb · outbound

This paper cites Visual instruction tuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Visual instruction tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.496181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.496181Z digest=sha256:6f39f9c9de08b47f592f2ae314ebb5a041b99d88114a804aa3e9d39ad7e7b8dd

Observation 642773eb-c65f-414d-ba00-0ccf0ba78eef · outbound

This paper cites MMBench: Is your multi-modal model an all- around player?, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks MMBench: Is your multi-modal model an all- around player?, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.613801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.613801Z digest=sha256:2f722621f7fa69390f527d65192d68dfc1b8a15324e1776d273245824c624352

Observation 517f68ac-7397-4396-b81b-ca04a9826e64 · outbound

This paper cites Decoupled weight decay regulariza- tion.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Decoupled weight decay regulariza- tion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:50.721583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:50.721583Z digest=sha256:c3e8bc740b5d379c6e49f5ce8d357430af019ed750c31293b160e97359726fc4

Observation 36e0cc30-eb5d-471d-84cc-5cd5421fb6a2 · outbound

This paper cites HarmBench: A standardized evaluation framework for automated red teaming and robust refusal.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks HarmBench: A standardized evaluation framework for automated red teaming and robust refusal

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.433158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:50.847333Z digest=sha256:a7c2f475ddc0601679440c739781d5986f93b17b3d48ca9d85d536895d093368

Observation 83054a01-f0cd-4ff9-a024-7392f6604c96 · outbound

This paper cites Pruning convolutional neural networks for resource efficient inference.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Pruning convolutional neural networks for resource efficient inference

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:58.072871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:50.993536Z digest=sha256:f0320feb0d123f388c9e454559b72bc29b51181624b065627b06da2e7acf92c6

Observation 777b2304-c637-43ff-89cf-f44664ae95d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Training language models to follow instructions with human feedback

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:57.645617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:51.108876Z digest=sha256:8ef54ff6565591d166757353f4c9cb6965371fcbc00d52398746e0e815d91c04

Observation 67ad9ef2-ede9-4589-83c2-ac872a61d6fc · outbound

This paper cites Representation noising: A defence mech- anism against harmful finetuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation noising: A defence mech- anism against harmful finetuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:57.275855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:51.280048Z digest=sha256:9a33702e3db0886ae7bcf26e7d402434c617e72b77a215b41affbf7433670359

Observation 3874d651-a443-4ef3-a1ff-56db695739b6 · outbound

This paper cites XSTest: A test suite for identifying exaggerated safety behaviours in large language models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks XSTest: A test suite for identifying exaggerated safety behaviours in large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.952210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:51.435864Z digest=sha256:c4ac1346ac4139e74d8560108c8e991188e933a183c15f77beeb0cd3fd9a790a

Observation b45e36db-5e0f-406e-a599-c1632c39a1f4 · outbound

This paper cites Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:51.590300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:51.590300Z digest=sha256:503f88b12168f04486dbb9def5e394cba6f6813002346d5665fccf87b0986242

Observation 0382d27c-9de5-4212-a464-d3152fe14255 · outbound

This paper cites Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Where culture fades: revealing the cultural gap in text-to-image generation.arXiv preprint arXiv:2511.17282, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:51.980020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:51.980020Z digest=sha256:b93d08f4208fa82af77609ab26637a30bf2e5ec637d59cfa3ac272cd2dccd6df

Observation f1ad0204-b8c7-497d-be98-9ce6e5cde755 · outbound

This paper cites TraceRouter: Robust safety for large foundation models via path-level intervention.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks TraceRouter: Robust safety for large foundation models via path-level intervention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.579897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:52.226014Z digest=sha256:dd0176998514dfbd9b94f31d3e77e73cf35f7de00c0b0440ca7faad379b08d1a

Observation c64ee951-f054-4f4e-8519-3ab80e966c19 · outbound

This paper cites Orthoeraser: coupled-neuron orthogonal projection for concept erasure.arXiv preprint arXiv:2603.11493, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Orthoeraser: coupled-neuron orthogonal projection for concept erasure.arXiv preprint arXiv:2603.11493, 2026

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.311850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.311850Z digest=sha256:80643cfb388ed109254ae7c2f78570e8e966d23a330aa4c9212a96f9fc246e8b

Observation 51dcc1f5-2ec3-4bba-abf7-f0e651ef2a17 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks A StrongREJECT for Empty Jailbreaks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.404593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.404593Z digest=sha256:424701195b46638c11543720348e14fece6eed22c2a1e9550b85cc48739e8c82

Observation 7018b5b6-ffcc-434f-8f23-acb594ded70c · outbound

This paper cites Dropout: A simple way to prevent neural networks from overfitting.Journal of Machine Learning Research, 15(56):1929–1958, 2014.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Dropout: A simple way to prevent neural networks from overfitting.Journal of Machine Learning Research, 15(56):1929–1958, 2014

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.471514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.471514Z digest=sha256:bafd2646ba314770eab18f1498544d37558e3f23f5d2215b2805f5846576b4f0

Observation 77401e9b-678a-40dc-9eaa-01507ef87998 · outbound

This paper cites Zico Kolter.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.559100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.559100Z digest=sha256:20714c9754fb480ca4b56550a4e2bc5047ddb6413e511deccf32614a29256f4e

Observation 3844070b-028c-4aed-b328-806cd2808c4f · outbound

This paper cites Zico Kolter.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Zico Kolter

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:56.252598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:52.637182Z digest=sha256:2d16f66b26a2f5214942650e06999ab70b7b4ba2cce98074849b0c58bf5d9b71

Observation 5881ff72-767d-43ea-ad4a-d7f368a58b0b · outbound

This paper cites SafeNeuron: Neuron- level safety alignment for large language models.arXiv preprint arXiv:2602.12158, 2026.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks SafeNeuron: Neuron- level safety alignment for large language models.arXiv preprint arXiv:2602.12158, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.703131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.703131Z digest=sha256:9793928eff018e1bb7ccdf3f715d6c0903b8a587ca2553fa23039ce769ed7e63

Observation 4d07861b-3234-4a57-b119-647180c55b74 · outbound

This paper cites Taxonomy, opportunities, and challenges of representation engi- neering for large language models.arXiv preprint arXiv:2502.19649, 2025.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Taxonomy, opportunities, and challenges of representation engi- neering for large language models.arXiv preprint arXiv:2502.19649, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:52.787762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:52.787762Z digest=sha256:43c8f418eb9696a9a0283438dd722149c087e18a3950046f62f619d2959dffe2

Observation b650a3e8-a2df-4c1a-b276-41f8ed9d33ba · outbound

This paper cites Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2023.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Jailbroken: How does LLM safety training fail? InAdvances in Neural Information Processing Systems, volume 36, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.871548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:52.858768Z digest=sha256:ddde5e741ae835a6b779c9c5ec357e52ad0ec926a599eabf7a032924be4b35d1

Observation 4f1293c3-bd93-43c6-a8d7-123bee059ef4 · outbound

This paper cites NeuroStrike: Neuron-level attacks on aligned LLMs.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NeuroStrike: Neuron-level attacks on aligned LLMs

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.461925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:52.926634Z digest=sha256:9c0c96df50e80ece76d5ddef6c7e6e2cec65e38405054cdf330d84bc1166105e

Observation 11b3413a-478c-4fd6-9f1d-e45d8ad267b3 · outbound

This paper cites NLSR: Neuron-level safety realignment of large language 9 models against harmful fine-tuning.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks NLSR: Neuron-level safety realignment of large language 9 models against harmful fine-tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:55.100559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:52.993493Z digest=sha256:12091d19c5f09e99713133e8a5d20df3d206938fba84ba7c100b7d6a8e3fde62

Observation b4639a0e-6f5b-4dbe-8f57-3ad4a739267c · outbound

This paper cites Representation bending for large language model safety.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Representation bending for large language model safety

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.924763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:53.071474Z digest=sha256:d999237582beabcb9c7f0ac2765e66fff277186ba1de713c7f40024930e590c6

Observation 433c468d-dbc8-42b1-8e23-0e716e9f8166 · outbound

This paper cites Weston, and Xian Li.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Weston, and Xian Li

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.675581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:53.303482Z digest=sha256:6d78633bb2823925847a24a81e0509a4287e16d19f6226413acc60897619e20e

Observation dc92f761-cf4f-40ed-91ff-dfb68f534e9c · outbound

This paper cites Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Understanding and enhancing safety mechanisms of LLMs via safety-specific neuron

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:54.473399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T00:18:53.415568Z digest=sha256:1ed5f2a22e7ff6869b6a7174d3d37449170ee04bdeb4125558371eb5a834d4ab

Observation 552e4e83-f8bc-4208-a42c-770f591127eb · outbound

This paper cites Improving Alignment and Robustness with Circuit Breakers.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Improving Alignment and Robustness with Circuit Breakers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:53.554407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:53.554407Z digest=sha256:cbdcfee0aae3b854c07e0a075dfe87f2274156f03bcd00b910c1000d3c26429e

Observation c28e9822-e015-46e2-99c6-67dbb638237f · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:53.661855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:53.661855Z digest=sha256:5dd8566600ac07620893eb0962877eed1564ba132f64ac22b7de99d56b144015

Pith citing papers

No inbound Pith citation observations are available.