Pith. sign in

Paper Citation Record · LEDGER

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2505.14300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14300 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:43:33.763079Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:56:07.082215Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:37:37.332332Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cfadd81-d068-4024-a70c-ecc59ea4ab69 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:28.543259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:28.543259Z digest=sha256:24f2494d4e383da0ca14a0bdc10ddebad06d9c83d42e8245870fe671d254e39f

Observation 5149911b-abee-4c46-9cbc-ff6d294fcf87 · outbound

This paper cites Obfuscated Activations Bypass LLM Latent-Space Defenses.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Obfuscated Activations Bypass LLM Latent-Space Defenses

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.275006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.275006Z digest=sha256:5646f63ce8f86d1f3aabe478fe0c9566ce278ffe1287fde438dad8a2a98bd3de

Observation 06bb95d0-789c-442c-a834-1485bcff7de0 · outbound

This paper cites Towards evaluations-based safety cases for AI scheming.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Towards evaluations-based safety cases for AI scheming

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.414746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.414746Z digest=sha256:575b252e9906cc76b0b43d42774db48c49d1624408deb4cfc8888fd596c0ab2b

Observation d161dd7e-dea4-4a28-9099-57735193d468 · outbound

This paper cites Autoencoders.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Autoencoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.498720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.498720Z digest=sha256:0c72be741724b2151b8947b3e792effa593ba45288494ea5eb6246e5e8dd8682

Observation 1a98e376-f4db-4820-a492-b0f9960b93bd · outbound

This paper cites Taken out of context: On measuring situational awareness in LLMs.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Taken out of context: On measuring situational awareness in LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.616157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.616157Z digest=sha256:76b800a0178b59a66f5f53b0f9f0365b68ff0f681504eb5abd13eee0e2ce6d09

Observation fb07f58a-a39f-4c14-91d0-5bac347e7115 · outbound

This paper cites Safety cases for frontier AI.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Safety cases for frontier AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.714748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.714748Z digest=sha256:4eae7180a037ef78c3ef1d7500e5cc24b2b5e57b15f5d7eee4a36bd75aa2c11f

Observation d262e28e-47f4-4895-8452-53709453a4fd · outbound

This paper cites Scheming AIs: Will AIs fake alignment during training in order to get power?.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Scheming AIs: Will AIs fake alignment during training in order to get power?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.785732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.785732Z digest=sha256:a1e29fcb1ec1064f4c72100c822beab4dce2f74cd82bc79b3cb8812bf9c56d8c

Observation f15c433a-b791-443d-8867-e8415ea2486a · outbound

This paper cites Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Towards Trustworthy and Aligned Machine Learning: A Data-centric Survey with Causality Perspectives

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:30.890006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:30.890006Z digest=sha256:a0fbdebcf0dd3dcdaa96682023afb3946b11ea7c403b69f4110f5a8076c520e3

Observation 27927f34-c2c6-4f08-bc6d-efe71bd82580 · outbound

This paper cites Backdoor defense, learnability and obfuscation, 2025.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Backdoor defense, learnability and obfuscation, 2025

Reference 10

Resolution
verified exact
doi, observed 2026-08-07T15:43:33.973035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:30.988869Z digest=sha256:adcc022606b077314d458d5570e13cbe93218d7b0fbbe76543852b962b505d31

Observation 458fbeb4-4769-4618-a27c-d68c9e721794 · outbound

This paper cites Safety Cases: How to Justify the Safety of Advanced AI Systems.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Safety Cases: How to Justify the Safety of Advanced AI Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.143723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.143723Z digest=sha256:95db7e2cd2c0c55fbb0d3c2301f6b12dc7b6656a41d850ab1820157217479322

Observation 3053807f-71d0-48d8-930e-5bfb833a9279 · outbound

This paper cites Industrial monitoring system, 2025.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Industrial monitoring system, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:36.646763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:31.253688Z digest=sha256:d9372ee3aba6bcc316f289110b4100b87608b48a33d5ef0de880a55856114b1f

Observation ed1e148c-abaa-4d1d-8bf4-ae87b7a95e22 · outbound

This paper cites Dynamic safety cases for frontier AI.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Dynamic safety cases for frontier AI

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:43:34.679398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:31.377089Z digest=sha256:cc0d8a1f3d66846f1e260d36f29cec59f63d67f991011efdb3b5ebc2da0fabf6

Observation eca34920-2c90-4037-aaea-21f704550070 · outbound

This paper cites Challenges with unsupervised LLM knowledge discovery.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Challenges with unsupervised LLM knowledge discovery

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.484748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.484748Z digest=sha256:fc83ebebffd3306f4c6e65f771dc53bc417f80f86c5a0b9a28e2e5d574078cab

Observation 60e2b358-386b-40d1-b04e-197bf857f57f · outbound

This paper cites Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.583205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.583205Z digest=sha256:ac93c6a6a52086980a79372d4af5da77fd68678e82c94e4d31a63382457c8e99

Observation b24a5d12-50fd-4626-bcc3-8ce9ecdfd7c5 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.645292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.645292Z digest=sha256:5d3c790db40338a715093dffedbec343497a8cf361ec07219a447c6109720860

Observation baccc249-5cf1-47a5-9688-517136c1f0f1 · outbound

This paper cites Alignment faking in large language models.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Alignment faking in large language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.745601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.745601Z digest=sha256:171a064d567f831caa600feb385a0e7a3019bf2e2f4d3b38bbbae19afd3d53e1

Observation b74ddfec-d8c7-476c-a6dd-91b303de2238 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.832025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.832025Z digest=sha256:e016f3524925676a2305390df37065da2a70bec26efbb5898bc2905dc1ef9dd0

Observation c6009f29-91d0-4614-81f5-447ec34b5200 · outbound

This paper cites Auto-Encoding Variational Bayes.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Auto-Encoding Variational Bayes

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.900336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.900336Z digest=sha256:85b6e118159713570d09d84966d074ce24eb55c4277d6db2743c552d80161786

Observation 61a54966-ee67-49b2-a9fb-331a09beded8 · outbound

This paper cites The Remarkable Robustness of LLMs: Stages of Inference?.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors The Remarkable Robustness of LLMs: Stages of Inference?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:31.973605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:31.973605Z digest=sha256:0cdb43abe09f32f9a359c646a1a7b380d0ed017ca500ec32821ec27d58e6e3e9

Observation df9ad0c5-094e-4e51-92d1-bc206ceb6ae0 · outbound

This paper cites Me, myself, and ai: The situational awareness dataset (sad) for llms.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Me, myself, and ai: The situational awareness dataset (sad) for llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:36.456599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:32.113134Z digest=sha256:4a96e6dc585d9c5e297ddb684ab5b6d576c2a9826dc1112eb32dd5c098b18301

Observation 0e0f2d5d-5f7f-49fd-9d31-7072f6bb4463 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:32.241219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:32.241219Z digest=sha256:966237af8c513412039da9a05405c1f1f3a7926ddcbaaebc69d99db4c817d24b

Observation fc00e302-4b40-4362-b740-9f094539f6e9 · outbound

This paper cites Decoupled Weight Decay Regularization.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Decoupled Weight Decay Regularization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:32.358776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:32.358776Z digest=sha256:71f8962b5ec5f9a0a8d1bcc5619e1efefff048283750e9ed56e15e57b3209f44

Observation 2e5e159f-8a11-45a7-8260-98f7d073d15f · outbound

This paper cites The "Beatrix'' Resurrections: Robust Backdoor Detection via Gram Matrices.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors The "Beatrix'' Resurrections: Robust Backdoor Detection via Gram Matrices

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:32.438935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:32.438935Z digest=sha256:b942393cef24985e2d5e71a1716192c0a29f8f5a28c0239710b95221bac7d471

Observation d8264b3d-2e10-4e6b-84ce-9cc0183e17f6 · outbound

This paper cites On the generalized distance in statistics.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors On the generalized distance in statistics

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:36.255458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:32.555486Z digest=sha256:071577f4a563abaa0ce0bc47067ce2aa1a624d2f5a7bcbd1d6588e228f8c775d

Observation 70868715-3306-4e10-b4ff-b6e9c50b84ab · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Frontier Models are Capable of In-context Scheming

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:32.674783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:32.674783Z digest=sha256:bfcf218864adad606e3b79aeac6f4c034a8f8a76adf49593c45acf5ae5523252

Observation 5f790046-e524-4b48-bad2-23c3587a4ddd · outbound

This paper cites Ai models can be dangerous before public deployment.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Ai models can be dangerous before public deployment

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:36.055262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:32.785394Z digest=sha256:a083458b9de16a993148b6d831aef7b793b9a1f0bfe0c9d66d24a47cfd2dd112

Observation 3853cfd3-2e59-471b-903c-a948c8436cb1 · outbound

This paper cites Metr’s gpt-4.5 pre-deployment evaluations.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Metr’s gpt-4.5 pre-deployment evaluations

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.836871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:32.938172Z digest=sha256:a7f452d1a506072fdd1dc9f1a6f15d71a85f8ba83e437c70a801547da1b27c20

Observation 1e2c06e0-4ff8-4b4a-b33f-420a00403184 · outbound

This paper cites CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.051071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.051071Z digest=sha256:a8a6f9a65dfcdeb822251da42e1044657bcd7a7e6421893d975254ce5e04046c

Observation 1477b547-3d38-4021-8299-0c3e9ddf5e5f · outbound

This paper cites Robust Backdoor Detection for Deep Learning via Topological Evolution Dynamics.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Robust Backdoor Detection for Deep Learning via Topological Evolution Dynamics

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.167602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.167602Z digest=sha256:11ce13d46d9a275564febab407a2c44aa56584a5d13f7dbf90110e187ed65480

Observation 9112b6eb-cb91-4c29-96db-c6593e38f6d1 · outbound

This paper cites Rail track monitoring system, 2025.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Rail track monitoring system, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.640638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:33.296688Z digest=sha256:16316bfb5a75ff178d57487bddeec2ed90575f30180e978c9a6bbb22d1b0d4bb

Observation 33bf3e96-e499-4ef8-9020-f5c5256b1d61 · outbound

This paper cites Universal Jailbreak Backdoors from Poisoned Human Feedback.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Universal Jailbreak Backdoors from Poisoned Human Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.389364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.389364Z digest=sha256:cd74919e646069a437c281d8c12870b2ae2ee3f44350bbfed2e2edde9e07d89b

Observation de4001fd-d646-49fe-a541-54d33e7daedc · outbound

This paper cites Aviation safety monitoring system, 2025.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Aviation safety monitoring system, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.419024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:33.492380Z digest=sha256:6fa6e3728f35e774537d6d22d7968685042fcdd5dd9493388e2f63db3aa62108

Observation 91657c50-6cce-4055-9586-22d287e48c43 · outbound

This paper cites The Black Swan: The Impact of the Highly Improbable.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors The Black Swan: The Impact of the Highly Improbable

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:43:35.233064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T15:43:33.599145Z digest=sha256:6bc8670b8af3185995e09ee76a2a8c03a591d7c2d3d8a9c65c94ad26f7e6cd03

Observation 9650b709-c852-4e4a-8ce9-b4e35eb16c1c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.696833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.696833Z digest=sha256:8c8285c5f9c243eefacf245d0c330bf0ac13a3313f0d0f22c1ab242a6ba34f64

Observation 968a4e6f-0bb3-495f-b052-a572df396cbc · outbound

This paper cites Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing.

Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:33.763079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:43:33.763079Z digest=sha256:3425549532271c646a5c4da144cac78d15de2cd69e65ca0d2ad8b07191a56a07

Pith citing papers

Observation ffe7e012-7818-4a65-af1a-76856157ffe7 · inbound

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models cites this paper.

Beyond Linear Probes: Dynamic Safety Monitoring for Language Models Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-13T00:17:01.038763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:41:48.620040Z digest=sha256:9c82383dd3f028e92582a026edfd821157caa546cecd7906b0da5710eecec337

Observation 6abccfcf-ec33-4be3-9721-92104184f470 · inbound

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation cites this paper.

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-13T00:17:01.038763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:45:13.320426Z digest=sha256:af1159f785ab363cf491a272f1a08d4abd3f49369720cd615d6ac19dfcb504b4

Observation 7a9b3364-da71-4aba-8598-302f7bfb565e · inbound

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation cites this paper.

Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-13T00:17:01.038763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:15:21.946819Z digest=sha256:a146ea19d7058c315a66996deb8dfd6b34ec09b4e20535499d2abeb07f20d509

Observation 41c58c6d-6c29-4331-ac13-93a62484fa0f · inbound

LLM Scheming Inversely Scales with Pretraining Language Coverage cites this paper.

LLM Scheming Inversely Scales with Pretraining Language Coverage Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T11:56:07.082215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:56:07.082215Z digest=sha256:b2042e66a045afacb4098fdcff702120cacad4ba43956f456690f51937a92043