Pith. sign in

Paper Citation Record · LEDGER

AI Control: Improving Safety Despite Intentional Subversion

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2312.06942.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06942 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:31:14.158845Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 56058db1-e979-4b2c-862e-a32b97257ba5 · inbound

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation cites this paper.

Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation AI Control: Improving Safety Despite Intentional Subversion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T18:31:14.158845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:31:14.158845Z digest=sha256:50499cbffa70c603fa09b6c1d771e6e4ae4a73bf55ea84aa039d56071f5930d8

Observation 58d6d5d9-4759-4f46-84ac-428afd55a46c · inbound

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management cites this paper.

A Frontier AI Risk Management Framework: Bridging the Gap Between Current AI Practices and Established Risk Management AI Control: Improving Safety Despite Intentional Subversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:48:44.440069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:48:44.440069Z digest=sha256:f52ca71b79f07de7d8ff306a612ce0c2cf5244f10ac447c0d3bab7be69341483

Observation 62406f3f-8d92-4d13-9b40-ac547bef0427 · inbound

Learning Safety Constraints for Large Language Models cites this paper.

Learning Safety Constraints for Large Language Models AI Control: Improving Safety Despite Intentional Subversion

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:30:15.698528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:30:15.698528Z digest=sha256:32915a9abce6fca7af71af59a3f20702e02f43fc9fe5f5092e6dfc17502b4656

Observation 6ee9e0b3-3fdd-4e45-a4b6-3791648bdc1f · inbound

Systematic Hazard Analysis for Frontier AI using STPA cites this paper.

Systematic Hazard Analysis for Frontier AI using STPA AI Control: Improving Safety Despite Intentional Subversion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:38:04.134407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:38:04.134407Z digest=sha256:f99e928032f284600912c07a0d4bf64dbbaba29b3bd0a0d86c62f3e421f35ccc

Observation 3811c15b-4fce-40af-af94-5b024e4c88ba · inbound

To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems cites this paper.

To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems AI Control: Improving Safety Despite Intentional Subversion

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:32:17.342323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:30:47.877793Z digest=sha256:beae897dd181472b105fab05535d92f2c20cfc74908e025ffc8a91dfc26d497d

Observation c75cd383-58c4-4964-8eca-3ea296dfcb39 · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models AI Control: Improving Safety Despite Intentional Subversion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.307636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.307636Z digest=sha256:030fbe23db93f5fef7f4d3afe04f077ea995c922a4a33c98ecefa53b451f19b3

Observation e61b53ec-69d8-4bec-8588-19b1ebf02a2b · inbound

A Red Teaming Roadmap Towards System-Level Safety cites this paper.

A Red Teaming Roadmap Towards System-Level Safety AI Control: Improving Safety Despite Intentional Subversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:19.176489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:19.176489Z digest=sha256:b85b889c21cae4e78b7ffebcc1640447bf8ba307641d4adc094b0626d4b674ca

Observation 171bc1d9-efa4-4d45-85cc-b610962d7f10 · inbound

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values cites this paper.

Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values AI Control: Improving Safety Despite Intentional Subversion

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:07.278246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:07.278246Z digest=sha256:92633c2c9f7f0245b436fb6240386a07afdb62ee403442e97a753a4a87a2f81e

Observation ab12fc43-486f-4daa-bac5-8049aee7dcf1 · inbound

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) cites this paper.

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) AI Control: Improving Safety Despite Intentional Subversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:57.283865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:57.283865Z digest=sha256:5c6b80331ed447be1ad02eb8996143d31ea92e17da5f471361032b3e5765f45c

Observation 6511c20f-7a93-40c4-b169-2f85f00cba12 · inbound

Subversion via Focal Points: Investigating Collusion in LLM Monitoring cites this paper.

Subversion via Focal Points: Investigating Collusion in LLM Monitoring AI Control: Improving Safety Despite Intentional Subversion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:52:29.430622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:52:29.430622Z digest=sha256:e39c934940a5d0670d0c3311e45202f79018da7a153e28af5c7f6c69e6df8a26

Observation ce29ce03-2324-465c-84a3-bcc948022b44 · inbound

Towards Measurement Theory for Artificial Intelligence cites this paper.

Towards Measurement Theory for Artificial Intelligence AI Control: Improving Safety Despite Intentional Subversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:27:33.196672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:27:33.196672Z digest=sha256:007ce05e5376999694483f700be7e93d912cb5d97e25f9e1af5c1d66fafe3121

Observation cdba531b-0cde-4c13-b6f1-c526f77578c9 · inbound

The bitter lesson of misuse detection cites this paper.

The bitter lesson of misuse detection AI Control: Improving Safety Despite Intentional Subversion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.580182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.580182Z digest=sha256:778753ddf65d2073d7501302cb54496e70fc28e15bcd5da235c63c03bb84a527

Observation a1a61a6d-dca5-4933-8d5b-37778dd21446 · inbound

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models cites this paper.

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models AI Control: Improving Safety Despite Intentional Subversion

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:08:22.995530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:08:22.995530Z digest=sha256:9011991fa9788e0a207c19dc0eab7aa965ddddfc0fbf391beceaaf442869dabc

Observation 319fba47-17c7-4813-9d53-370ada5c8e78 · inbound

Investigating Crossing Perception in 3D Graph Visualisation cites this paper.

Investigating Crossing Perception in 3D Graph Visualisation AI Control: Improving Safety Despite Intentional Subversion

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T10:22:55.800253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:22:55.800253Z digest=sha256:d0478143cfa4aa6ed953b51de017e23392c457f8d0059e03c4c1dc3236ad73a5

Observation 68cc17b8-adda-4376-bcc8-23d1f798ffc3 · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:49.397384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:49.397384Z digest=sha256:0345e764c76149df76a41d378896ad7828bac3b8cd8283dce22de0b93afafa73

Observation 99166c4d-a50d-4b60-8e9e-6240a4da226a · inbound

NEST: Nascent Encoded Steganographic Thoughts cites this paper.

NEST: Nascent Encoded Steganographic Thoughts AI Control: Improving Safety Despite Intentional Subversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:21:32.317208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:21:32.317208Z digest=sha256:5946d083342c0a27371c99d2e42e056c1fb6d62172261d02ae667fdd826836ba

Observation 93d90a3a-5b31-448d-826b-c405c9aaa3ac · inbound

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability cites this paper.

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability AI Control: Improving Safety Despite Intentional Subversion

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T22:48:22.975610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T22:46:39.337887Z digest=sha256:231d6aff2e4c1e6ace00a9de33e4a523e2cd8f3a0f30e7ec110270d860994dde

Observation a3b4a598-2138-491b-b663-57008d050bb5 · inbound

TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol cites this paper.

TraceGuard: Structured Multi-Dimensional Monitoring as a Collusion-Resistant Control Protocol AI Control: Improving Safety Despite Intentional Subversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T11:41:46.131575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T11:41:46.131575Z digest=sha256:a51fbd86023774c332db7d6fa5c749b9ecf50e7426f5855529f1dd4ae4b910b5

Observation 01c8868e-6d6e-4843-bb2b-b98d861f58f2 · inbound

Detecting Safety Violations Across Many Agent Traces cites this paper.

Detecting Safety Violations Across Many Agent Traces AI Control: Improving Safety Despite Intentional Subversion

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T21:20:52.795352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:38:41.835794Z digest=sha256:0d9fd4ba4542e670d9b01a7f05a37896a63205b5ab88de5321bbf97bfd093ee3

Observation b2e5325a-4b53-44b5-a717-93cb8c4cdc0e · inbound

Geographic Blind Spots in AI Control Monitors: A Cross-National Audit of Claude Opus 4.6 cites this paper.

Geographic Blind Spots in AI Control Monitors: A Cross-National Audit of Claude Opus 4.6 AI Control: Improving Safety Despite Intentional Subversion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:35:19.300181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T08:30:47.680521Z digest=sha256:ace1cbd8fc2783500c430bb41df6f2e21ea8384cc84ce753ccc83d0c739145be

Observation 3838f414-d5da-4cf9-bfa6-7abff81c1ba1 · inbound

From Admission to Invariants: Measuring Deviation in Delegated Agent Systems cites this paper.

From Admission to Invariants: Measuring Deviation in Delegated Agent Systems AI Control: Improving Safety Despite Intentional Subversion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.099656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:18:51.614295Z digest=sha256:56f8cf7ad28bfc445a5849865e1252873c92d7c8f8886be900a345ba9d7d38a2

Observation 9471e4d9-9796-4e58-9cfa-2046ce6c6c6e · inbound

ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data cites this paper.

ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data AI Control: Improving Safety Despite Intentional Subversion

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:10.985688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:55:41.400468Z digest=sha256:f7db126ca30c0f02f34d52febe7209539d8f5f7663e8e48f6b71bfc6f71c741f

Observation bcad1dac-5128-4161-9b99-b3d046604445 · inbound

Estimating Tail Risks in Language Model Output Distributions cites this paper.

Estimating Tail Risks in Language Model Output Distributions AI Control: Improving Safety Despite Intentional Subversion

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.142882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T12:26:41.797166Z digest=sha256:57f1334ecad4db38704751bdb7d356283ff0b1c8ec6c3822db698b485312490d

Observation 4a94edb4-7645-49b7-b1b2-242218782e9e · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use AI Control: Improving Safety Despite Intentional Subversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T23:16:16.810684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:6bd7ea41100c5ee4594453976b06242c4812b820425047317daf8a83de5d74d8

Observation 8a47d8b9-ae9d-4b04-9b44-ea27e042e8cf · inbound

Automated alignment is harder than you think cites this paper.

Automated alignment is harder than you think AI Control: Improving Safety Despite Intentional Subversion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.157741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:47:05.526632Z digest=sha256:6946a1f1652e316caee991cac70d2eb84d0baceda6a81a856d035c974489a5f9

Observation 121cd7ad-b89d-4e29-a212-ebf9b5346617 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AI Control: Improving Safety Despite Intentional Subversion

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:07.940397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T05:50:28.114140Z digest=sha256:1ccf5960eb15b183107f3e82128380b7c03d96425af7f70121f9fa4783444a1a

Observation 1ae322a6-a659-413e-8b5e-f6e772c49926 · inbound

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety cites this paper.

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety AI Control: Improving Safety Despite Intentional Subversion

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:06:42.857383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T06:05:27.736494Z digest=sha256:292dc286ee5be731e284ed78e48d9d648739c8cb1d7dc94674468cef243ad389

Observation b1ea782d-195a-4483-aa0d-7b2de7c793fe · inbound

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems cites this paper.

Operationalizing Reconstructive Authority: Runtime Construction, Dependency Resolution, and Execution Gating in Autonomous Agent Systems AI Control: Improving Safety Despite Intentional Subversion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:29:59.632637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T17:27:05.545088Z digest=sha256:d1003287905784ccd3d0298f4df5365bd399ae21602390a255b7b628b41e94b3

Observation d55c12cd-22ac-4e05-86b7-2c42c2763fb5 · inbound

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning cites this paper.

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning AI Control: Improving Safety Despite Intentional Subversion

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:48.019328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:29:34.096277Z digest=sha256:62b10ec7c891cbc94a7a0175420fe568930bbba95704a3f14f7f041dd9f398fc

Observation 47bb62e4-2c29-4758-b8c1-cbe4fa99cf4a · inbound

AI Integrity: Defending Against Backdoors and Secret Loyalties cites this paper.

AI Integrity: Defending Against Backdoors and Secret Loyalties AI Control: Improving Safety Despite Intentional Subversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:59:54.643757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-04T14:56:53.806480Z digest=sha256:183dc771baf3a0ef11fe0a5e2a7339fb3093e84b18437f31fedc1c524900c7a2

Observation 0b7bb4b9-9f84-4cb1-9e51-fa515ead3f77 · inbound

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling cites this paper.

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling AI Control: Improving Safety Despite Intentional Subversion

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:16:24.582255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T14:29:43.379840Z digest=sha256:f42895b8b13409d3a770145291efafa6bd9479d9709f1e9b4313d9c0c346332a

Observation b6802ae8-8481-4060-bb5d-272673ad8924 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AI Control: Improving Safety Despite Intentional Subversion

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.124671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:37814f8a7e2f8c5e6641fc0d85e800eeaf3e79619321fd0c39dc691cccab2488

Observation 65be4607-7185-4b38-b2e8-c593527d6f27 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AI Control: Improving Safety Despite Intentional Subversion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:333f6c9a505ac0d263a9586baae581b763fc0b79bd9178e2cf05d531a1033194

Observation 79ec6dc4-84aa-44a7-9073-31b257e5b5b3 · inbound

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing cites this paper.

A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing AI Control: Improving Safety Despite Intentional Subversion

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.565397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:17:01.173495Z digest=sha256:0b98cd6528308e886feb493caa50223740d901bee3e68204a93a560fc3cf6a1a

Observation 57645995-3d94-4668-9cc4-c0cac1a03d0c · inbound

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents cites this paper.

TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T17:07:12.748737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:11:29.317063Z digest=sha256:15ad03d690947c3b77a58e23451bb163d3c8bc50f58b960e9c53263267527e4d

Observation 81c34412-a419-4955-acf7-384e675bfbd2 · inbound

The Distributed Detectability Band Against Marginal-Preserving Attacks cites this paper.

The Distributed Detectability Band Against Marginal-Preserving Attacks AI Control: Improving Safety Despite Intentional Subversion

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.863875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T12:58:22.056355Z digest=sha256:f4dd4b9c7dca3c49dba3e47d7a183aab67771ecfea8f25abe3cef0b1da696b26

Observation 478e43f2-a038-493c-9bc1-047ce3c342db · inbound

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms cites this paper.

"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms AI Control: Improving Safety Despite Intentional Subversion

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.296740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T09:51:16.969884Z digest=sha256:c5135f7394f676851c8b6ad5f79854afca6719c7b86387fb7483eb1b8a190d7d

Observation f90f6bfe-30d3-41f1-82d4-0a6db6d86e14 · inbound

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems cites this paper.

Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems AI Control: Improving Safety Despite Intentional Subversion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.222877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T00:42:20.220668Z digest=sha256:3bbac2d3f633e8e8c68b186ef34471e8bca5d1304a028a3ecb56234c6c93c985

Observation db9b77e0-160e-4860-bd4e-4fc46418e204 · inbound

Online Safety Monitoring for LLMs cites this paper.

Online Safety Monitoring for LLMs AI Control: Improving Safety Despite Intentional Subversion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:07.488126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T12:53:49.681641Z digest=sha256:2bf8c4562034a4980d4bbc45ae4c34b534aae43e825a1b9efc50d3fbe48bee80

Observation e4e7d8dc-ac18-43b1-a7dc-ca15667e423d · inbound

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages cites this paper.

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages AI Control: Improving Safety Despite Intentional Subversion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T04:00:44.072640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T04:00:44.072640Z digest=sha256:7655349dcb5e6d265f255f7725f893c76ae6ae2124fee44f71a38a21384c36da

Observation 763b59fa-e211-41c2-96fb-65078c47324f · inbound

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages cites this paper.

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages AI Control: Improving Safety Despite Intentional Subversion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:31:14.785797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:31:14.785797Z digest=sha256:f52e1c190c2a7988fd4222885e5612ff5ec6a45e69e376e0b24cc30377564d8f

Observation 79439b91-bf29-490d-b568-a651c6184055 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-10T18:37:31.158825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-10T18:29:50.731038Z digest=sha256:1c04765b7c1213c2f71cb7b5a3e2e8f18aa3b1630226771cb73fdab1160b7498

Observation 9401af59-69f5-44bf-9c14-9fa2d3c67436 · inbound

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents cites this paper.

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:49:46.954375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:49:46.954375Z digest=sha256:a822f027c89ce2302ce23be572ff0f882edc3e62c2013c9124c1a4cddb4df7b9

Observation bd8cd22c-d745-4bb6-be33-69863c55f995 · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring AI Control: Improving Safety Despite Intentional Subversion

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-10T00:56:41.075656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:e11c36af2d002dd46e822cc0e4f20faf211f2e51154fe0587d9c24b838eab80f

Observation b6f05b4e-2134-4ea0-b1a8-32bde78b494b · inbound

GDM AI Control Roadmap cites this paper.

GDM AI Control Roadmap AI Control: Improving Safety Despite Intentional Subversion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:53:51.953667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:53:51.953667Z digest=sha256:d2e48785b15be1beafffae53afd17d419bf57d85486f7670a368ab47f05793b7

Observation 568a96db-126c-4bad-940b-7a630e2c6074 · inbound

Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives cites this paper.

Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives AI Control: Improving Safety Despite Intentional Subversion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:14:54.057238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:14:54.057238Z digest=sha256:2e7a0335040968abde481d13f8d55ca3dafa2f62f3553a8e82efe38049e070df

Observation 7aec30b4-4ff6-4a7c-ace2-d681684cb965 · inbound

Democratizing Agent Deployment Safety: A Structural Monitoring Approach cites this paper.

Democratizing Agent Deployment Safety: A Structural Monitoring Approach AI Control: Improving Safety Despite Intentional Subversion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:47:11.743438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:47:11.743438Z digest=sha256:edb1c040c7532bcc0d004c66ab44bbd7bf206980a5e3152b01c53ee955989932

Observation d38789b5-9a42-4d64-a416-229fdcf8eb46 · inbound

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D cites this paper.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D AI Control: Improving Safety Despite Intentional Subversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T12:49:57.527663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:49:57.527663Z digest=sha256:ac5e8f349b19a2abf3f40c3bab1c07ba9d3a30a25bbe6838d889bc201ca6a729

Observation 85258d82-146b-4d7f-a9a2-c0777fb33a3c · inbound

Code Monitor Red Teaming for Public-Test-Passing Code cites this paper.

Code Monitor Red Teaming for Public-Test-Passing Code AI Control: Improving Safety Despite Intentional Subversion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T09:12:25.068962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:12:25.068962Z digest=sha256:b1af163ebff5ab37505ddbc7a9854f2b07598127c09ad2bc285d7db1d0a07835

Observation 777c2d43-5634-4cbb-84bf-9866060db016 · inbound

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents cites this paper.

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T00:13:37.456860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:13:37.456860Z digest=sha256:9482c434a5cd07ea3d74b14c9ba5ee345c2c681c2983b3e5d919be1081f78bef

Observation 725e6c89-5343-40d3-b6a0-1ac7c27400f8 · inbound

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence cites this paper.

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence AI Control: Improving Safety Despite Intentional Subversion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T11:29:22.444200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T11:29:22.444200Z digest=sha256:8ad71e169a1680622cb59bb6c091697bf07597c7cc052673a7802a4c2a093b93

Observation ea43ee38-bee8-4063-a38c-75ce56b4ed8f · inbound

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents cites this paper.

Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents AI Control: Improving Safety Despite Intentional Subversion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:47:58.574890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:47:58.574890Z digest=sha256:4a4086301b3f0b0e3f62a4ff692def19afb3ecd102960e415db684a8218c8b0f