Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:16:02.443284Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.21518.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T07:16:02.443284Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
16 of 16 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 727bf8bf-af72-4d5e-8520-818ce48fb764 · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd254004-2404-4b60-834f-e030f56bfef1 · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Alignment faking in large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3b73b7-7a06-47cc-8ae7-adb020cc92d9 · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5048bffa-aef4-4ffd-ad6d-f7f24bdf424a · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e24397f-67c3-4b5f-90be-a7d332231bde · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Frontier Models are Capable of In-context Scheming
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b08e2da-563c-40b3-a4f7-76b323766029 · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.18653/v1/2023.findings-acl.847
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab57785a-77a0-487e-865e-b31a9c51fd0a · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://proceedings.neurips.cc/paper_files/paper/2023/ hash/ed3fea9033a80fea1376299fa7863f4a-Abstract-Conference.html
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f1a18fc-b149-4f32-8c72-bd94008c3f9a · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e31b94f-e9fc-4cb8-ba5f-57f9a7dccafc · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.18653/v1/2024.findings-acl.624
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e547ab-049c-4974-b733-6a5a693b180a · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Alignment faking in large language models
Reference 1923
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2141d33e-16bc-4322-b92f-3a77324a2b8d · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://doi
Reference 1960
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9aac6618-fef5-47ba-bcfd-16ad0c6d6c1e · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.1201/9780429246593
Reference 1994
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7930f030-0e8e-4504-b9f0-dc94e3372f7a · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Constitutional AI: Harmlessness from AI Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786d3824-a9e1-4d4a-80fc-22dd5c6e814c · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.1145/3605764.3623985
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4db23d-2d30-4c91-8bf3-1ad6c4e352de · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 97091a5177d8dc64b1da8bf3e1f6fb54-Abstract-Datasets_and_Benchmarks_Track.html
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c2b679-5937-4d5c-b354-19eda83950bc · outbound
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.