Pith. sign in

Paper Citation Record · LEDGER

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2607.21518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21518 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:16:02.443284Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 727bf8bf-af72-4d5e-8520-818ce48fb764 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.349786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.349786Z digest=sha256:56fa136a39a95193aaa2fbc9575d85c41d7b9c609f38070a321e34fbddb5e079

Observation fd254004-2404-4b60-834f-e030f56bfef1 · outbound

This paper cites Alignment faking in large language models.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Alignment faking in large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.388415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.388415Z digest=sha256:e111b2d90959c383ef668c6b413ebbadca2094daf9e839d6f0577e2f26a5eaa1

Observation 0d3b73b7-7a06-47cc-8ae7-adb020cc92d9 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.400998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.400998Z digest=sha256:37a76455d0dbfd089a1de06a36e8aa41bf2339563096c901a04c3281e37a44c6

Observation 5048bffa-aef4-4ffd-ad6d-f7f24bdf424a · outbound

This paper cites Auditing language models for hidden objectives.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.407247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.407247Z digest=sha256:064318e346c84ed15488a66beee73611a0c835cbe405d66caa636c3929e90aad

Observation 1e24397f-67c3-4b5f-90be-a7d332231bde · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Frontier Models are Capable of In-context Scheming

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.419877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.419877Z digest=sha256:95e5c1ba1e61e3721d234b77c3c4b09c2fef78de23149e376c0a74cd3fc33713

Observation 2b08e2da-563c-40b3-a4f7-76b323766029 · outbound

This paper cites doi: 10.18653/v1/2023.findings-acl.847.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.18653/v1/2023.findings-acl.847

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.426157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.426157Z digest=sha256:528d0dfa526172a691b378a2879eb5a25c3056a5fda386af640dc22f2e7e965f

Observation ab57785a-77a0-487e-865e-b31a9c51fd0a · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2023/ hash/ed3fea9033a80fea1376299fa7863f4a-Abstract-Conference.html.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://proceedings.neurips.cc/paper_files/paper/2023/ hash/ed3fea9033a80fea1376299fa7863f4a-Abstract-Conference.html

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.431628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.431628Z digest=sha256:b773d2d5eb1542741a036fb49d646a18503ebb231cd5f6106bcb73650bcd4242

Observation 0f1a18fc-b149-4f32-8c72-bd94008c3f9a · outbound

This paper cites Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.438374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.438374Z digest=sha256:8c545c4b9c7d8fae28e096b1410cc84d68eb594d2450b3d7130db539ef0baf3a

Observation 6e31b94f-e9fc-4cb8-ba5f-57f9a7dccafc · outbound

This paper cites doi: 10.18653/v1/2024.findings-acl.624.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.18653/v1/2024.findings-acl.624

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.443284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.443284Z digest=sha256:11498b0f55976bf88eee3bf63ab6675ced13bc64b65fc38619aec26e09316450

Observation b6e547ab-049c-4974-b733-6a5a693b180a · outbound

This paper cites Alignment faking in large language models.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Alignment faking in large language models

Reference 1923

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.381976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.381976Z digest=sha256:efec11a536c57e43ad8d0121086ac12ead9e536ebd7120c367fa13b1fd2d9a2f

Observation 2141d33e-16bc-4322-b92f-3a77324a2b8d · outbound

This paper cites URL https://doi.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://doi

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.364315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.364315Z digest=sha256:1291253a9bd091d2e9dcba8fd874dd4795286e56af7af9425df7910e37bc553a

Observation 9aac6618-fef5-47ba-bcfd-16ad0c6d6c1e · outbound

This paper cites doi: 10.1201/9780429246593.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.1201/9780429246593

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.376566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.376566Z digest=sha256:51a235141a5e954cfa1c7434ecff553800d3bbabff21e7a84f279e13f7ccf3f4

Observation 7930f030-0e8e-4504-b9f0-dc94e3372f7a · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Constitutional AI: Harmlessness from AI Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.357924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.357924Z digest=sha256:c8f422af81d3547545355cd81d2ca81c4368532f255ffe4860946a91ff7d1ca7

Observation 786d3824-a9e1-4d4a-80fc-22dd5c6e814c · outbound

This paper cites doi: 10.1145/3605764.3623985.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation doi: 10.1145/3605764.3623985

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.395357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.395357Z digest=sha256:296e7e61ce1197e5e916fcbae32c2a328538e36bb4f7a9bafbf26dabdbbf3d87

Observation 3a4db23d-2d30-4c91-8bf3-1ad6c4e352de · outbound

This paper cites URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 97091a5177d8dc64b1da8bf3e1f6fb54-Abstract-Datasets_and_Benchmarks_Track.html.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ 97091a5177d8dc64b1da8bf3e1f6fb54-Abstract-Datasets_and_Benchmarks_Track.html

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-01T07:16:02.370023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.370023Z digest=sha256:273d0ee5aaad7799bae30aa7a8f64ace6eb5145067e4a67e6c7f3d87548b2991

Observation 33c2b679-5937-4d5c-b354-19eda83950bc · outbound

This paper cites Auditing language models for hidden objectives.

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation Auditing language models for hidden objectives

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T07:16:02.413345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:16:02.413345Z digest=sha256:62b9d1da032248752586606c6ccd9d5dd4eb3ea6ff02bb0ce79e378ae061add7

Pith citing papers

No inbound Pith citation observations are available.