Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:44:44.310314Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2412.16325.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:44:44.310314Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:35:11.412956Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T11:35:12.372679Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b137f163-56a9-4707-82a7-a87ba70b3d47 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Unsolved problems in ml safety
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 48444c6d-dd7f-4327-a6b8-ece2bbb5690c · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Toward trustworthy ai development: Mechanisms for supporting verifiable claims
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5238e710-8330-473f-94cf-75ecb29db56f · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Deception analysis with artificial intelligence: An interdisciplinary perspective
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 17402991-8706-4046-a5e1-330864d132cb · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Unmasking the shadows of ai: Investigating deceptive capabilities in large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9de0a34c-9312-4e5a-926b-2b7024e769dd · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Human-level play in the game of diplomacy by combining language models with strategic reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 40e46928-84a1-43a1-b842-8bc1206a3565 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Hendricks, M.R
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e7614278-62d3-4b53-9339-8f2fdf2c71f5 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Collective constitutional ai: Aligning a language model with public input
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 14f108e6-a78e-4dc9-b621-26c06a763c70 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Constitutional ai: Harmlessness from ai feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9e72724f-ee72-4532-a3bf-77e75543b47e · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Truthful ai: Developing and governing ai that does not lie
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 36de0d67-2e0f-4dcf-a358-0a91332d7a02 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Language models represent beliefs of self and others
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d3dc3d51-5a6d-4ffd-9a91-64bd9cba8d47 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Premakumar, Michael Vaiana, Florin Pop, Judd Rosenblatt, Diogo Schwerz de Lucena, Kirsten Ziman, and Michael S
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ff581c8b-40d4-421b-a9f1-d956110cc66a · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Predicting vs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e01b7c9-552a-4324-8a96-967aaa4617c5 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7294465e-5088-4713-9a1a-950442ea7f3f · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Brethel-Haurwitz, Elise M
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9697ab3b-33d0-4106-b3e7-9d92fd04f44e · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Brethel-Haurwitz, Elise M
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2cd2f173-7e91-4d50-9852-68eaf49cec22 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Do altruists lie less? Journal of Economic Behavior & Organization, 157:560–579, 2019
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9270e5fe-131a-4111-99f3-7de25208f163 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap O’Connell, Shawn A
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5d9e9d4c-2ef2-4b3b-a71f-3c7fcb8d5f3e · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6bb15a6f-7c1f-4d01-ac4f-121e1c4b12c9 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Jonason, Minna Lyons, Holly M
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 386ceeb9-3847-4a9e-8ad8-43030e1d798e · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Towards empathic deep q-learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2bcabce1-db1d-422a-b3d7-290a2bd7e52b · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Modeling others using oneself in multi-agent reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ab1e7bda-58a3-48f1-9c74-df8bccc7fb8e · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Zou, Colin Raffel, Chris Callison-Burch, Yan Cao, Dzmitry Bahdanau, Gregory Diamos, and Jacob Steinhardt
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 312a6920-8674-41d1-a3f9-9d86e988c482 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Path-specific objectives for safer agent incentives
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 51aa5b73-1ead-4862-8ab3-b26d7c3b3d51 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Ortega, Elizabeth Barnes, and Shane Legg
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bc973fb2-c6b5-45cc-8d2c-81b2914c8c86 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Honesty is the best policy: Defining and mitigating ai deception
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d1e973be-c953-4d8b-961b-8dfa99b0c82c · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap The history and risks of reinforcement learning and human feedback
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5954792c-479b-4402-b0a2-a0a36e484fb0 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Deception abilities emerged in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4f6b15b4-5811-40ec-b669-2a3cb30c1726 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Xing, Hao Zhang, Joseph E
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8ec5ae5c-d92c-4b1d-97ae-c186b2a92284 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Physical-deception: An implementation of multi-agent deep deterministic policy gradient in pytorch to solve the physical deception environment from openai, 2023
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bfcaeb9b-4a37-4fa4-bc8a-bdc87e678ef1 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Multi-agent actor-critic for mixed cooperative-competitive environments
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5dfa4582-d139-4fe9-90ca-2106764c518b · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Human Compatible: Artificial Intelligence and the Problem of Control
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06c1aae-01bf-4c97-9480-4a8a436d9615 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Risks from learned optimization in advanced machine learning systems
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation de3ef56c-131a-44ec-9dc8-2079ec1f7e83 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Training a helpful and harmless assistant with reinforcement learning from human feedback
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 57346059-457a-4b6c-b1ba-8b23c2432b3a · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Ziegler, Tim Maxwell, Newton Cheng, et al
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f182b524-7d34-4a95-8516-bc3dc2c32418 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c995afbc-b5ea-47a3-b2a6-53afb6342ed8 · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Chain of thought prompting elicits reasoning in large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7cd99554-e09e-404d-8e41-8a34ea973aae · outbound
Towards Safe and Honest AI Agents with Neural Self-Other Overlap Only respond with the room name, no other text
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 64a96ae2-078d-45b4-b7f8-2ba2f1f6dd26 · inbound
Contemplative Artificial Intelligence Towards Safe and Honest AI Agents with Neural Self-Other Overlap
Reference 506
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.