Pith. sign in

Paper Citation Record · LEDGER

The Alignment Problem from a Deep Learning Perspective

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2209.00626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.00626 v8

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:17:59.786686Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

63
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fe5749d7-f45c-4d5b-ac2d-231543115114 · inbound

Scaling Laws for Reward Model Overoptimization cites this paper.

Scaling Laws for Reward Model Overoptimization The Alignment Problem from a Deep Learning Perspective

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:04:53.341962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:04:53.129737Z digest=sha256:f09465fa437f98d94dbe0833c2a73a3ac6f187953e5afbe42f5a09bb269ddacc

Observation becee435-783a-4ffe-8552-f9f561b98347 · inbound

Sparse Autoencoders Find Highly Interpretable Features in Language Models cites this paper.

Sparse Autoencoders Find Highly Interpretable Features in Language Models The Alignment Problem from a Deep Learning Perspective

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T06:44:02.226998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-24T06:42:23.274826Z digest=sha256:237f645455a49f0608de4eb6d7dd1793b0193c26545d8dbedc9e7eaa1fae0feb

Observation 80067b9f-aa3c-4a40-b1c4-1dc44c00c48a · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey The Alignment Problem from a Deep Learning Perspective

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:52.783423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:c7be894ce0f0d8b0f91b46df22c269965be6c2dbfb43ba4dd35f0a4970bc1270

Observation 3fdd8724-59e7-4815-a791-096d637ecd14 · inbound

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models cites this paper.

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models The Alignment Problem from a Deep Learning Perspective

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T13:15:10.700429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T13:15:10.632115Z digest=sha256:9195569a5e1534493008390c73bc7ad396ea47762edee471cc425815c5b7db3f

Observation 7f903738-a102-44b9-93bd-b147b5d04099 · inbound

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models cites this paper.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models The Alignment Problem from a Deep Learning Perspective

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T14:43:30.216305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:78ab7eef671b5f35baea2ad183957c5433c7e512d0451523481105cd00c99792

Observation ba9ab7f7-df82-4e25-b00d-d10996e49f89 · inbound

Safety case template for frontier AI: A cyber inability argument cites this paper.

Safety case template for frontier AI: A cyber inability argument The Alignment Problem from a Deep Learning Perspective

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T22:01:11.483062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T22:01:11.483062Z digest=sha256:f6fd05b7a5a16672b81034ec67443e0068ca1bae7c5038f361050656f05fbd24

Observation 2d055cf9-5346-4e77-be64-37b7d71acfff · inbound

Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset cites this paper.

Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset The Alignment Problem from a Deep Learning Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:52:34.659958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:52:34.659958Z digest=sha256:e12af0e11fc831b8e5b0116968d7eefef5403781e0a5761fa84e7d30b3a743f7

Observation 88cbea12-254c-4fed-b589-e85e52bdf171 · inbound

Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies cites this paper.

Can an AI Agent Safely Run a Government? Existence of Probably Approximately Aligned Policies The Alignment Problem from a Deep Learning Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:45:49.995737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:45:49.995737Z digest=sha256:ce4a83ec18436535f9ea660b2e392af9ca8fc4a9abf331f04ef8090357f4a441

Observation 9275644b-34da-4f51-bbb5-5d4707c86caa · inbound

Open Problems in Machine Unlearning for AI Safety cites this paper.

Open Problems in Machine Unlearning for AI Safety The Alignment Problem from a Deep Learning Perspective

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:10.486442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:10.486442Z digest=sha256:7e510f4fb110bce267232d89d2e417bb90fdedca56d65c20b01e9c819c1a8428

Observation 35b96a1d-149d-4731-a7c2-12cf6d6dc780 · inbound

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development cites this paper.

Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development The Alignment Problem from a Deep Learning Perspective

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T05:34:03.897439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T05:34:03.897439Z digest=sha256:b6109e0dfff1e87803bd83c9d6101f05fb159b6d08173fd4edc5687941277698

Observation e0b1aa67-3e4e-4af0-b51d-799c32909168 · inbound

Compromising Honesty and Harmlessness in Language Models via Deception Attacks cites this paper.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Alignment Problem from a Deep Learning Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.420807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.420807Z digest=sha256:a9b325c7658281006042601f1af478a0432e76170aea8a74132a3d8b44fbb08b

Observation 69e0e3d4-2040-4c0b-8b68-1c21e8c6a2e8 · inbound

From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs cites this paper.

From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs The Alignment Problem from a Deep Learning Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:01.897413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:28:01.897413Z digest=sha256:1b176faa67fb85905ac3fd0cc4f0ad4da30872a67db74d95938e1f3ccacb867b

Observation 8e0d26c6-80e1-46c1-9551-0cc53537e299 · inbound

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components cites this paper.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components The Alignment Problem from a Deep Learning Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.158123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.158123Z digest=sha256:1f7a3efb3df52abab05f07e8704c912fbef176446936cff11ec6e95e168faff8

Observation de07081a-edd7-4381-b169-5490e630346d · inbound

Will artificial agents pursue power by default? cites this paper.

Will artificial agents pursue power by default? The Alignment Problem from a Deep Learning Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.329174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.329174Z digest=sha256:55f5e4bccc3a8e06969d70f02e6cdbe26b555bfae1b87a3f55354bbcb21a0d69

Observation d9f40715-cc08-47b5-80ec-c53063625578 · inbound

Deontically Constrained Policy Improvement in Reinforcement Learning Agents cites this paper.

Deontically Constrained Policy Improvement in Reinforcement Learning Agents The Alignment Problem from a Deep Learning Perspective

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:59.350948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:51:59.350948Z digest=sha256:2defda69f870a33237045fb5d9a1113425d44ad1ee39e9c418e84c32976ebf71

Observation 972a95bd-1328-4030-a171-307a97df5f0c · inbound

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) cites this paper.

Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack) The Alignment Problem from a Deep Learning Perspective

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:57.474356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:57.474356Z digest=sha256:6864650fcd2241d95c84851c89d5e7d0f573e0adca4e6bdc8af71343a2b3e41c

Observation c8ab42da-82a3-43ea-a7ff-b988c3df64ca · inbound

Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective cites this paper.

Evolving Prompts In-Context: An Open-ended, Self-replicating Perspective The Alignment Problem from a Deep Learning Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:29:28.487492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:29:28.487492Z digest=sha256:432b5f18aacc13f099ee6aeeed54b9e39baf8182a2914bd2161cd89a76b2fc27

Observation c32a77d8-2d1a-4f83-8f61-4c5584cf15a4 · inbound

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models cites this paper.

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models The Alignment Problem from a Deep Learning Perspective

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:22.984820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:22.984820Z digest=sha256:5ccd3fce22865faeaf4ec5925a4dac4a5d8d1e44e27c6ba24cf596b15a080d10

Observation 29342e78-4d90-4680-9744-bc205f44e80d · inbound

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact cites this paper.

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact The Alignment Problem from a Deep Learning Perspective

Reference 219

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:11.445678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:11.445678Z digest=sha256:cbc3421af7a0ff331f4802b42efaf1586e77558059295dcfae490363099a9a18

Observation efb17172-59eb-40dd-a6ba-dc2aaa5117ae · inbound

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language cites this paper.

Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language The Alignment Problem from a Deep Learning Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:18:15.417317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:18:15.417317Z digest=sha256:2afe6d4622027bfdd61560a5a2b7d3922f4d97699924297f0bf6adb7121f3e40

Observation 7579f224-aec7-4c7a-b527-bf7b68bc9725 · inbound

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework cites this paper.

Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework The Alignment Problem from a Deep Learning Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:36.793291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:39:36.793291Z digest=sha256:77c10ed7e70e1286462105b1ed4133e7a7453c8e183b86f592234abbf47014a5

Observation 9290f995-4ac4-4161-91df-229157a44fe6 · inbound

On the Inevitability of Left-Leaning Political Bias in Aligned Language Models cites this paper.

On the Inevitability of Left-Leaning Political Bias in Aligned Language Models The Alignment Problem from a Deep Learning Perspective

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:41:34.328537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:41:34.328537Z digest=sha256:f25427052e657231ec1dd796a68e75f02548ffb3978fc0308b30eb802dba5e90

Observation 33df2d56-405c-4584-b48d-f0f05d6b4a74 · inbound

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns cites this paper.

Mechanistic Exploration of Backdoored Large Language Model Attention Patterns The Alignment Problem from a Deep Learning Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T18:42:02.769943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:42:02.769943Z digest=sha256:586f2a1aa7ab637b74dc91c51e9f9766d14038279b8af3c5fc14e587f16866f4

Observation e369c779-39c4-4540-be9b-f88018797938 · inbound

Human-AI Complementarity: A Goal for Amplified Oversight cites this paper.

Human-AI Complementarity: A Goal for Amplified Oversight The Alignment Problem from a Deep Learning Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:22.518530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:22.518530Z digest=sha256:d4fcfe0500beb2d09e763f1b1dd51392df482365615e1e0a061872698c5785f6

Observation bad157e3-c7d8-4752-9875-64b18cc0ca4e · inbound

Language Model Circuits Are Sparse in the Neuron Basis cites this paper.

Language Model Circuits Are Sparse in the Neuron Basis The Alignment Problem from a Deep Learning Perspective

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T06:35:33.921059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:35:33.921059Z digest=sha256:a14dcd2d1009b4f27c75c91dbc8c769515f7bb6c3bd5f5c55c5dcc30af0f6d4c

Observation 2b2b6fd6-4280-4d26-866d-b1890dbd293a · inbound

An Onto-Relational-Sophic Framework for Governing Synthetic Minds cites this paper.

An Onto-Relational-Sophic Framework for Governing Synthetic Minds The Alignment Problem from a Deep Learning Perspective

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:59:53.036090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T08:57:58.369793Z digest=sha256:ca722043db791a90f3eeb03a1a5befb59038d9e64ad9949f3d5652cff51994c4

Observation e1c90d48-7394-4748-8c3a-a1470b975659 · inbound

Framing Effects in Independent-Agent Large Language Models: A Cross-Family Behavioral Analysis cites this paper.

Framing Effects in Independent-Agent Large Language Models: A Cross-Family Behavioral Analysis The Alignment Problem from a Deep Learning Perspective

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:01:25.299330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T18:01:02.298383Z digest=sha256:94d6462d78b9c43453d873bb6bf1841696ef3a95ee9afffbe9a1ca12e7efa1dd

Observation 47bb1925-c2b9-463a-800c-14aee472f533 · inbound

Safety, Security, and Cognitive Risks in World Models cites this paper.

Safety, Security, and Cognitive Risks in World Models The Alignment Problem from a Deep Learning Perspective

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:38:22.257456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T22:35:46.126714Z digest=sha256:6a214fb178571b775b5a5e34078fb1e1dda59ddcbc97f0112d8aa3e98416a929

Observation 8ba639fd-1849-487f-bc2b-ae7d9f8abb23 · inbound

Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry cites this paper.

Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry The Alignment Problem from a Deep Learning Perspective

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:08:09.498955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T19:08:06.518978Z digest=sha256:c88a37dfe87bbdb832c5878a33a65de733acf557df5247db9e47df5d0928567b

Observation 1c211139-d19d-4576-a4c8-b30f3669bb99 · inbound

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories cites this paper.

Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories The Alignment Problem from a Deep Learning Perspective

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:10.083360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T05:48:44.687520Z digest=sha256:be96b4823f73202e8266eb003d3192acd60a212d126ea86adcaa851c108b0003

Observation b9088d91-92c3-476b-b250-2d987df8f3f9 · inbound

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem cites this paper.

Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem The Alignment Problem from a Deep Learning Perspective

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:04:17.665895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T23:02:34.564375Z digest=sha256:4d213bd94e8e5f6a6e5719b2aa10f2876a5874327794c899439a8a046c993387

Observation dd349a6b-154b-45e5-8c58-c9bdb20405b2 · inbound

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training cites this paper.

Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie Training The Alignment Problem from a Deep Learning Perspective

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:32:24.282885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T06:30:51.812541Z digest=sha256:8c95467762b306e1609b0e534a207c237c6c74dad83e921ab89465819a1bcbda

Observation 5e13343c-e132-49c4-a373-33625775d819 · inbound

Who Owns This Agent? Tracing AI Agents Back to Their Owners cites this paper.

Who Owns This Agent? Tracing AI Agents Back to Their Owners The Alignment Problem from a Deep Learning Perspective

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T17:28:47.944759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T17:26:23.647993Z digest=sha256:1d9e9ef86845b041e9630ccabace93182ba66663afdddf32556957d4cf7bb4b9

Observation 17b2db02-300d-4e4e-bc92-da4f815dc981 · inbound

Understanding Goal Generalisation in Sequential Reinforcement Learning cites this paper.

Understanding Goal Generalisation in Sequential Reinforcement Learning The Alignment Problem from a Deep Learning Perspective

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:20.720789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-25T04:49:50.034743Z digest=sha256:26c6994c196c025fa4d46315699e356f01c6779a5b0ef6aab0ebb2d956883bb1

Observation eec751e3-a391-4e71-8792-7a21e330318b · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model The Alignment Problem from a Deep Learning Perspective

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.156294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:dc4d7c0ffbe6798237061518c693bbe97f8ed0ec9557ab1c1ecb75b9e1c6fb56

Observation cba9c881-9c24-4a4d-a12b-239c951161cf · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model The Alignment Problem from a Deep Learning Perspective

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:d8591d5ad5a8de10ded909f415f97ccbcfb3e9acbaa01111702887c3ab15cca0

Observation 5e41e396-90c9-4ddf-9ede-91913fab68b7 · inbound

Misaligned AI as a New Insider Risk cites this paper.

Misaligned AI as a New Insider Risk The Alignment Problem from a Deep Learning Perspective

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.843497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T23:20:11.068720Z digest=sha256:00ddabb3befaedc124443da5dc8bda87f1ea368e064400ba40b2a61a99de5c24

Observation f216dbdc-5887-40b9-8baf-b10e1ef4225c · inbound

Enhancing AI Interpretability with Localised Architectures cites this paper.

Enhancing AI Interpretability with Localised Architectures The Alignment Problem from a Deep Learning Perspective

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:37:22.506038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T20:21:00.330552Z digest=sha256:9face9fae09ac8c29531b2aba3b38b6a34283d76c0c718491535b58d269416b7

Observation 1f89b684-e6f5-48a8-91a1-ef694013a808 · inbound

The Agentic Web Requires New Normative Infrastructure cites this paper.

The Agentic Web Requires New Normative Infrastructure The Alignment Problem from a Deep Learning Perspective

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-28T02:41:32.525097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T11:55:10.777827Z digest=sha256:02114a247c84642fda707075792d137a8282eba013449386a1070ecb528f06db

Observation f7ed8e7a-2c63-4b96-9d4a-aaf0ffaec409 · inbound

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing cites this paper.

Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing The Alignment Problem from a Deep Learning Perspective

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.660392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T01:28:38.810889Z digest=sha256:00cb5da22d355d4e28389b803b903b2326ff1d8fb608edea32d003181d340edf

Observation 48be50e9-231b-48f4-a39b-3da11af23294 · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor The Alignment Problem from a Deep Learning Perspective

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.207972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-30T06:53:19.737408Z digest=sha256:900959f17032356e9ea1f9cdc9d2a983c10cee290b628598a6dc8088b571d074

Observation 8aa9d0fd-2017-4d11-8eaf-cfb0e2b86185 · inbound

Safety from Honesty in a Disinterested AI Predictor cites this paper.

Safety from Honesty in a Disinterested AI Predictor The Alignment Problem from a Deep Learning Perspective

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T17:02:19.189814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T17:02:19.189814Z digest=sha256:b750297e472f4672d7c8e163014470ceecb56aec8b93e18ac331850d1e455d43

Observation d36bd4a5-e4af-4fdc-be91-b373b5c89ff1 · inbound

A Scalable Approach to Evaluating Moral Sensitivity in LLMs cites this paper.

A Scalable Approach to Evaluating Moral Sensitivity in LLMs The Alignment Problem from a Deep Learning Perspective

Reference 135

Resolution
unresolved
no resolver link, observed 2026-07-12T05:44:33.099337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T05:44:33.099337Z digest=sha256:f8e5abf31082f9a2d5294858a952bb83fe41698c2d85b0ac34e49d6dab83b9e5

Observation 610c9187-8023-460f-92c3-0dab0ec79e0c · inbound

User identity conditions moral wrongness ratings in non-reasoning large language models cites this paper.

User identity conditions moral wrongness ratings in non-reasoning large language models The Alignment Problem from a Deep Learning Perspective

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:46:01.555826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-09T05:37:58.875536Z digest=sha256:fe2fac749ff12a245229d9a3caf8d213e57b009fd5ebca4f31a2eddfad41301e

Observation 48ff409a-ace0-4170-99e4-352dff6e8235 · inbound

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring cites this paper.

Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring The Alignment Problem from a Deep Learning Perspective

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:56:41.011024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-10T00:52:47.537142Z digest=sha256:0fd56c6cacfe074bac5a65b5a0c8adb597b535a5c4d8e8d61e3896ab34e05e71

Observation 53f8cf84-4132-43b9-b21f-e9005a1172ba · inbound

Hardware Mechanisms to Dynamically Throttle AI Performance cites this paper.

Hardware Mechanisms to Dynamically Throttle AI Performance The Alignment Problem from a Deep Learning Perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T16:16:14.799670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:16:14.799670Z digest=sha256:294384a8004670ce37c470d257cab5fd8f5d3f49827f1246fa7eddb2c8f5ec16

Observation f26ac867-d237-4dfe-8199-f00c18ea6625 · inbound

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF cites this paper.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF The Alignment Problem from a Deep Learning Perspective

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:41.660778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:41.660778Z digest=sha256:fd2999b823ad9a44a840a07564cc017b58933fef92daa24f7bc3c47b308d68e9

Observation 8a82ba0b-24ce-44b2-bb39-f914401fc253 · inbound

Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives cites this paper.

Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives The Alignment Problem from a Deep Learning Perspective

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T05:36:11.664750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:36:11.664750Z digest=sha256:848b8091946a5453f527dfd24717729fd42bf446267f5ee61dbaab830fded34d

Observation 1814b975-231d-4f06-9a20-ce47150dea6f · inbound

Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction cites this paper.

Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction The Alignment Problem from a Deep Learning Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T11:13:43.700315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T11:13:43.700315Z digest=sha256:c773cf194d23c91ad19ef5dd91a75af7e726721f6423c9b1f499edaaedca37bb

Observation 745f7d37-64b8-4d71-ac37-90849517e8c3 · inbound

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes cites this paper.

Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes The Alignment Problem from a Deep Learning Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:17:59.786686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:17:59.786686Z digest=sha256:0c81b1813221a786829ba9dec15ee5bb4bd3ff119a2d142cd2e3a2bb53c14f5a