Pith. sign in

Paper Citation Record · LEDGER

Misalignment from Treating Means as Ends

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2507.10995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10995 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:34:15.255851Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88265400-60cf-49f2-811c-ecbbfbd15654 · outbound

This paper cites Faulty reward functions in the wild.

Misalignment from Treating Means as Ends Faulty reward functions in the wild

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.644297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.445657Z digest=sha256:ff14c0643e17cd4a11fee507d308b2e03d24d7ab0a9b2a5e0810c39ed22ef32e

Observation 79375c00-6796-4b4a-af29-0df029ce93dc · outbound

This paper cites Potential-based shaping in model-based reinforcement learning.

Misalignment from Treating Means as Ends Potential-based shaping in model-based reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.634052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.564815Z digest=sha256:2e89f0c757a3d84507617f53afc86b612c8956c2f82d13299fb60a7661aa895f

Observation 9afe300b-8795-4d1f-a755-f6937a1eae89 · outbound

This paper cites Discrete dynamic programming.

Misalignment from Treating Means as Ends Discrete dynamic programming

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.624257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.683133Z digest=sha256:7ae31d1bc12bfb7061d07bc830519037483e81420f1ffa016fa125026abee0b1

Observation 7fb69216-d876-4780-a825-12f87281853f · outbound

This paper cites The superintelligent will: Motivation and instrumental rationality in advanced artificial agents.

Misalignment from Treating Means as Ends The superintelligent will: Motivation and instrumental rationality in advanced artificial agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.614442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:13.845517Z digest=sha256:bed97ebe6a269fb9a483153ac26d72d216a5615af2d9534144aef7437e9d33e2

Observation 1dda2239-00c0-4048-afc8-c3eddabe3697 · outbound

This paper cites Deep reinforcement learning from human preferences.

Misalignment from Treating Means as Ends Deep reinforcement learning from human preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:14.007127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:14.007127Z digest=sha256:d85c156eab7dcfe76cd1d74169f60ec758127b8649a28b294430ee8787dae64d

Observation f4e68bce-d562-4453-8159-dfc05d25b0db · outbound

This paper cites an unresolved cited work.

Misalignment from Treating Means as Ends Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:34:15.597861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.115875Z digest=sha256:9f442354f1faa331af0c14538886bd9ac924921c694d39354cfda059d54cf2b8

Observation ee0d7ebe-8c38-4f58-8b6e-871135d1b94f · outbound

This paper cites Exploration-guided reward shaping for reinforcement learning under sparse rewards.

Misalignment from Treating Means as Ends Exploration-guided reward shaping for reinforcement learning under sparse rewards

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.588281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.267435Z digest=sha256:77e879e1ce28fe97c949398c3725be1a4b0073b113f77b91c7ae25181bd5623d

Observation 59a02da4-7372-4db5-be5e-a2cb6fa5455a · outbound

This paper cites Dynamic potential-based reward shaping.

Misalignment from Treating Means as Ends Dynamic potential-based reward shaping

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.577539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.381757Z digest=sha256:a817204704b2259812a92e840208a158117e89c6c36c40d60288d1e4d358f7e8

Observation b7db942a-cc12-495c-bbd7-0c354d537e01 · outbound

This paper cites What is it you really want of me? generalized reward learning with biased beliefs about domain dynamics.

Misalignment from Treating Means as Ends What is it you really want of me? generalized reward learning with biased beliefs about domain dynamics

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.567489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.590020Z digest=sha256:2e47cae8186ffa0a1df8da47875cb8aea24716e27b7d7a91a0d87c49866ee9a2

Observation 3acafe36-0836-4621-9bd8-5d3d59776dc1 · outbound

This paper cites Reward shaping in episodic reinforcement learning.

Misalignment from Treating Means as Ends Reward shaping in episodic reinforcement learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.557064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:14.759732Z digest=sha256:eca94266ae8705e802a8d687c9a2a6ace6e1e5cc7630018c6f67622edf1e8d03

Observation a4264b5e-a178-477f-8c19-fd9242abc934 · outbound

This paper cites The off-switch game.

Misalignment from Treating Means as Ends The off-switch game

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:14.925169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:14.925169Z digest=sha256:6732d0be6c309762f75941e803db5ca94a2d9f1c475afaee64e28ea79e4355c4

Observation 13b04838-d379-48d7-a2a2-0906e7b465d3 · outbound

This paper cites Exposure and response prevention for obsessive-compulsive disorder: A review and new directions.

Misalignment from Treating Means as Ends Exposure and response prevention for obsessive-compulsive disorder: A review and new directions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.541323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.097323Z digest=sha256:9d3650f5d22d8d92ac067ba2bb318f9ecc5d2cfbe2a6875ddeeb316a80160082

Observation f96126d9-8445-4a64-b98d-6a9385688096 · outbound

This paper cites Teaching with rewards and punishments: Reinforcement or communication? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 37, 2015.

Misalignment from Treating Means as Ends Teaching with rewards and punishments: Reinforcement or communication? In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 37, 2015

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.531224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.181926Z digest=sha256:a8241255622f2a50ea9b54d322137a71e652430544053b7cf60d460a90279844

Observation 385de4a5-9185-4930-a313-3f4e81c05f78 · outbound

This paper cites People teach with rewards and punishments as communication, not reinforcements.

Misalignment from Treating Means as Ends People teach with rewards and punishments as communication, not reinforcements

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.521675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.185020Z digest=sha256:b6dacfb0626c27ddd639111ec445dbcc9567eab3478376288ba5a6d296d19b07

Observation 9385e40f-fbf7-4088-b447-795c1d6b83c6 · outbound

This paper cites Horn and Charles R.

Misalignment from Treating Means as Ends Horn and Charles R

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.512021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.188111Z digest=sha256:02886953aeacf36bb5a6ffddb58355c40eb66b127712a3748d52636b67954e54

Observation 09d72ca4-d9c5-4a43-8820-b8070f38aa5b · outbound

This paper cites Reward learning from human preferences and demonstrations in Atari.

Misalignment from Treating Means as Ends Reward learning from human preferences and demonstrations in Atari

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.502179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.191190Z digest=sha256:b080ddfef1d16dfb74c7f0a9de5a4a02c0058504f8fe461a6b35b39fe4089352

Observation 34e377bc-249d-4306-b0b4-3c007a454b7a · outbound

This paper cites Interactively shaping agents via human reinforcement: The tamer framework.

Misalignment from Treating Means as Ends Interactively shaping agents via human reinforcement: The tamer framework

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.492914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.193878Z digest=sha256:a3245e05764e9fbf17f199a139332fc1e6be2d71971e50226bc7294d60b19acf

Observation 3942d4f3-4cca-4c13-a3ca-f86360ca32ce · outbound

This paper cites How humans teach agents: A new experimental perspective.

Misalignment from Treating Means as Ends How humans teach agents: A new experimental perspective

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.483473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.196610Z digest=sha256:1a85061944a04f6397d0e77e0b411998c203564045821ccaf384f3c544652d55

Observation b05217fc-f97d-452d-9eea-6474cbc25f63 · outbound

This paper cites Models of human preference for learning reward functions.

Misalignment from Treating Means as Ends Models of human preference for learning reward functions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.199300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.199300Z digest=sha256:42d1ed20b5f11046d82195f929024cbfaace06a9b243ec2c8998931a7623ae32

Observation 3db4851c-dadd-4dbc-b68b-79336cf351f3 · outbound

This paper cites Learning optimal advantage from preferences and mistaking it for reward.

Misalignment from Treating Means as Ends Learning optimal advantage from preferences and mistaking it for reward

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.473562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.202726Z digest=sha256:5194d8f7518007619feadedb1273ba5e614a15e3046a2ad515c1d3df27d5928c

Observation cd331bdb-b089-4b25-9a86-9a209c9fd58f · outbound

This paper cites BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping.

Misalignment from Treating Means as Ends BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:34:15.299575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.205426Z digest=sha256:06955add1abf5d3cb596073668e7e8b64643f72267afb5318b9587414529279a

Observation 2c0656be-52b7-4a18-9432-4a76b93da383 · outbound

This paper cites an unresolved cited work.

Misalignment from Treating Means as Ends Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:34:15.463949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.208350Z digest=sha256:5c0d6a0cf97a82f9af9ecfdb9f1f768edbfa513e22adf62c8f28dfc5cc949620

Observation 6383da11-7435-4a03-bc2a-2451044abaae · outbound

This paper cites Choice between partial trajectories: Disentangling goals from beliefs, 2024.

Misalignment from Treating Means as Ends Choice between partial trajectories: Disentangling goals from beliefs, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.454002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.211065Z digest=sha256:8d94ff68d702bca45ccd7a7af0a56eb3368562c4abed2b6a5b461bbf07f1bab5

Observation 0400189d-aec9-4f9d-8d28-b305e0bf115d · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping.

Misalignment from Treating Means as Ends Policy invariance under reward transformations: Theory and application to reward shaping

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.444941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.214008Z digest=sha256:7a4ede4fa1460296bd84446bf34288e8fd96bb95ab03a7b29b9c6fc51fa469e0

Observation 457b866a-e7c5-48ae-bcb4-c4308c19a0ab · outbound

This paper cites Anthropic’s new AI model threatened to reveal engineer’s affair to avoid being shut down.

Misalignment from Treating Means as Ends Anthropic’s new AI model threatened to reveal engineer’s affair to avoid being shut down

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.435244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.217137Z digest=sha256:c586af6b7eaa6791068eeef1abbb0a9ab72f4b2c4e8a72e3cdfbe4ad51b2473d

Observation 2b3f10f5-6b0a-408d-84b1-8e72cf5ce63c · outbound

This paper cites The basic AI drives.

Misalignment from Treating Means as Ends The basic AI drives

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.424662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.219920Z digest=sha256:2d744e76311c2f310eeac56f9169a2fb17f4f9e989a6c3c72c784d0332ec67a9

Observation 8d5bda20-19af-46a3-a813-f9933b54f384 · outbound

This paper cites Learning to drive a bicycle using reinforcement learning and shaping.

Misalignment from Treating Means as Ends Learning to drive a bicycle using reinforcement learning and shaping

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.413722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.222968Z digest=sha256:b0c4919019b66e4efe5fdefc8c7f224c072782959015fd4ff1284c42895e05cb

Observation 45e6ca2a-9107-4ec4-9254-7dea463fec9a · outbound

This paper cites AI is learning to escape human control.

Misalignment from Treating Means as Ends AI is learning to escape human control

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.402944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.225939Z digest=sha256:e6924b6a2140a8de108a1d64a8158f986b9900c719b9309ee9727cdb39c33682

Observation 4e66a997-67d8-41d3-be94-1999cabb3f9e · outbound

This paper cites Human-compatible artificial intelligence, 2022.

Misalignment from Treating Means as Ends Human-compatible artificial intelligence, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.393424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.229124Z digest=sha256:a7088cd4d0de3dfe32b3f410d8fb77b543fe5f2216edb727743ce8f5de082478

Observation ae0afa82-5d18-4aab-948f-d298f630f731 · outbound

This paper cites Artificial intelligence: a modern approach.

Misalignment from Treating Means as Ends Artificial intelligence: a modern approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.383047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.232151Z digest=sha256:44df9e13481f8e7ae0663bffb32e58002b9a2170606d46ad8d5eec182b29c24e

Observation 3c1f24bb-19f9-4cfa-ba51-389bb9f34d54 · outbound

This paper cites Where do rewards come from.

Misalignment from Treating Means as Ends Where do rewards come from

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.372851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.235350Z digest=sha256:6529b6d183990bdce67477c24a4381ef8feaac4d533152debb78cee4fd5287a4

Observation 87166d99-fb94-4a85-aa04-c247f9d643e1 · outbound

This paper cites Intrinsically motivated reinforcement learning: An evolutionary perspective.

Misalignment from Treating Means as Ends Intrinsically motivated reinforcement learning: An evolutionary perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.238058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.238058Z digest=sha256:cfdc890e7b132570c539a39525e613b83178bd2f04230074834129ede5b84517

Observation 9bff7598-8344-439e-b985-d32ea35a874f · outbound

This paper cites Corrigibility.

Misalignment from Treating Means as Ends Corrigibility

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.357143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.240721Z digest=sha256:0a03462df4e132e871d66599ee95d3cd0a7bdc82dcf27d32de1b45d5ff7b05de

Observation 028f26ba-7548-48ff-958d-1dcd59cf1be7 · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Misalignment from Treating Means as Ends Reinforcement learning: An introduction, volume 1

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.243611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.243611Z digest=sha256:ca5db221de7644fa9e8d02280df9dda1f4f1ac42c04db4be4910f96be18ae513

Observation 6f36d324-050b-43f8-8cc8-0a64e926a041 · outbound

This paper cites Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance.

Misalignment from Treating Means as Ends Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.340391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.246831Z digest=sha256:4582233fab0394aba3d8f200bc85cd734b5b70c5a6ec369e7016f1ac67fa4789

Observation 16481d20-58c4-41ae-a253-d4e787523bb0 · outbound

This paper cites a ngberg, Mikael B \.

Misalignment from Treating Means as Ends a ngberg, Mikael B \

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.329928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.250509Z digest=sha256:2c36fd972dc29d43340265b84da2eb73b6f13251fa82f2541edcc123d99fc422

Observation 63e2764a-c98a-46d2-a69c-2090a722c3fb · outbound

This paper cites Principled methods for advising reinforcement learning agents.

Misalignment from Treating Means as Ends Principled methods for advising reinforcement learning agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:34:15.319816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:34:15.253266Z digest=sha256:7f82aa4aa2ee3a08ad26e042bc0360723e7305228c35ccb77d9cdb3e5ebf104a

Observation f12742d6-154c-4225-adfa-8cd93a2728e0 · outbound

This paper cites Reward Shaping via Meta-Learning.

Misalignment from Treating Means as Ends Reward Shaping via Meta-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:34:15.255851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:34:15.255851Z digest=sha256:bddf984966a2d031dc1ce92076f4bbe385f76e354b04b658b05515b9e53ef938

Pith citing papers

No inbound Pith citation observations are available.