Pith. sign in

Paper Citation Record · LEDGER

Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:1908.04734.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04734 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:58.520812Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7f8677c2-5b30-41ac-a233-fa9c0bdc1e58 · inbound

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation cites this paper.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.520812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.520812Z digest=sha256:4eea4e3c5c9f6ffb3337372b3a43280886d0e4e60978dee7c9f562c9188b73e9

Observation d85b3b8b-e0fe-4a47-8372-e790844142c7 · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:01:10.371076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T14:07:44.717260Z digest=sha256:a8af01d6c61816f6572b061687da6f5395bb4133b08fdaedf79840beefdb13f7

Observation df910814-39e0-40df-9a3c-f5d7c16a6846 · inbound

AI Alignment via Incentives and Correction cites this paper.

AI Alignment via Incentives and Correction Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.686310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:51:31.357544Z digest=sha256:6c6ee2ddb3fb795c28f36261dedbeeb960c2aa9d2556f4520d0c1b7fde6728ee

Observation 20557408-85e8-4ef3-bea9-6014f7709f51 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.877863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:c6547d4d66e1be5431715c6bae349faa65a2c358522cc688079d88dece9df9ba

Observation 6450a2a7-a3cb-4be2-9690-80086c5e63eb · inbound

Hide to Guide: Learning via Semantic Masking cites this paper.

Hide to Guide: Learning via Semantic Masking Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.987148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T12:16:12.108715Z digest=sha256:0c45e0d8ced715a6258ac2561d571ebb77e9fac10d3d5b93df715ca8d10f28e9

Observation 40b214d8-0f47-4531-811d-777aff5d58fc · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:37:30.510880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:565339841fc0da3aacdd66ba6870809660986af57e7f1ea5a227dd2fcdd955f8

Observation 7a88bb85-e460-4ef1-b92f-2f796a95c04a · inbound

Reframing AGI Confrontation with Off Earth Autonomy cites this paper.

Reframing AGI Confrontation with Off Earth Autonomy Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:35:34.306248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:18:55.451342Z digest=sha256:7adda434fdfa5570eab58f72713d57479c387ce3628b5968d505744d78ab91ec

Observation f1e3aded-70be-4758-abd1-013060a60e25 · inbound

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting cites this paper.

Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.264264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T20:40:11.059493Z digest=sha256:3dea7a4d749b6075c1eb9c18042d12414bd55e76a7b0974b9cf97da76786f425