Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:12.792191Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2508.21443.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:12.792191Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
33 of 33 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 99c67e3f-93a2-4f96-ad99-6f4c30902898 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Human- level control through deep reinforcement learning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f849639-2619-4e83-a01b-bd67c1b4310c · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Benchmarking deep reinforcement learning for continuous control,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e88dae6b-4ef8-43df-a770-9381aeba42f5 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Research priorities for robust and beneficial artificial intelligence,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8b4efd2d-524a-437c-aca8-8bd0763ecbee · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Concrete Problems in AI Safety
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b1a652-9c0a-4afd-8728-4f4773efaf9d · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5870618-a950-4566-b73e-3cc5c7324a69 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b4b4b84-1086-4b56-a809-c853c4c5293f · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Bertsekas, Reinforcement Learning and Optimal Control
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ee9982-9bc0-4cee-8352-15813493aff4 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Reinforcement learning with non-ergodic reward in- crements: robustness via ergodicity transformations,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5b91cb89-0bfe-46c0-b32c-929af9a18041 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Linearly-solvable markov decision problems,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9a1d85f1-1d73-4396-985a-fb0d9b0e7f7e · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A theory of regularized Markov decision processes,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e77cda3b-c52f-4351-84de-5f764dbd789d · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Approximate modified policy iteration and its application to the game of tetris
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3b2f7ae1-5ec8-449b-90df-fff9d6b77d6b · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Munchausen reinforcement learning,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fe9d1c02-fa7a-405a-aacf-8bc83e35b612 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Maximum a Posteriori Policy Optimisation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37475928-ea80-4620-b75d-a97fed392256 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Leverage the average: an analysis of KL regularization in reinforcement learning,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d340f80d-c6b6-4000-a8ef-b7a9cd3b862a · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Incremental multi-step q-learning,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9d80eb6e-7a80-493b-94db-6a6d18fb2dd8 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Multi- step reinforcement learning: A unifying algorithm,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d44e8e4c-aef5-45a9-9d6d-861a865a1590 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Multi-Bellman operator for convergence of $Q$-learning with linear function approximation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 176002bd-bd9d-46f3-8e54-c5231c0efe9b · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A novel multi-step q-learning method to improve data efficiency for deep reinforcement learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9ead2d0e-8c2c-4029-a97a-f9adbda5bed2 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Rainbow: Combining improvements in deep reinforcement learning,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8b2937fa-e06b-47d8-9fb3-80102323141c · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Policy invariance under reward transformations: Theory and application to reward shaping,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b115652d-26c2-4aaa-9bf6-ad36573eec73 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Self- supervised online reward shaping in sparse-reward environments,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cfc613d6-3c2c-4231-a5ce-395deac298a2 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning On learning intrinsic rewards for policy gradient methods,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 63a95a3f-c7cf-41c3-8739-cd423be75372 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Microfoun- dations of discounting,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 51f104c9-7e90-4b61-b2c8-39faabb455d8 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning The time interpretation of expected utility theory
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d7c7bd7-6180-4222-b328-314c3c7ec392 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning On-policy deep reinforcement learning for the average-reward criterion,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fca3c8ce-c7e5-424c-a0f9-526e21c89b17 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A provably-efficient model-free algo- rithm for infinite-horizon average-reward constrained markov decision processes,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1f56c440-556a-4953-9a7a-3bd54f3c865a · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Robust average-reward markov decision processes,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5631c70d-22e9-478c-b473-ccc365ac65df · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Proof of the ergodic theorem,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 7942132c-9b81-4527-b2ba-087db936646b · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8eb43f76-1e2c-411e-b29c-d9923c1b49fc · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Finite-time analysis of natural actor- critic for POMDPs,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f10314ca-72c1-4a23-894e-793bd3b5506b · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning The ergodicity solution of the cooperation puzzle,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6e6a2025-1cc9-4354-bdb9-4830d8a86c19 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning OpenAI Gym
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3fdd08-294a-437c-bd82-77dfd1264619 · outbound
Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Playing Atari with Deep Reinforcement Learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.