Pith. sign in

Paper Citation Record · LEDGER

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning

As of 16 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2508.21443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21443 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:12.792191Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 99c67e3f-93a2-4f96-ad99-6f4c30902898 · outbound

This paper cites Human- level control through deep reinforcement learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Human- level control through deep reinforcement learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.727422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:09.753161Z digest=sha256:43171f3a8681f46ef7583e24249ff1f1ec74d04fe92ed99c92a595a030e59886

Observation 4f849639-2619-4e83-a01b-bd67c1b4310c · outbound

This paper cites Benchmarking deep reinforcement learning for continuous control,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Benchmarking deep reinforcement learning for continuous control,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.518368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:09.806178Z digest=sha256:cccd9057c46edc572959bc8a5834439352ed27642b826b3024a589af0e820dc0

Observation e88dae6b-4ef8-43df-a770-9381aeba42f5 · outbound

This paper cites Research priorities for robust and beneficial artificial intelligence,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Research priorities for robust and beneficial artificial intelligence,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.375126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:09.942629Z digest=sha256:a7a74fa3cbf021737aeb6eadb6cae9da23ee5590c12170b4bcf65d776f1509f5

Observation 8b4efd2d-524a-437c-aca8-8bd0763ecbee · outbound

This paper cites Concrete Problems in AI Safety.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Concrete Problems in AI Safety

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:10.073115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:10.073115Z digest=sha256:5385acc746efcf4c5ea3cb01c50636d2c08f33dd6109b7dd8ec96a79c0f512d0

Observation 29b1a652-9c0a-4afd-8728-4f4773efaf9d · outbound

This paper cites Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:10.172323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:10.172323Z digest=sha256:a0ceb8e4b5a357818d85ecb327d119dc8b57a4d58c30aef501e17556a35b5ae9

Observation f5870618-a950-4566-b73e-3cc5c7324a69 · outbound

This paper cites an unresolved cited work.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:10.288361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:10.288361Z digest=sha256:299bcbf8c0b30f6038e5cd41b566b86ca560e4456a902f0e96dd1e5160d76d5e

Observation 0b4b4b84-1086-4b56-a809-c853c4c5293f · outbound

This paper cites Bertsekas, Reinforcement Learning and Optimal Control.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Bertsekas, Reinforcement Learning and Optimal Control

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:10.405961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:10.405961Z digest=sha256:a978b63c8d5ac3c228278b2126c9ab91a477ca9a6076d4bbe0c0521e364357ad

Observation f2ee9982-9bc0-4cee-8352-15813493aff4 · outbound

This paper cites Reinforcement learning with non-ergodic reward in- crements: robustness via ergodicity transformations,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Reinforcement learning with non-ergodic reward in- crements: robustness via ergodicity transformations,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.263825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:10.518051Z digest=sha256:0a512ac04b9fd1a5afeef28fc253a3beb42d58e527205dac859ac4c6d7fccb1f

Observation 5b91cb89-0bfe-46c0-b32c-929af9a18041 · outbound

This paper cites Linearly-solvable markov decision problems,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Linearly-solvable markov decision problems,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.162010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:10.621757Z digest=sha256:285266685aca9c06a96f8866256eaa0ff326c16eb7ee2987a45d5371b92b885a

Observation 9a1d85f1-1d73-4396-985a-fb0d9b0e7f7e · outbound

This paper cites A theory of regularized Markov decision processes,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A theory of regularized Markov decision processes,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:16.032142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:10.719957Z digest=sha256:5c5c1c7d28f0902e2182c3c81cb00a00b9014503670dc9c3ccb959ef4304060d

Observation e77cda3b-c52f-4351-84de-5f764dbd789d · outbound

This paper cites Approximate modified policy iteration and its application to the game of tetris.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Approximate modified policy iteration and its application to the game of tetris

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.881505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:10.847778Z digest=sha256:db788d51afee4b7f31e30f74f83e74b81aaf8ce41286ac4f2fe929ba631009a4

Observation 3b2f7ae1-5ec8-449b-90df-fff9d6b77d6b · outbound

This paper cites Munchausen reinforcement learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Munchausen reinforcement learning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.762341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:10.991122Z digest=sha256:96bad52a86f7677a76e00298d89474aa28b757df7f4162273da88a7c12928e2c

Observation fe9d1c02-fa7a-405a-aacf-8bc83e35b612 · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Maximum a Posteriori Policy Optimisation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:11.102414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:11.102414Z digest=sha256:9fe349c8a2a7d4a4222640e89d0d51a8ba5bb52feb9426888f52001cd92e4ed8

Observation 37475928-ea80-4620-b75d-a97fed392256 · outbound

This paper cites Leverage the average: an analysis of KL regularization in reinforcement learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Leverage the average: an analysis of KL regularization in reinforcement learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.633431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.163630Z digest=sha256:c881cb6fb1b3c4d086efb0008ce69fb2efaeb6df9fa62a2ac09403ed4efc3765

Observation d340f80d-c6b6-4000-a8ef-b7a9cd3b862a · outbound

This paper cites Incremental multi-step q-learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Incremental multi-step q-learning,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.516046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.223547Z digest=sha256:e06ee2d6040b8f839caa09e72b16b5ea08aebecff233d6f066cafc3d63cf770a

Observation 9d80eb6e-7a80-493b-94db-6a6d18fb2dd8 · outbound

This paper cites Multi- step reinforcement learning: A unifying algorithm,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Multi- step reinforcement learning: A unifying algorithm,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.406979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.298837Z digest=sha256:e47b44fec4bd48faae8c54b79448a6f78616b6404a396563abf976956f53b243

Observation d44e8e4c-aef5-45a9-9d6d-861a865a1590 · outbound

This paper cites Multi-Bellman operator for convergence of $Q$-learning with linear function approximation.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Multi-Bellman operator for convergence of $Q$-learning with linear function approximation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:13.133448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.426915Z digest=sha256:92333c04601c9eb5dfe090b0eb519ae9bc48112c416c51bcf6dd19921e2d1d5b

Observation 176002bd-bd9d-46f3-8e54-c5231c0efe9b · outbound

This paper cites A novel multi-step q-learning method to improve data efficiency for deep reinforcement learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A novel multi-step q-learning method to improve data efficiency for deep reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.292218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.557222Z digest=sha256:70128023f196a577a6d49e7d9b0854a6d7febd4f7d6840c712527e8e726821aa

Observation 9ead2d0e-8c2c-4029-a97a-f9adbda5bed2 · outbound

This paper cites Rainbow: Combining improvements in deep reinforcement learning,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Rainbow: Combining improvements in deep reinforcement learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.176986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.622988Z digest=sha256:fee5b072e580bd5f9a32aaf2996dfff1a591679d4390b32cd373e4b83f232d89

Observation 8b2937fa-e06b-47d8-9fb3-80102323141c · outbound

This paper cites Policy invariance under reward transformations: Theory and application to reward shaping,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Policy invariance under reward transformations: Theory and application to reward shaping,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:15.033813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.691962Z digest=sha256:4789c30add8d5e40e6ea452ccff930445a9ed5aeb5c7caa1c1efc98d1f454445

Observation b115652d-26c2-4aaa-9bf6-ad36573eec73 · outbound

This paper cites Self- supervised online reward shaping in sparse-reward environments,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Self- supervised online reward shaping in sparse-reward environments,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.907088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.767253Z digest=sha256:f1a9e934844eb5ca84f1f34d1137f5b6b87577a89533673975dbc58166ce5f31

Observation cfc613d6-3c2c-4231-a5ce-395deac298a2 · outbound

This paper cites On learning intrinsic rewards for policy gradient methods,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning On learning intrinsic rewards for policy gradient methods,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.705725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.856828Z digest=sha256:33da3832485731d5d01b95ff1aad6d4720a339472769b7f89e97d23d092cc6c3

Observation 63a95a3f-c7cf-41c3-8739-cd423be75372 · outbound

This paper cites Microfoun- dations of discounting,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Microfoun- dations of discounting,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.508866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:11.931008Z digest=sha256:c894a3760ad255352c7ce88bd01ccfd038f2e748b8e70012099e72cbd8219ff4

Observation 51f104c9-7e90-4b61-b2c8-39faabb455d8 · outbound

This paper cites The time interpretation of expected utility theory.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning The time interpretation of expected utility theory

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-05T14:23:12.989244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.015329Z digest=sha256:959e58ae047d114e7467abd3fe9276bf77c7b644db532b068b9279b7c85e26ce

Observation 0d7c7bd7-6180-4222-b328-314c3c7ec392 · outbound

This paper cites On-policy deep reinforcement learning for the average-reward criterion,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning On-policy deep reinforcement learning for the average-reward criterion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.318428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.109419Z digest=sha256:a82b0ea2f5c8d0732ea134a3c24896030fbe3583981e6095e1b24a81e5a5c5d9

Observation fca3c8ce-c7e5-424c-a0f9-526e21c89b17 · outbound

This paper cites A provably-efficient model-free algo- rithm for infinite-horizon average-reward constrained markov decision processes,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning A provably-efficient model-free algo- rithm for infinite-horizon average-reward constrained markov decision processes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.193869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.185584Z digest=sha256:07dc37b6c0b83bd9480fe7dcb785f0c12326209dd9b4551a0d0f55461b97b30a

Observation 1f56c440-556a-4953-9a7a-3bd54f3c865a · outbound

This paper cites Robust average-reward markov decision processes,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Robust average-reward markov decision processes,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:14.026163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.260190Z digest=sha256:684031245fcd0415c599368648b91f95d23f7c9d99ef3867aee4f631e303c342

Observation 5631c70d-22e9-478c-b473-ccc365ac65df · outbound

This paper cites Proof of the ergodic theorem,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Proof of the ergodic theorem,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:13.832657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.291154Z digest=sha256:5115567d3050a39b76777d5a206736d769d7d244247105ba202a1c5c86b3c1b1

Observation 7942132c-9b81-4527-b2ba-087db936646b · outbound

This paper cites an unresolved cited work.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-05T14:23:13.646460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.377413Z digest=sha256:b0a69b1c33aebd6b2c4597ca6a1256b005c4ed1c629e33d0282292fa154a04e8

Observation 8eb43f76-1e2c-411e-b29c-d9923c1b49fc · outbound

This paper cites Finite-time analysis of natural actor- critic for POMDPs,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Finite-time analysis of natural actor- critic for POMDPs,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:13.471063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.443375Z digest=sha256:28a757f66876f16af1a9440c8808d4b3b1d5c8e59a49749fe38fb986c5af899c

Observation f10314ca-72c1-4a23-894e-793bd3b5506b · outbound

This paper cites The ergodicity solution of the cooperation puzzle,.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning The ergodicity solution of the cooperation puzzle,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:13.314948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T14:23:12.612348Z digest=sha256:4fe8a55758fccde71ed4e7de0e000cbfb2d0bc67a0138fa8d98c13f6ecba5e11

Observation 6e6a2025-1cc9-4354-bdb9-4830d8a86c19 · outbound

This paper cites OpenAI Gym.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning OpenAI Gym

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:12.734177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:12.734177Z digest=sha256:4a8767065d07b80db7c8519d5621cb5110f22017b65971c4d6ad0906b52e221d

Observation 2d3fdd08-294a-437c-bd82-77dfd1264619 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Beyond expected value: geometric mean optimization for long-term policy performance in reinforcement learning Playing Atari with Deep Reinforcement Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:12.792191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:12.792191Z digest=sha256:b57aafd193e64dd1dc24944358cc4a1d707c2541b2696896b1da58c32362f5d3

Pith citing papers

No inbound Pith citation observations are available.