Pith. sign in

Paper Citation Record · LEDGER

Divergence-Augmented Policy Optimization

As of 11 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2501.15034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15034 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T14:46:40.998711Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f844ae2-f8d4-4b4c-b01c-44b4f3316422 · outbound

This paper cites TensorFlow: A system for large-scale machine learning.

Divergence-Augmented Policy Optimization TensorFlow: A system for large-scale machine learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.885943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.885943Z digest=sha256:45a6fb6b4fb5907690bbfa37c98a98ec61539c3fa92bb1db350c5e92230a0e62

Observation 97bbab76-2bba-4b81-aa95-deac999c1547 · outbound

This paper cites Taming the Noise in Reinforcement Learning via Soft Updates.

Divergence-Augmented Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.901446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.901446Z digest=sha256:23b9c322c06a417f1294345301ccc9ed2686ea2bd63789366319e1ea5d8af82c

Observation 0260c940-db19-4d58-924b-e0e36da41efd · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.217949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.917022Z digest=sha256:4bb98d5f782b8903c34e103493b0d119d6dd04427fe5860a5a892792ce9b154b

Observation 6419e812-9cde-47ee-9b83-2cba3d7c5183 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.934019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.934019Z digest=sha256:f7c70705e71de4665236b32c023f72cfe2e1ff1dd4c1ad1259d1efbaf8aa8f86

Observation 64beb63e-5657-47de-b85e-6a571a04c30a · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

Divergence-Augmented Policy Optimization Equivalence Between Policy Gradients and Soft Q-Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.941589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.941589Z digest=sha256:2a6623fb504762716c1dad30a50a48ae8717b40d75685dd49af375bb725fb2de

Observation ff3e350a-0b57-491c-bcbf-5751fcd76c4c · outbound

This paper cites Sample Efficient Actor-Critic with Experience Replay.

Divergence-Augmented Policy Optimization Sample Efficient Actor-Critic with Experience Replay

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.945463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.945463Z digest=sha256:48effa1e6bb920e42a7812e20e317ef7d27a7edbdd072de84cedfed384174c53

Observation abe0769c-773e-427d-ae32-1c1cfe4ae64a · outbound

This paper cites episodic life.

Divergence-Augmented Policy Optimization episodic life

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:41.171930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.954832Z digest=sha256:4c35c1583f622479bc74ff16150eb5117ad49c2b5b71639cdb4f42782f1fcd6f

Observation fad0416a-92c2-4dc8-9ea5-99e148600d82 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.160877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.958324Z digest=sha256:36f6ffb6c226f5a4271cbe3e6d1742f9f02c216425554c9ddafb5abc4c52dcff

Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · outbound

This paper cites Divergence-Augmented Policy Optimization.

Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:46:41.005958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.962416Z digest=sha256:7403e0618eed253e687996a7a5356ccf691f5cdcfc3fa0a55369909b6f0080cf

Observation 7cf5fd38-49d3-4d3f-bc41-cb127a3de1d9 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Divergence-Augmented Policy Optimization A unified view of entropy-regularized Markov decision processes

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.937680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.937680Z digest=sha256:84dfbc65b4bb492786ff100589175a1f0d139026f6c84d79eacd9138f3ccd4bc

Observation 917e466e-b7f6-44c7-bca1-937fc3c786cb · outbound

This paper cites Distributed Prioritized Experience Replay.

Divergence-Augmented Policy Optimization Distributed Prioritized Experience Replay

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.920694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.920694Z digest=sha256:5c9414dd8ac84d42225be50697b4c73bfff0e70404d5fd37c077a9921bac9e7a

Observation 4d6d90b9-b282-40ac-9659-9c9d353d11e7 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 2002

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.203702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.924989Z digest=sha256:291f5074fad23935aabf2a9725d89e0b99d2ecd1adca65f82a974d0de9bf398c

Observation d2847209-d254-481e-8b15-a2ae258b2ed3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Divergence-Augmented Policy Optimization Adam: A Method for Stochastic Optimization

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.929280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.929280Z digest=sha256:285b2fccf97ace6d771fa6c6f8ab85f47f40cec1cae9a01955704331847cc420

Observation 6467b71b-4615-43d1-9bfb-192d6363263d · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

Divergence-Augmented Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.896067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.896067Z digest=sha256:735ffdf565d71a98b39bcf384a428e7221463f6c062c2d818e05f00aad3d78d0

Observation a8aab43f-8290-4770-b58b-b32373e9bc34 · outbound

This paper cites human starts.

Divergence-Augmented Policy Optimization human starts

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:41.182627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.950996Z digest=sha256:17dc479a95f75b772a094fb7687fe19dcf40b6bd49f4dfd730f19f4720457a36

Observation 6e128694-c735-47d5-9d9c-d983c8bf1b58 · outbound

This paper cites Reinforcement Learning with Deep Energy-Based Policies.

Divergence-Augmented Policy Optimization Reinforcement Learning with Deep Energy-Based Policies

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.905786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.905786Z digest=sha256:9554ac1cc9bbfbc603909a741572d729c440d199d94b102cd3b36aa9ce13b985

Observation 77170a9f-57d9-4f00-9079-2733b297d719 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Divergence-Augmented Policy Optimization Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.910483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.910483Z digest=sha256:0b46903fdf77aab91ed1f0ada25d4abbacc96803901e7b2baadc66891e8243ee

Observation 90ca5d42-3891-436d-8d06-30a494554142 · outbound

This paper cites Constrained Policy Optimization.

Divergence-Augmented Policy Optimization Constrained Policy Optimization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.891011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.891011Z digest=sha256:b3e1f97821568e4d39a08b983d91c01eb53765fe7a6328a19cd1d89cbdbe1e65

Pith citing papers

Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · inbound

Divergence-Augmented Policy Optimization cites this paper.

Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:46:41.005958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T14:46:40.962416Z digest=sha256:7403e0618eed253e687996a7a5356ccf691f5cdcfc3fa0a55369909b6f0080cf