Pith. sign in

Paper Citation Record · LEDGER

Divergence-Augmented Policy Optimization

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2501.15034.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15034 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T14:46:40.998711Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f844ae2-f8d4-4b4c-b01c-44b4f3316422 · outbound

This paper cites TensorFlow: A system for large-scale machine learning.

Divergence-Augmented Policy Optimization TensorFlow: A system for large-scale machine learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.885943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.885943Z digest=sha256:9369f06d9f13d2fb89a0c4fd8e959df8fdc2290639399b3bc8ab00cb556edf0d

Observation 97bbab76-2bba-4b81-aa95-deac999c1547 · outbound

This paper cites Taming the Noise in Reinforcement Learning via Soft Updates.

Divergence-Augmented Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.901446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.901446Z digest=sha256:c5d998ff27c86df0fdbce90ff8f1b8de0f9b251a0943df12531921f70cfd306f

Observation 0260c940-db19-4d58-924b-e0e36da41efd · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.217949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.917022Z digest=sha256:0639670079ad3bcd523a243c5746d60b164afa9861bbf124c7831cd521fb79f2

Observation 6419e812-9cde-47ee-9b83-2cba3d7c5183 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.934019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.934019Z digest=sha256:ac9090563e43f41a772c84d80a17756daf2ff686052408fa1bf0e50b756aec15

Observation 64beb63e-5657-47de-b85e-6a571a04c30a · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

Divergence-Augmented Policy Optimization Equivalence Between Policy Gradients and Soft Q-Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.941589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.941589Z digest=sha256:6dd27256aa5bb6f5792193bbb005326043bcb23becc59bb0ffa23f547427105c

Observation ff3e350a-0b57-491c-bcbf-5751fcd76c4c · outbound

This paper cites Sample Efficient Actor-Critic with Experience Replay.

Divergence-Augmented Policy Optimization Sample Efficient Actor-Critic with Experience Replay

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.945463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.945463Z digest=sha256:45a48d3c20ee4d09f8f18cebcba66a5b33881806ca7e6a694ee114212f5aa44c

Observation abe0769c-773e-427d-ae32-1c1cfe4ae64a · outbound

This paper cites episodic life.

Divergence-Augmented Policy Optimization episodic life

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:41.171930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.954832Z digest=sha256:0ae01ab4567c81d9525767e7da919725ed268844682f4847c5b46da674e995cf

Observation fad0416a-92c2-4dc8-9ea5-99e148600d82 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.160877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.958324Z digest=sha256:a81b63a1ca4f4b208d221cfdb974df88a0e380c3bd27b7eb8e5dc110e0387985

Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · outbound

This paper cites Divergence-Augmented Policy Optimization.

Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:46:41.005958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.962416Z digest=sha256:fca3ce4d9e4459766cfd6af687d5d9c8333311a9fa509324240717ee01838e91

Observation 7cf5fd38-49d3-4d3f-bc41-cb127a3de1d9 · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Divergence-Augmented Policy Optimization A unified view of entropy-regularized Markov decision processes

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.937680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.937680Z digest=sha256:22fede44053337c09f4fcdefc92f88ac473f9de06ac9811d6cd23210aa7b9b35

Observation 917e466e-b7f6-44c7-bca1-937fc3c786cb · outbound

This paper cites Distributed Prioritized Experience Replay.

Divergence-Augmented Policy Optimization Distributed Prioritized Experience Replay

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.920694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.920694Z digest=sha256:dc34646b9170a08436c2fad02be7abfa08015716783e65949542c67a4398db30

Observation 4d6d90b9-b282-40ac-9659-9c9d353d11e7 · outbound

This paper cites an unresolved cited work.

Divergence-Augmented Policy Optimization Unresolved cited work

Reference 2002

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:46:41.203702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.924989Z digest=sha256:2eb6f0f5068dbe9305c60b35be463604c9b5652e96fc89d9990a7edf3d688492

Observation d2847209-d254-481e-8b15-a2ae258b2ed3 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Divergence-Augmented Policy Optimization Adam: A Method for Stochastic Optimization

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.929280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.929280Z digest=sha256:4b0e390ec9120a484313d92816aef52ebd210a4681823bdfaf0d650a11174540

Observation 6467b71b-4615-43d1-9bfb-192d6363263d · outbound

This paper cites IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures.

Divergence-Augmented Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.896067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.896067Z digest=sha256:cafb4eee0792444f49f51f009f71d370c5b2ed974d3b23875039aa91f7895add

Observation a8aab43f-8290-4770-b58b-b32373e9bc34 · outbound

This paper cites human starts.

Divergence-Augmented Policy Optimization human starts

Reference 2013

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:46:41.182627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.950996Z digest=sha256:9600880ec0e170108bcc4c63a74aa58cb436cc5196e2dc2d2409e4106fe3e4d6

Observation 6e128694-c735-47d5-9d9c-d983c8bf1b58 · outbound

This paper cites Reinforcement Learning with Deep Energy-Based Policies.

Divergence-Augmented Policy Optimization Reinforcement Learning with Deep Energy-Based Policies

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.905786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.905786Z digest=sha256:4f179e87c497982eeea2b2dd9a8ac1ae07d884c54eaeddf935c013c7cc729c2e

Observation 77170a9f-57d9-4f00-9079-2733b297d719 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Divergence-Augmented Policy Optimization Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.910483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.910483Z digest=sha256:ffa1d4742efc93a64fdd3d9d31ab78d4fd3919720b2181ecb2e9a7e7d4744c3a

Observation 90ca5d42-3891-436d-8d06-30a494554142 · outbound

This paper cites Constrained Policy Optimization.

Divergence-Augmented Policy Optimization Constrained Policy Optimization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:40.891011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:40.891011Z digest=sha256:30d0daf702e13aa333d01f0e739d2366024c713a085a88ec9a558e63d013ff72

Pith citing papers

Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · inbound

Divergence-Augmented Policy Optimization cites this paper.

Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T14:46:41.005958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T14:46:40.962416Z digest=sha256:fca3ce4d9e4459766cfd6af687d5d9c8333311a9fa509324240717ee01838e91