Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2501.15034.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:46:40.962416Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-10T14:46:40.998711Z
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7f844ae2-f8d4-4b4c-b01c-44b4f3316422 · outbound
Divergence-Augmented Policy Optimization TensorFlow: A system for large-scale machine learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bbab76-2bba-4b81-aa95-deac999c1547 · outbound
Divergence-Augmented Policy Optimization Taming the Noise in Reinforcement Learning via Soft Updates
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0260c940-db19-4d58-924b-e0e36da41efd · outbound
Divergence-Augmented Policy Optimization Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6419e812-9cde-47ee-9b83-2cba3d7c5183 · outbound
Divergence-Augmented Policy Optimization Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64beb63e-5657-47de-b85e-6a571a04c30a · outbound
Divergence-Augmented Policy Optimization Equivalence Between Policy Gradients and Soft Q-Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3e350a-0b57-491c-bcbf-5751fcd76c4c · outbound
Divergence-Augmented Policy Optimization Sample Efficient Actor-Critic with Experience Replay
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe0769c-773e-427d-ae32-1c1cfe4ae64a · outbound
Divergence-Augmented Policy Optimization episodic life
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fad0416a-92c2-4dc8-9ea5-99e148600d82 · outbound
Divergence-Augmented Policy Optimization Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · outbound
Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7cf5fd38-49d3-4d3f-bc41-cb127a3de1d9 · outbound
Divergence-Augmented Policy Optimization A unified view of entropy-regularized Markov decision processes
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917e466e-b7f6-44c7-bca1-937fc3c786cb · outbound
Divergence-Augmented Policy Optimization Distributed Prioritized Experience Replay
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d6d90b9-b282-40ac-9659-9c9d353d11e7 · outbound
Divergence-Augmented Policy Optimization Unresolved cited work
Reference 2002
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2847209-d254-481e-8b15-a2ae258b2ed3 · outbound
Divergence-Augmented Policy Optimization Adam: A Method for Stochastic Optimization
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6467b71b-4615-43d1-9bfb-192d6363263d · outbound
Divergence-Augmented Policy Optimization IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8aab43f-8290-4770-b58b-b32373e9bc34 · outbound
Divergence-Augmented Policy Optimization human starts
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e128694-c735-47d5-9d9c-d983c8bf1b58 · outbound
Divergence-Augmented Policy Optimization Reinforcement Learning with Deep Energy-Based Policies
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77170a9f-57d9-4f00-9079-2733b297d719 · outbound
Divergence-Augmented Policy Optimization Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ca5d42-3891-436d-8d06-30a494554142 · outbound
Divergence-Augmented Policy Optimization Constrained Policy Optimization
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fd24c70-8541-40a7-a58c-b2f77bbd94c0 · inbound
Divergence-Augmented Policy Optimization Divergence-Augmented Policy Optimization
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.