Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2404.03372.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:56:22.453068Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:23:15.210988Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 53de0872-a891-4fab-9ada-976fdf1488b8 · inbound
Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Elementary Analysis of Policy Gradient Methods
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fd8aad-413b-403f-9e3f-ae8aa635afec · inbound
Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives Elementary Analysis of Policy Gradient Methods
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b09d0da5-99b1-4cb7-a56b-ed62587608c8 · inbound
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum Elementary Analysis of Policy Gradient Methods
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78f1a5c5-0921-403e-ae92-c241984b705b · inbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Elementary Analysis of Policy Gradient Methods
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2d089cb-e58d-4fa2-8d19-30d3af83bed7 · inbound
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness Elementary Analysis of Policy Gradient Methods
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6193d972-ea4e-4cff-b657-5a0c79578cce · inbound
On the Policy Convergence of Policy Mirror Descent Methods Elementary Analysis of Policy Gradient Methods
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 030f074d-1a7c-4837-b5c5-ef81206316d4 · inbound
Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes Elementary Analysis of Policy Gradient Methods
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.