Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:21:29.954534Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.05953.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:21:29.954534Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1af85dc5-728c-4ef0-b73e-d409453afb69 · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Continuous control with deep reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a224fb8b-1410-48fe-9a42-989f10280e04 · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6091738d-e883-4ec0-96d1-66569de17fed · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65f17fcb-8643-478e-b787-55678502f56e · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c0e0e6c-6364-41b6-b6ab-16f999295e79 · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work
Reference 2002
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3bff57d4-77ba-4980-b962-88df74cf03af · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work
Reference 2006
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e3dabbf-8f89-4c53-80c7-b30bff4f7e6c · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Policy Gradient in Partially Observable Environments: Approximation and Convergence
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7531016-6d4f-4552-8e16-26323277461b · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Journal of Optimization Theory and Applications 153, 688–708
Reference 2012
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4271e805-bac3-4047-b529-7c02f71a29ea · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Safe Exploration in Continuous Action Spaces
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23eb6839-ec15-461b-983d-65b88833db7d · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Advances in Neural Information Processing Systems (NeurIPS) 33, 8378–8390
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 701ec355-c83d-4a44-a034-63684c786d0b · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Policy Optimization for Constrained MDPs with Provable Fast Global Convergence
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dee27c5-8433-4924-9ad7-512cb59a7cf7 · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5d26c38f-73b2-439c-9214-e1e4b74089e0 · outbound
Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes 11506–11533
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.