Pith. sign in

Paper Citation Record · LEDGER

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes

As of 7 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.05953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05953 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:21:29.954534Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1af85dc5-728c-4ef0-b73e-d409453afb69 · outbound

This paper cites Continuous control with deep reinforcement learning.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Continuous control with deep reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.928004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.928004Z digest=sha256:b1162b82952509e1a9b5bfbaa790ccf5ad837a5ec731bef4ab84716c32362493

Observation a224fb8b-1410-48fe-9a42-989f10280e04 · outbound

This paper cites Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Last-Iterate Convergence of General Parameterized Policies in Constrained MDPs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.937676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.937676Z digest=sha256:056ae5e1d90505a1e95171337ebcbdd0bbb2453a077ee03fa7201128da3e043d

Observation 6091738d-e883-4ec0-96d1-66569de17fed · outbound

This paper cites an unresolved cited work.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:21:30.156809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.954534Z digest=sha256:5ab7fafe63f8224b2b2c5ed09b84d16f4c8dc925610cd9a8ef6f3425e5ac18d7

Observation 65f17fcb-8643-478e-b787-55678502f56e · outbound

This paper cites an unresolved cited work.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:21:30.226107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.733240Z digest=sha256:e15e4a7c108a870f6781fab503fd2d5dd47f18fa9bda3a182db273510f868786

Observation 1c0e0e6c-6364-41b6-b6ab-16f999295e79 · outbound

This paper cites an unresolved cited work.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work

Reference 2002

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:21:30.184116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.945891Z digest=sha256:ccce215f991c35b1a1206e4d98e5b6d982fe85ebccdd4590cdf26303b86fb511

Observation 3bff57d4-77ba-4980-b962-88df74cf03af · outbound

This paper cites an unresolved cited work.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work

Reference 2006

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:21:30.197055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.941700Z digest=sha256:27d74b1d91e0dee0edca20b515e5aaeeebbda2f01cac281cc7d3b067d24211ad

Observation 8e3dabbf-8f89-4c53-80c7-b30bff4f7e6c · outbound

This paper cites Policy Gradient in Partially Observable Environments: Approximation and Convergence.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Policy Gradient in Partially Observable Environments: Approximation and Convergence

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.414918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.414918Z digest=sha256:b0914a83a12470b704633611331aa6055d78dc80acd6d3bc28f0ac3dafcc1e1f

Observation d7531016-6d4f-4552-8e16-26323277461b · outbound

This paper cites Journal of Optimization Theory and Applications 153, 688–708.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Journal of Optimization Theory and Applications 153, 688–708

Reference 2012

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:21:30.240836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.487809Z digest=sha256:e557780ca9f29773a13941772c8e3048b295122fee71841f01fffddd8b933c8e

Observation 4271e805-bac3-4047-b529-7c02f71a29ea · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Safe Exploration in Continuous Action Spaces

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.646248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.646248Z digest=sha256:eec86df12d328d09ebe7b16dbce090165a591c82926b0245e9c5203a30aaf5b1

Observation 23eb6839-ec15-461b-983d-65b88833db7d · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) 33, 8378–8390.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Advances in Neural Information Processing Systems (NeurIPS) 33, 8378–8390

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.828608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.828608Z digest=sha256:8387b8993278c0e8c068f3cd571739a23866e3a6be7c718fae511d0f17e43a7a

Observation 701ec355-c83d-4a44-a034-63684c786d0b · outbound

This paper cites Policy Optimization for Constrained MDPs with Provable Fast Global Convergence.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Policy Optimization for Constrained MDPs with Provable Fast Global Convergence

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T10:21:29.933220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:21:29.933220Z digest=sha256:7bc9eaafce14fc77d9686d4f2490e3797ba54d2c4294ed1b8b36a6e8a56c494c

Observation 0dee27c5-8433-4924-9ad7-512cb59a7cf7 · outbound

This paper cites an unresolved cited work.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:21:30.170599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.950115Z digest=sha256:c79bf448e2a23d681c9e8ea7b1b9d2b577c66f0a0b895c2b04b09da26fe1a3d7

Observation 5d26c38f-73b2-439c-9214-e1e4b74089e0 · outbound

This paper cites 11506–11533.

Learning Deterministic Policies with Policy Gradients in Constrained Markov Decision Processes 11506–11533

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:21:30.211143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:21:29.920865Z digest=sha256:aae6c5a0348f9267126af4a42b892c0d754b234844c20e71846b4a9f28b390ad

Pith citing papers

No inbound Pith citation observations are available.