Pith. sign in

Paper Citation Record · LEDGER

Equivalence of stochastic and deterministic policy gradients

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.23244.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23244 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:29.884758Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6ed50c2-e24a-4539-9dc0-1ebead028663 · outbound

This paper cites Williams, Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learn- ing.

Equivalence of stochastic and deterministic policy gradients Williams, Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learn- ing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:28.346715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:28.346715Z digest=sha256:670df77dad593e000e4bc72b278bd0f872574a78f28efcec0f46c80883125f1f

Observation d7b99730-2a05-464a-98a2-b5e48986f3ca · outbound

This paper cites Amari, Natural Gradient Works Efficiently in Learning.

Equivalence of stochastic and deterministic policy gradients Amari, Natural Gradient Works Efficiently in Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:28.490621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:28.490621Z digest=sha256:d5986299d26c5378057938a3d06ed4b81ce8ccd8a9025f3365cfe5ad2498292e

Observation 86a1a3bf-59e6-40c6-9bdd-9c67d6f0e67b · outbound

This paper cites Sutton, D.

Equivalence of stochastic and deterministic policy gradients Sutton, D

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:28.631769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:28.631769Z digest=sha256:0154501087f7329b5d2f3a0785f934daf2f7baef0b6c067e74e7d0baab37b88a

Observation caf38370-71ce-4142-9d98-2f0cbe907d0d · outbound

This paper cites Kakade, A Natural Policy Gradient.

Equivalence of stochastic and deterministic policy gradients Kakade, A Natural Policy Gradient

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:28.814513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:28.814513Z digest=sha256:38722df982f28dc161f62f3a1d70fe8b79fa91b9b807c6d3b915c1074960e413

Observation 17efe37e-0539-433c-b43e-37868a62a573 · outbound

This paper cites Todorov, Lineary-solvable Markov decision problems.

Equivalence of stochastic and deterministic policy gradients Todorov, Lineary-solvable Markov decision problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:31.254023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:28.895088Z digest=sha256:4ab26da2e75c394697093f6123a57c04b97c105b83820d75833d9715b286ee99

Observation 471feaae-8410-4ff8-bbd0-12e4ecab025a · outbound

This paper cites Peters and S.

Equivalence of stochastic and deterministic policy gradients Peters and S

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:28.986703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:28.986703Z digest=sha256:f726cdc19e12a96ee4d9c8ffedb828402b918773c2e5fa1a64671c67595d08be

Observation 52f34415-5466-4eb2-8a9a-156aedd6bc2c · outbound

This paper cites Silver, G.

Equivalence of stochastic and deterministic policy gradients Silver, G

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:29.096487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:29.096487Z digest=sha256:719d4d726bdad2d8cad29407f9237bd9661075bde57eb43fe79f7ca2668a1711

Observation 588d8104-6aac-40ed-a952-be76780cce76 · outbound

This paper cites Todorov, Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in MuJoCo.

Equivalence of stochastic and deterministic policy gradients Todorov, Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in MuJoCo

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:31.013844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:29.183672Z digest=sha256:bc552e9d0d7bb054ec95ade5ea8cf83798cbd87d0e009e440f369b6857a483ce

Observation 1a6fe948-f573-4248-badd-215aead4470d · outbound

This paper cites Sutton and A.

Equivalence of stochastic and deterministic policy gradients Sutton and A

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:29.294517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:29.294517Z digest=sha256:32eca3b1ad2601b41180ec6d4e3b6ede049db37a5656f65fa0bbc538f012408f

Observation 73d7de37-da59-4057-bbfb-296ba197f135 · outbound

This paper cites Ciosek and S.

Equivalence of stochastic and deterministic policy gradients Ciosek and S

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:30.805288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:29.412389Z digest=sha256:f5e9ca927e80bad5b9e3c94567c8c5590af1cb05c6d9f5f571c89385c2b194d2

Observation f436505e-3d12-431c-b2fb-6c4f353a63f1 · outbound

This paper cites Szepesv´ ari, Policy gradients.

Equivalence of stochastic and deterministic policy gradients Szepesv´ ari, Policy gradients

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:30.593400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:29.518730Z digest=sha256:f9d8933abfc5aae515555762316772e57232998b6f6917cac125342c4e4db434

Observation 60a59dcf-e2db-486a-96da-73c7d86a585e · outbound

This paper cites Bhandari and D.

Equivalence of stochastic and deterministic policy gradients Bhandari and D

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:29.624070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:29.624070Z digest=sha256:e1f24f67c67b014399d56e4bd0e3cd74c9cec2712ce89f8566b62f08debf7570

Observation 13b0b48c-9e18-45db-bfeb-aa67359a237e · outbound

This paper cites an unresolved cited work.

Equivalence of stochastic and deterministic policy gradients Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:30.363099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:29.770370Z digest=sha256:4ec911bfd7ec1bf5a7b4d85c08a52148509b9a4eb1ba10d7bde4c23fc4b13026

Observation 3f0185fd-7f2f-43ec-b125-f5110267dc70 · outbound

This paper cites Todorov, Dynamical System Optimization.

Equivalence of stochastic and deterministic policy gradients Todorov, Dynamical System Optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:30.094762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:29.884758Z digest=sha256:0be2b9b8ba90324b29568ee10a3738934b23da71fb67a6afcb654852c39a300f

Pith citing papers

No inbound Pith citation observations are available.