Pith. sign in

Paper Citation Record · LEDGER

A general class of surrogate functions for stable and efficient reinforcement learning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2108.05828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.05828 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:14:33.989387Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:48:17.553752Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fba64ec1-bb3b-4173-bca5-b9c122c15fac · inbound

Fast Convergence of Softmax Policy Mirror Ascent cites this paper.

Fast Convergence of Softmax Policy Mirror Ascent A general class of surrogate functions for stable and efficient reinforcement learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:14:33.989387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T18:14:33.989387Z digest=sha256:f49de0adef1e036b065571292d39e00bc688d37dbd690e695fb11610348a27e6

Observation f4a52692-0c4d-462f-9324-32a22e6970d2 · inbound

Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G cites this paper.

Multi-Agent Reinforcement Learning in Wireless Distributed Networks for 6G A general class of surrogate functions for stable and efficient reinforcement learning

Reference 200

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:46.683421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:46.683421Z digest=sha256:bdbee54fe8dcbd85da7817438265b3eb57ae253340e79779cee3af5128817e42

Observation 84b6513c-37d4-49fb-a202-acc5094a8723 · inbound

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives cites this paper.

Dissecting Discrete Soft Actor-Critic: Limitations and Principled Alternatives A general class of surrogate functions for stable and efficient reinforcement learning

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T17:06:39.818733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T17:05:36.100114Z digest=sha256:de1f0b2457b4a663359cb1ac8c5fc3803f686168d5f2f64990c2f34b2690960f

Observation 1fc7e9a6-ab16-4e5b-a227-5eba737a471f · inbound

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs cites this paper.

Augmented Lagrangian Method for Last-Iterate Convergence for Constrained MDPs A general class of surrogate functions for stable and efficient reinforcement learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:27:29.105915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T07:27:12.302109Z digest=sha256:1a79368a02566e148068649f5c78c023ef3aa4ba872da14a880ac765b26a28b6

Observation 7f7f6c33-825d-4603-a0c2-dc63c6bd6fa9 · inbound

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation cites this paper.

Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation A general class of surrogate functions for stable and efficient reinforcement learning

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:48:17.555556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-20T12:44:29.147095Z digest=sha256:5489331f766dd01c5c55c7fa34d45052a48c12d9769e082f98d5057a187bf219

Observation 652097ef-63fc-4676-9b7f-82e914a7b81a · inbound

On the Policy Convergence of Policy Mirror Descent Methods cites this paper.

On the Policy Convergence of Policy Mirror Descent Methods A general class of surrogate functions for stable and efficient reinforcement learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:17:12.534820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T04:17:12.534820Z digest=sha256:0358f02328d257287a3fc8406103de305f575c21725877403eaa688e92f22edf