Pith. sign in

Paper Citation Record · LEDGER

First-order Policy Optimization for Robust Markov Decision Process

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2209.10579.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.10579 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:15.771773Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T10:33:18.772880Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ad6ae68f-dc42-47b6-84b5-6ddabd41e6d9 · inbound

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form cites this paper.

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form First-order Policy Optimization for Robust Markov Decision Process

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:03:30.920726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-23T21:58:56.180393Z digest=sha256:9709a48ced54b348a705b86febfa425652b62b31bd735b40a4db299cff1b90cd

Observation 6b2800b2-64ec-4e46-b8c4-4e056478a7aa · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning First-order Policy Optimization for Robust Markov Decision Process

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.213481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.213481Z digest=sha256:1bf1419c3328c4849bc5d66883b6ab0c3a062d5166b1b1277614bf01f74c520f

Observation 02e14e7c-de26-479c-bdd9-1439190d22ba · inbound

Sample Complexity for Markov Decision Processes and Stochastic Optimal Control with Static Risk Measures cites this paper.

Sample Complexity for Markov Decision Processes and Stochastic Optimal Control with Static Risk Measures First-order Policy Optimization for Robust Markov Decision Process

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:50:51.101008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:31:33.732451Z digest=sha256:3d4c55acf0d917ed12512cd791988457397b207ee34b47659b82547e828827a7

Observation ca97eebf-6bc5-4ee6-9fb9-206d4d7de8ec · inbound

Value Mirror Descent for Reinforcement Learning cites this paper.

Value Mirror Descent for Reinforcement Learning First-order Policy Optimization for Robust Markov Decision Process

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:49.576825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:07:41.863349Z digest=sha256:c2bae54a42045d414ec0a36bf0ba9fe48f740b8b277559ce398c346be78886c6

Observation 6f144f27-1a45-4766-a00a-be694147343a · inbound

Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions cites this paper.

Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions First-order Policy Optimization for Robust Markov Decision Process

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:05.312687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T22:01:46.060931Z digest=sha256:6697961c320e5062ae5c7a8336ac35c9578bfc865d4e739ff336f613cd10ca92

Observation 748102b9-a20c-4ab4-8810-057dbc871299 · inbound

Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework cites this paper.

Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework First-order Policy Optimization for Robust Markov Decision Process

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:29.369993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:01:07.589351Z digest=sha256:0201e92af1045fe390e52460b2a0caa72fbdd98f8ac87d797686c7f2ad1a3c51

Observation c6747c55-653f-44c3-8ec3-2862ce2e3095 · inbound

Robust Markov Decision Processes on Continuous State Spaces cites this paper.

Robust Markov Decision Processes on Continuous State Spaces First-order Policy Optimization for Robust Markov Decision Process

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:18.774758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T10:31:29.175140Z digest=sha256:cbf9c5222aef5746617505a449b39ab01f919c9d783147b22869f3646937432c

Observation 9e666d68-b161-47fe-8196-cd229dfebfd1 · inbound

Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework cites this paper.

Non-Asymptotic Convergence of Stochastic Iterative Algorithms: A Lyapunov Framework First-order Policy Optimization for Robust Markov Decision Process

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.683876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T23:11:02.699220Z digest=sha256:67a2a4bac6ff092588277387a540a86f2dadf3c4691249dc3a312947dfd0eaf5

Observation 718eca6b-6355-4dcd-8927-5fdaaa3dab37 · inbound

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions cites this paper.

Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions First-order Policy Optimization for Robust Markov Decision Process

Reference 142

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:15.771773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:15.771773Z digest=sha256:94f969730470ddf8645cad2cb6d176541c137c1a2db3a7065e53bf4b72b3420a