Pith. sign in

Paper Citation Record · LEDGER

Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2212.00603.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.00603 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:42:09.194938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:55:59.140544Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32ab1a26-20b5-443f-9542-807e6181f9e3 · inbound

Near-Optimal Sample Complexity for MDPs via Anchoring cites this paper.

Near-Optimal Sample Complexity for MDPs via Anchoring Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T22:48:32.824182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:48:32.824182Z digest=sha256:3c6af2ed480af38fdba256526a37a9d9b7230c7b6218acc055f59a2e63527e8c

Observation 1dbf06e3-abd3-41d1-b432-3413b6dc6477 · inbound

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs cites this paper.

A Computationally Efficient Algorithm for Infinite-Horizon Average-Reward Linear MDPs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:42:09.194938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:42:09.194938Z digest=sha256:e70be7d88102034e1db1c3a60b1928fa84645741e34b53111d4158a12c35ccb4

Observation ff63a073-7486-48e5-a640-243d1f8a9835 · inbound

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis cites this paper.

Model-Free Robust Average-Reward Reinforcement Learning with Sample Complexity Analysis Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:45:55.187895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:45:55.187895Z digest=sha256:0fcfcce16760c559463138217023f20cac2746e6d7db580379d4bf7f5bde111a

Observation a5de33fd-8f62-47de-b381-294e48cdd827 · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.972128Z digest=sha256:1ff2192a10e7c8f63b05f6968c57814c4eb96051111164efffe70773f30681dc

Observation afdf2dd7-7116-48d9-98bc-5e58ec3549fd · inbound

Thresholds for sensitive optimality and Blackwell optimality in stochastic games cites this paper.

Thresholds for sensitive optimality and Blackwell optimality in stochastic games Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:06:31.959616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:06:31.959616Z digest=sha256:bf09e481d5337ba610fcbc8c861618084ae0289fce4884c251b93dbc24c94da4

Observation f54689f6-5d42-400b-a9bf-f0e1dfc99477 · inbound

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model cites this paper.

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:57.865812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:57.865812Z digest=sha256:0cf72a855df926342d6fd7b7a5b321a123fdd35f23954aebc19a86b63e326021

Observation 287162d5-da44-4a40-a323-5ea354c2632d · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.135525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:9b6b2d8807dcc59ae0515d07b38fc3e69a4eda673bae7d53dc42d91062e9830e

Observation 74e3f11d-0540-4342-bec9-7f8b67a8098c · inbound

Learning in Markovian bandits with non-observable states and constrained decision epochs cites this paper.

Learning in Markovian bandits with non-observable states and constrained decision epochs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.142261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:f876af9b3cc9b808d33f59866b8c1ba51aff9b92887e4b096c360a6d5067629d