Pith. sign in

Paper Citation Record · LEDGER

Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2212.00603.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.00603 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:54:11.972128Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T18:55:59.140544Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a5de33fd-8f62-47de-b381-294e48cdd827 · inbound

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning cites this paper.

Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:11.972128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:11.972128Z digest=sha256:9a5cd989fd9e823bdd5de0c459a9dedfc5ad3ccd9a09edfaf3c5e94aa4c72a96

Observation f54689f6-5d42-400b-a9bf-f0e1dfc99477 · inbound

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model cites this paper.

A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:57.865812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:57.865812Z digest=sha256:15cabf0dbe44032afd009ae91e661dcb381ab214d792c7377f520dc9b23c9e4b

Observation 287162d5-da44-4a40-a323-5ea354c2632d · inbound

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies cites this paper.

A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:50:57.135525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:47:47.782246Z digest=sha256:77e7e34a283f4c9d176346c847da1532125cdc2e61cd0ca7b9e3c9307b6f81ec

Observation 74e3f11d-0540-4342-bec9-7f8b67a8098c · inbound

Learning in Markovian bandits with non-observable states and constrained decision epochs cites this paper.

Learning in Markovian bandits with non-observable states and constrained decision epochs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.142261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T01:20:21.181843Z digest=sha256:06861148031a025d3349ef5c6d0d0da4325ddbf5efc023052c8cbecae9053e61