Pith. sign in

Paper Citation Record · LEDGER

Asymptotically optimal regret in communicating Markov decision processes

As of 9 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.18064.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18064 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b97480b8-281b-492d-a57e-531ff7779aa3 · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Asymptotically optimal regret in communicating Markov decision processes The regret lower bound for communicating Markov Decision Processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.317140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.317140Z digest=sha256:0961246376205965c3a8a5497f3933d5d5659a80b1eae72e1b81fea7f8bc6c2b

Observation c1177cdb-2798-4a64-9447-616a5d833687 · outbound

This paper cites Thompson Sampling: An Asymptotically Optimal Finite Time Analysis.

Asymptotically optimal regret in communicating Markov decision processes Thompson Sampling: An Asymptotically Optimal Finite Time Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.699755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.699755Z digest=sha256:25374a457024d0b2250ae81fcb14a73cdcff96fdfcc7a710abbbf4ae3f74725a

Observation 59780be3-1494-4337-9bc4-08bd81a43143 · outbound

This paper cites Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities.

Asymptotically optimal regret in communicating Markov decision processes Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.867553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.867553Z digest=sha256:d66ac587d58d4aafbb23b97cbb4da42770edbd62347d41c3474334a74d2fe978

Observation c8f64276-0b8f-4855-882f-859cd0eab264 · outbound

This paper cites 2 2.1.1 Randomized policies, their gain, bias & gap functions.

Asymptotically optimal regret in communicating Markov decision processes 2 2.1.1 Randomized policies, their gain, bias & gap functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.230041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:53.055380Z digest=sha256:2ac314312e047457fd6d5d1f6a1ef8752505496c61e11af4d1bc508dd0a7ebd7

Observation 4f545917-448a-4995-90a0-eae6717909f2 · outbound

This paper cites Shipra Agrawal and Navin Goyal.

Asymptotically optimal regret in communicating Markov decision processes Shipra Agrawal and Navin Goyal

Reference 1988

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:41:54.076804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:52.015182Z digest=sha256:30cbd28f06d287879e1a79f5bb1d99a7aad95647e98fff4c975ff03fe3e60e77

Observation e787acf3-7951-4e31-b64d-2be6d3e1ae14 · outbound

This paper cites OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September.

Asymptotically optimal regret in communicating Markov decision processes OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.432856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:52.388109Z digest=sha256:bc4ea91ae9827e59262e921cc664d2472cbbc37ca20e3b715696f8fbce2eb39a

Observation aa5e2e6b-1d62-480e-ada3-dfe14b0cbdb4 · outbound

This paper cites Optimism in Reinforcement Learning and Kullback-Leibler Divergence.

Asymptotically optimal regret in communicating Markov decision processes Optimism in Reinforcement Learning and Kullback-Leibler Divergence

Reference 2010

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:41:53.465716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:52.449262Z digest=sha256:1b399e9a3e7c3bec8c44bd0fab0feb89d859e13bfd2549afd80adec25ada3ed5

Observation 4dcb854f-c078-4950-baf7-133786970846 · outbound

This paper cites Analysis of Thompson Sampling for the multi-armed bandit problem.

Asymptotically optimal regret in communicating Markov decision processes Analysis of Thompson Sampling for the multi-armed bandit problem

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.078339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.078339Z digest=sha256:5262b53046052670b0bcf4f3114243c591e34efebe551a849af19110bb7f66bd

Observation f77ad02a-93d7-4232-9cd3-b69d95b5a342 · outbound

This paper cites Improved Analysis of UCRL2 with Empirical Bernstein Inequality.

Asymptotically optimal regret in communicating Markov decision processes Improved Analysis of UCRL2 with Empirical Bernstein Inequality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.529881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.529881Z digest=sha256:fc2d6b1f329fc474f2194d8599b6914989941fd3ee0fa9e82b15694d520e0458

Observation 25ad5aad-bcea-4ec1-a0a9-b4747667d607 · outbound

This paper cites Regret Analysis in Deterministic Reinforcement Learning.

Asymptotically optimal regret in communicating Markov decision processes Regret Analysis in Deterministic Reinforcement Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:41:53.279499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:52.936313Z digest=sha256:9595b71bee9a229417fa8127e20ae2eca34932303351ca104af8bd285c98a135

Observation 7f096932-4b03-4d01-afc6-f1398f33e713 · outbound

This paper cites _eprint: 2502.06480.

Asymptotically optimal regret in communicating Markov decision processes _eprint: 2502.06480

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:41:53.838157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:41:52.215124Z digest=sha256:b764d60f05fc7f681d1bae1fd89d91c3c0e7a6945a871f41301e73ca4325368d

Pith citing papers

No inbound Pith citation observations are available.