Pith. sign in

Paper Citation Record · LEDGER

Asymptotically optimal regret in communicating Markov decision processes

As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.18064.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18064 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b97480b8-281b-492d-a57e-531ff7779aa3 · outbound

This paper cites The regret lower bound for communicating Markov Decision Processes.

Asymptotically optimal regret in communicating Markov decision processes The regret lower bound for communicating Markov Decision Processes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.317140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.317140Z digest=sha256:fddfa8dc1b9e12072550107e1ae61b4eef7e52336781cc1ff6f9da7d8162f05f

Observation c1177cdb-2798-4a64-9447-616a5d833687 · outbound

This paper cites Thompson Sampling: An Asymptotically Optimal Finite Time Analysis.

Asymptotically optimal regret in communicating Markov decision processes Thompson Sampling: An Asymptotically Optimal Finite Time Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.699755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.699755Z digest=sha256:4245f03ce01e11683b23233c8435cda86c8b70438e66faa99338eb25ec134464

Observation 59780be3-1494-4337-9bc4-08bd81a43143 · outbound

This paper cites Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities.

Asymptotically optimal regret in communicating Markov decision processes Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.867553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.867553Z digest=sha256:47952b55cc6ffb9cce4fda7e904b087f99d1f12d5cb389b8fda5fd5b17b138d3

Observation c8f64276-0b8f-4855-882f-859cd0eab264 · outbound

This paper cites 2 2.1.1 Randomized policies, their gain, bias & gap functions.

Asymptotically optimal regret in communicating Markov decision processes 2 2.1.1 Randomized policies, their gain, bias & gap functions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.230041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:53.055380Z digest=sha256:c553596eb3cb615359f8e5fa4585cc783650e083f1e8db517a95afa5a1d393d1

Observation 4f545917-448a-4995-90a0-eae6717909f2 · outbound

This paper cites Shipra Agrawal and Navin Goyal.

Asymptotically optimal regret in communicating Markov decision processes Shipra Agrawal and Navin Goyal

Reference 1988

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T14:41:54.076804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:52.015182Z digest=sha256:e8e5e402a4458c8a16d7ffdb4c025baeef6c3490961c4a8d5535db6c243545c7

Observation e787acf3-7951-4e31-b64d-2be6d3e1ae14 · outbound

This paper cites OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September.

Asymptotically optimal regret in communicating Markov decision processes OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:41:54.432856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:52.388109Z digest=sha256:d84fdcd60050812bbf637001017eebbdf42e0531881c2da97352e43472a56954

Observation aa5e2e6b-1d62-480e-ada3-dfe14b0cbdb4 · outbound

This paper cites Optimism in Reinforcement Learning and Kullback-Leibler Divergence.

Asymptotically optimal regret in communicating Markov decision processes Optimism in Reinforcement Learning and Kullback-Leibler Divergence

Reference 2010

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:41:53.465716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:52.449262Z digest=sha256:e24e394e5f3598c2b8da2087ae5de45ae53f458e955be7d1ccba2a7c9eac1231

Observation 4dcb854f-c078-4950-baf7-133786970846 · outbound

This paper cites Analysis of Thompson Sampling for the multi-armed bandit problem.

Asymptotically optimal regret in communicating Markov decision processes Analysis of Thompson Sampling for the multi-armed bandit problem

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.078339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.078339Z digest=sha256:cbd159d2d657902712d3b3f5d82ddc67a660464ae2e55bde0f891d6824db79f1

Observation f77ad02a-93d7-4232-9cd3-b69d95b5a342 · outbound

This paper cites Improved Analysis of UCRL2 with Empirical Bernstein Inequality.

Asymptotically optimal regret in communicating Markov decision processes Improved Analysis of UCRL2 with Empirical Bernstein Inequality

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:41:52.529881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:41:52.529881Z digest=sha256:06c3e93f0553ae11deb42ae61d45d00d1ea6d39732e4df5fe276c2aa6c99807e

Observation 25ad5aad-bcea-4ec1-a0a9-b4747667d607 · outbound

This paper cites Regret Analysis in Deterministic Reinforcement Learning.

Asymptotically optimal regret in communicating Markov decision processes Regret Analysis in Deterministic Reinforcement Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T14:41:53.279499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:52.936313Z digest=sha256:935546dcc2b7b48f0d8704e3c6b770388f05bbd1ee7b84cb19246879f59d0837

Observation 7f096932-4b03-4d01-afc6-f1398f33e713 · outbound

This paper cites _eprint: 2502.06480.

Asymptotically optimal regret in communicating Markov decision processes _eprint: 2502.06480

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:41:53.838157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:41:52.215124Z digest=sha256:4bd1a38166f0d09bf64f84e4dc2fe10fb3fcfd6330d73481e50975fe9e7edbbe

Pith citing papers

No inbound Pith citation observations are available.