Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2505.18064.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:41:53.055380Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
11 of 11 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b97480b8-281b-492d-a57e-531ff7779aa3 · outbound
Asymptotically optimal regret in communicating Markov decision processes The regret lower bound for communicating Markov Decision Processes
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1177cdb-2798-4a64-9447-616a5d833687 · outbound
Asymptotically optimal regret in communicating Markov decision processes Thompson Sampling: An Asymptotically Optimal Finite Time Analysis
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59780be3-1494-4337-9bc4-08bd81a43143 · outbound
Asymptotically optimal regret in communicating Markov decision processes Near-optimal Optimistic Reinforcement Learning using Empirical Bernstein Inequalities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f64276-0b8f-4855-882f-859cd0eab264 · outbound
Asymptotically optimal regret in communicating Markov decision processes 2 2.1.1 Randomized policies, their gain, bias & gap functions
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f545917-448a-4995-90a0-eae6717909f2 · outbound
Asymptotically optimal regret in communicating Markov decision processes Shipra Agrawal and Navin Goyal
Reference 1988
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e787acf3-7951-4e31-b64d-2be6d3e1ae14 · outbound
Asymptotically optimal regret in communicating Markov decision processes OptimisminReinforcementLearningand Kullback-Leibler Divergence.2010 48th Annual Allerton Conference on Communication, Control, and Computing (Allerton), pages 115–122, September
Reference 2006
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aa5e2e6b-1d62-480e-ada3-dfe14b0cbdb4 · outbound
Asymptotically optimal regret in communicating Markov decision processes Optimism in Reinforcement Learning and Kullback-Leibler Divergence
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4dcb854f-c078-4950-baf7-133786970846 · outbound
Asymptotically optimal regret in communicating Markov decision processes Analysis of Thompson Sampling for the multi-armed bandit problem
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77ad02a-93d7-4232-9cd3-b69d95b5a342 · outbound
Asymptotically optimal regret in communicating Markov decision processes Improved Analysis of UCRL2 with Empirical Bernstein Inequality
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25ad5aad-bcea-4ec1-a0a9-b4747667d607 · outbound
Asymptotically optimal regret in communicating Markov decision processes Regret Analysis in Deterministic Reinforcement Learning
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f096932-4b03-4d01-afc6-f1398f33e713 · outbound
Asymptotically optimal regret in communicating Markov decision processes _eprint: 2502.06480
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.