Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T01:20:21.181843Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2606.27448.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T01:20:21.181843Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 07e44a8e-f684-4c2a-8b50-4baeb1cbc319 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs neurips.cc/paper_files/paper/2019/file/88fee0421317424e4469f33a48f50cb0-Paper.pdf
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0270dfa6-9c1c-4e49-8df8-678260715a3c · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs The regret lower bound for communicating Markov Decision Processes
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 767d66c0-aad8-4698-a560-33ac382859f2 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Thompson Sampling in Non-Episodic Restless Bandits
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82796a65-172f-4a47-b398-e61fd1541736 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Logarithmic weak regret of non-bayesian restless multi-armed bandit
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0644c29f-677f-4677-a41f-3abe3660a766 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Learning in a changing world: Restless multiarmed bandit with unknown dynamics.IEEE Transactions on Information Theory, 59(3):1902–1916,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e3f11d-0540-4342-bec9-7f8b67a8098c · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Near Sample-Optimal Reduction-based Policy Learning for Average Reward MDP
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a7a8a8ea-be33-49cf-90fe-5610276eef58 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs As a direct consequence of Lemma A.1, the regret can be approximated by ∑T−1 t=1 (g∗−R(t)), where g∗= maxa∈Aga
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3aa586-26cf-464e-b470-16fd06458b68 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Using a change of measure (Appendix C.1, see Theorems C.1 and C.2), we can bind the behavior of the learning algorithm inM1 to its behavior onM2
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f299c58-78d3-45a8-becd-32efe5e06374 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs 31 Initial, trueMa Ma Possible alternativeMε a Mε a s∞ ε R= 1 Figure 6: An illustration of the transformationMε a
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f814b7-1563-4d38-bc34-51b131b4e9c5 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338e0e30-e0fc-47be-86b7-9481704e84a1 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs The dominant termC0 b √ |A|Tlog(T)scales exclusively with the span bound
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5851e5-a193-4011-b2c7-7a4270cbdbe0 · outbound
Learning in Markovian bandits with non-observable states and constrained decision epochs Unresolved cited work
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.