Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2506.06521.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:09:21.939768Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T18:10:52.644951Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-09T06:45:40.553981Z
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 752a21c9-eded-434d-be7b-0f168f5d613f · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Navigating to the best policy in markov decision processes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e9af9e9-27e2-47d8-8e8b-e7c0ff6e7545 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic online regret bounds for undiscounted reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f7ec633-71fd-414e-bb52-31eb330fa6ce · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Finite-time analysis of the multiarmed bandit problem
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1560998c-4618-4634-a281-39ba3a04691c · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal regret bounds for reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5655449-346d-46ee-9f2c-2fe69a778fac · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Minimax regret bounds for reinforcement learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54fec24-9f68-4cf5-aa49-5f8fe58a873c · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs REGAL: A Regularization based Algorithm for Reinforcement Learning in Weakly Communicating MDPs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c29630a4-0f78-40e9-887c-b334b8b794a4 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3663fe-4d78-4094-b5ae-d7cf90cf1a0d · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Top-k off-policy correction for a reinforce recommender system
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1ac95b1-9f22-4f96-ae91-10a8d6604ec1 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-Aware Sparse Linear Bandits
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cec2c58e-e241-40a0-9c94-fdd2d08dfd2b · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Policy certificates: Towards accountable reinforcement learning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6fa8dfe2-4779-4705-9abb-c97da5ade988 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond value-function gaps: Improved instance-dependent regret bounds for episodic reinforcement learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60b87a59-dc4d-49ed-a24e-ebb846250806 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-dependent bounds for two-player markov games
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a168d2be-fb72-4c2d-a587-6dbfc37b7b0a · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Cascaded gaps: Towards logarithmic regret for risk-sensitive reinforcement learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d22e8a7b-3d5e-480c-8a43-fe6b483cb5ad · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Efficient bias-span-constrained exploration-exploitation in reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec4381ac-2eb4-48cf-9940-8968b0f5725d · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Logarithmic regret for reinforcement learning with linear function approximation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0076f097-b362-4663-b39d-45d547528fa1 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tackling heavy-tailed rewards in reinforcement learning with function approximation: Minimax optimal and instance-dependent regret bounds
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c146e2b-e269-42fb-9c88-5f62ed713de8 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is q-learning provably efficient? Advances in neural information processing systems, 31, 2018
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf290a2-8bd0-4d6a-927c-49934c46ac58 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reward-free exploration for reinforcement learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4001edbd-21f0-4c19-b55c-fc39d77c45e5 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Planning in markov decision processes with gap-dependent sample complexity
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e9893598-862c-4221-92c3-14e79de722bd · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved regret analysis for variance-adaptive linear bandits and horizon-free linear mixture mdps
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3b9b6b2-9766-4c14-9ef0-847e3dce72d3 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Asymptotically efficient adaptive allocation rules
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e475c0b6-34e8-4d90-b0ca-53a9292c24d2 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0374572d-fe69-4408-9c01-0759e382bb4f · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Continuous control with deep reinforcement learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a8208e-7d51-4449-b65d-cd43cbc729b9 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Deep reinforcement learning for dynamic treatment regimes on medical registry data
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfa9340b-fa18-408b-bcbb-1d0b634b2f37 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Adaptive Sampling for Best Policy Identification in Markov Decision Processes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82315765-0338-43ad-bf3a-22ed3b6f355b · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Empirical Bernstein Bounds and Sample Variance Penalization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da871a9a-82bc-4c25-9641-0600ed1a9997 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Ucb momentum q-learning: Correcting the bias without forgetting
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e4070e2-adca-46e5-9b69-c550d4538112 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning for optimized trade execution
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 840d6741-43ff-467c-9513-0fb223349c7e · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On instance-dependent bounds for offline reinforcement learning with linear function approximation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83035fc7-3068-41b5-9a09-6380cec19577 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Exploration in structured reinforcement learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 143457cc-037a-4666-9f5a-426af5f8d651 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Why is posterior sampling better than optimism for reinforcement learning? In International conference on machine learning, pages 2701--2710
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc82e5fe-f4e1-49a8-9b08-93ef8f346b0c · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning in linear mdps: Constant regret and representation selection
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fdcbba15-75cf-4a2c-8f4e-d664249465b1 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Mastering the game of go with deep neural networks and tree search
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0afb3a-986d-4ccd-8bb4-e9c44f172cfb · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Non-asymptotic gap-dependent regret bounds for tabular mdps
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f03b79-cce4-4b1a-a750-b79a831414fd · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning: An introduction, volume 1
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9bbdc6-dc6b-45b5-b181-dd2192b8093f · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Variance-aware regret bounds for undiscounted reinforcement learning in mdps
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7757793-65e0-4df2-a5bc-6aca707c4cd1 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic linear programming gives logarithmic regret for irreducible mdps
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87b74cce-14e2-47e2-a05b-8814cb258c4c · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near instance-optimal pac reinforcement learning for deterministic mdps
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 646c21e6-60f6-4282-ba5a-dbe02e9614a8 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Optimistic pac reinforcement learning: the instance-dependent view
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 665fbd0f-d338-43cc-89fe-2f62a84c2986 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Reinforcement learning with logarithmic regret and policy switches
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4fba1ab-3745-4ef1-9a60-f3f745eabfc8 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Instance-dependent near-optimal policy identification in linear mdps via online experiment design
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c96a5531-24de-4071-ac6b-bb1c7376ee50 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs First-order regret in reinforcement learning with linear function approximation: A robust estimation approach
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 343a71d3-ab56-4d4c-a37e-13f8cab7b505 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Beyond no regret: Instance-dependent pac reinforcement learning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 587e52c9-d96d-4887-80a9-2f0250ffdba0 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs On gap-dependent bounds for offline reinforcement learning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c2295a7-13c1-4934-a9f2-7106bdeb9912 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Near-optimal randomized exploration for tabular markov decision processes
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee3fccd1-e11e-4db4-9784-a506b0ca06b9 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Fine-grained gap-dependent bounds for tabular mdps via adaptive multi-step bootstrap
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c561d791-c61e-41ee-922e-0c6d2e04dd84 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Q-learning with logarithmic regret
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e09e3872-2afb-4001-84da-90d806ce96b6 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fdf739b-4ba0-4d61-acb9-a387c1a54368 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Regret minimization for reinforcement learning by evaluating the optimal bias function
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03367650-80df-4da6-97d3-4005db1289cb · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Almost optimal model-free reinforcement learningvia reference-advantage decomposition
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a62ee762-7be0-4634-a9f7-806d703c4041 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb10c7a5-ddd7-49d9-9810-6017f98d517a · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Improved variance-aware confidence sets for linear bandits and linear mixture mdp
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 437ff0ee-9e1e-4079-b4c7-1dd3ac6a4efa · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 744f8b4a-6ecc-44b7-bd4a-b4cd2893f646 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Settling the sample complexity of online reinforcement learning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f3bd36-1d48-4a40-8458-ac15aa3701e9 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a74689ed-21ec-46a7-a988-4e5560f21227 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46abb972-bf09-441b-a206-dd59d2d18436 · outbound
Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs Sharp variance-dependent bounds in reinforcement learning: Best of both worlds in stochastic and deterministic environments
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229aedfe-0a81-41e4-9f90-efbe3fc4d041 · inbound
On-line Learning in Tree MDPs by Treating Policies as Bandit Arms Sharp Gap-Dependent Variance-Aware Regret Bounds for Tabular MDPs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.