Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:27.181643Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.19854.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T11:40:27.181643Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4801e300-7b59-41b0-8bfe-f7a0fcf0f73d · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c269259-52a8-433e-bf90-68ac84d15294 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Minimax regret bounds for reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fce1e969-1212-4c88-9cc1-30a58ceec6a6 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Regal: a regularization based algorithm for reinforcement learning in weakly communicating mdps
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab63c10-7f08-477c-abc8-699b27b6aeab · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Brafman and Moshe Tennenholtz
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9174a010-34e2-4312-9c41-61633472f8a2 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Provably Efficient Exploration in Policy Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e614e620-efcd-49f5-8143-c4dec2944f3a · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Implicit finite-horizon approximation and efficient optimal algorithms for stochastic shortest path.Advances in Neural Information Processing Systems, 34, 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15128eb3-93ae-4a2b-b3f8-ae7da0dcd47e · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Sample complexity of episodic fixed-horizon reinforcement learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00b6e90a-50bb-482a-ba23-3db55c9c7e5c · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31b74b9f-edf2-4c00-9f0f-fa38aafa09cd · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Policy certificates: Towards accountable reinforcement learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c7bf4ac-ee8b-40aa-8ac7-de0c812c5dc2 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5faac2ce-e8c0-48ab-846a-259c55210303 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near optimal exploration-exploitation in non- communicating markov decision processes
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e095135e-2139-42b5-83c1-2f0f2157a5a9 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal regret bounds for reinforcement learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c6d83dd-0910-48af-a758-a4812b4c8f5b · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Open problem: The dependence of sample complexity lower bounds on planning horizon
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c7db996-bc4e-4ab8-8353-0bc018a3330a · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is Q-learning provably efficient? InAdvances in Neural Information Processing Systems, pages 4863–4873, 2018
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98353e53-0d49-439a-8d18-a6ecee66b4d0 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PhD thesis, University of London London, England, 2003
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 489bb316-9a4c-4425-a6ed-eabe06adb878 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-optimal reinforcement learning in polynominal time
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7e42a6-9880-4103-857a-5b5834278656 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-bayesian exploration in polynomial time
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2da9399-ed90-4a4c-970b-8993b1be5470 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Pac bounds for discounted mdps
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59bf1383-2265-46ad-8a9c-9bb2823128b5 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Breaking the sample complexity barrier to regret-optimal model-free reinforcement learning.Advances in Neural Information Processing Systems, 34, 2021
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ca7571-374a-4447-bd9f-0c3aaceb0e04 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free learning for Markov decision processes and games: Stochasti- cally bounded rewards and improved bounds
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff56f70c-8e43-440a-ada9-b1c6b75b78d2 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Settling the horizon-dependence of sample complexity in reinforcement learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd057107-90c4-4b3a-a149-94e9ae10c6db · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Empirical Bernstein bounds and sample variance penalization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03d8f37b-4764-4631-8a16-006a14384333 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Ucb momentum q-learning: Correcting the bias without forgetting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f199159-8fd5-4a68-8c4a-4f7a24c25b5c · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence A Unifying View of Optimism in Episodic Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25829e35-b5ec-41be-90bf-7e226d8fc955 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence (more) efficient reinforcement learning via posterior sampling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245af5ec-f3da-421a-ba75-7cf266ec160d · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Why is posterior sampling better than optimism for reinforcement learning? InProceedings of the 34th International Conference on Machine Learning-Volume 70
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a3e0ad-2587-48ca-ac34-ade72daa83ed · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Towards Tractable Optimism in Model-Based Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245c824c-5e34-434d-898a-a73e470f401f · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Nearly horizon-free offline reinforcement learning.Advances in neural information processing systems, 34, 2021
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2f5be01-be32-4e1e-b9a4-f79376a246d1 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Worst-case regret bounds for exploration via randomized value functions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8d873f-f2c7-4b34-90fe-a89e81992529 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Non-asymptotic gap-dependent regret bounds for tabular MDPs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b6acb7-2315-4367-8bdb-d6e9152453a7 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence PAC model-free reinforcement learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc7cf24-c48e-4411-acde-1175eb5923a7 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence An analysis of model-based interval estimation for markov decision processes.Journal of Computer and System Sciences, 74(8):1309–1331, 2008
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59520a41-9a07-4041-af11-e2acddf13094 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Model-based reinforcement learning with nearly tight exploration complexity bounds
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18ebd5e-b905-4ede-a069-f6526fb93e1a · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-Aware Regret Bounds for Undiscounted Reinforcement Learning in MDPs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f2af213-6957-4163-b8e0-5c1528a40e63 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Stochastic shortest path: Minimax, parameter-free and towards horizon-free regret.Advances in Neural Information Processing Systems, 34, 2021
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d572841b-2c0e-452e-bc99-aa6a402735a8 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is long horizon reinforcement learning more difficult than short horizon reinforcement learning? InAdvances in Neural Information Processing Systems, 2020
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041767f7-ac81-47bd-a7a1-123a5d521c32 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Near-Optimal Randomized Exploration for Tabular Markov Decision Processes
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de55941f-e943-4f1b-875e-15a01520adff · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence $Q$-learning with Logarithmic Regret
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0baaf33-cc64-4e9b-afe5-77e15b5ab775 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Is reinforcement learning more difficult than bandits? a near-optimal algorithm escaping the curse of horizon
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c354c13d-6c92-473b-8046-e54aba83750e · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Horizon-free reinforcement learning in polynomial time: the power of stationary policies
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06a19116-357c-426b-8a81-1be9e439aa0f · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Variance-aware confidence set: Variance- dependent bound for linear bandits and horizon-free bound for linear mixture mdp
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e022117d-be36-46b2-9812-c2b709538011 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence Almost optimal model-free reinforcement learning via reference-advantage decomposition
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bf24377-46c4-429f-a828-bbbcb4807ab2 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The horizon-truncation lemma implies that the optimal H-step value is withinOpυqof the optimalH 1-step value
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f34ac99-42d8-4c7e-927e-d1aa34d12138 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The proof is a backward induction on h
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65ed91f3-fcdf-41d2-8e0d-e25b251464ab · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence The process is stopped when the trajectory reaches the unlearned set Ok or when the frozen count of a learned pair doubles
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4e9a66-26ff-4174-b235-29c98665c644 · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence This requires a variance closure argument for the optimistic values and the optimistic gaps
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d6f6be-6c0f-41bb-ac7e-04de21c57a2d · outbound
Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence V ˚ H1`1
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.