Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:54:13.924295Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:1909.02506.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:54:13.924295Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2da5c4f8-32bb-4c1c-bdd2-7da15c1aafc3 · outbound
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2f559d39-bb65-4437-b035-86153a337c30 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Efficient optimal learning for contextual bandits
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5765909d-2b49-40bf-9eca-ff4ab1211e87 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Dynamic programming and optimal control , volume 1
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 60f2e34c-8475-4f9c-b361-b942afa71e2e · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Approximate Dynamic Programming: Solving the curses of dim ensionality, volume
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9aa0314a-1dfb-46f4-9f31-76c614c0cf63 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Human-level control through deep reinforcement learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3b5f717-4263-45b2-b6bf-bd38c4a7d0b2 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Unifying count-based exploration and intrinsic motivation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 985da6e6-a9ab-442c-ba7e-e40460a9059c · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Mastering the game of go with deep neural networks and tree search
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 60df2f6d-336c-4dba-845a-2f26467e27b2 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Ma stering the game of go without human knowledge
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ab1668e3-a944-41da-9d4f-f4008e991b56 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6d76ad9-7b13-42ce-85ed-7ef61dd5d296 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Is Q-learning provably efficient? In Proceedings of Advances in Neural Information Processing S ystems (NeurIPS) , pages 4863–4873,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7541a362-e33a-4c63-b5bc-e4617e947fd7 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Tighter Problem-Dependent Regret Bounds in Reinforcement Learning without Domain Knowledge using Value Function Bounds
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6ff1699-9e4d-4039-a20c-1bfea857f4f2 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Min imax regret bounds for reinforcement learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6457e25c-b60e-4e10-aab5-a021349f3c44 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Pac model-free reinforcement learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 53dad104-a1ff-48bf-b438-5895be4eede5 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Speedy q-learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3710c4c3-634f-44ba-8bd7-1d9bd8e7c6f3 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning rates for q-lear ning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 012aa6c1-db3b-4f91-b3da-ced392c3c3af · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Variance re duced value iteration and faster al- gorithms for solving markov decision processes
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f99b3821-9f17-4986-b7ab-706eb62a1154 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank An analysis of bid-price con trols for network revenue manage- ment
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 257226e7-a698-4515-b100-85dfb2a514f3 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Dynamic bid prices in revenue management
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b31dc2dc-f577-4717-b60e-64d439a354da · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Asynchronous metho ds for deep reinforcement learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4a48be75-ec48-4686-883c-ce9ec514f066 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank The malmo platform for arti- ficial intelligence experimentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a09365b0-df07-499c-9f80-b48ac7c68536 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Gambling in a rigged casino: The adversarial multi-armed bandit problem
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0aa6861f-a9e4-4934-9c9d-2c2a5a5c2caa · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning from delayed rewards
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9b4beea-ee82-4b7a-9ff5-064f5027c09f · outbound
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 543ac82f-4627-4459-b740-67a81400fc24 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Asynchronous stochastic approximation and q -learning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa802055-7200-4088-b23e-ee942157aadc · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Regu- larized policy iteration with nonparametric function spaces
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation be1358dd-9d80-4b2a-aa15-aa4a969a9253 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Finite-sample analysis of least- squares policy iteration
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1bdd4641-48d3-470b-b6f8-d777a05c04ac · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ada3838-90ba-47e3-a4c1-a75e3b3722d7 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Finite-time bounds for fit ted value iteration
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 056c1862-9c91-4e10-a7de-62d4e0d4e10f · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Information-theoretic consideratio ns in batch reinforcement learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0002a454-f72e-466d-ac60-e94f78eff01d · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Stable function approximation in dynamic pro gramming
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 70db911d-5bc8-49d7-9b49-3b148c7fdae3 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Provably Efficient Reinforcement Learning with Linear Function Approximation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d6e05e9-b7c9-4ff2-abcb-f222d4938260 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On oracle-efficient pac rl with rich observations
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c6df8ebc-569a-4987-95c3-27c4c4c98a16 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, M iroslav Dudik, and John Langford
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 96047cdd-2dd3-4e8d-b827-f1a3f13fc13c · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Efficient reinforcement learnin g in deterministic systems with value function generalization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4ea3b8f1-d38c-4c25-8cb4-a165513d1543 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank N ear-optimal time and sample complexities for solving discounted markov decision process with a ge nerative model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7efce2da-627c-4d07-a012-eb34ddf29243 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Glo bal convergence of policy gradient methods for linearized control problems
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48344bfe-79b9-435f-a0a5-561279b0772f · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Stochastic approximation with cone-contractive operators: Sharp $\ell_\infty$-bounds for $Q$-learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7573e80-b58a-4011-874b-6b5b620b634d · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Taming the monster: A fast and simple algorithm for contextual bandits
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d86242a7-485e-42aa-905c-954229a22d59 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret Bound
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 51440e06-2d73-446b-9d59-e3b16f2cbdea · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On learning sets and functions
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d12b1613-a872-4dc4-b587-ae450c1e7b5c · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Decision theoretic generalizations of the pac mo del for neural net and other learning applications
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05a60069-8a0e-4796-9b03-781f3fdfc24a · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Rates of convergence in the central limit the orem for empirical processes
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c2f41ed6-565d-470f-a3d2-221fc359dfa0 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On general minimax theorems
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4eb87988-5959-4ab2-b515-148935a61a47 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank On minimum volume ellipsoids containing part of a give n ellipsoid
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7aa974c-3be4-41e0-9f41-d2f8ae213fef · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Sphere packing numbers for subsets of the bo olean n-cube with bounded vapnik- chervonenkis dimension
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c84bc67-97d4-4741-84f9-54cdb046dbd1 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Springer Science & Business Media, 2013
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aced5f6a-38ae-4bb1-adfe-07eddab1dd2a · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Convergence of stochastic processes
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1d8be10-b6d3-4b4e-9381-4490a608b6b0 · outbound
$\sqrt{n}$-Regret for Learning in Markov Decision Processes with Function Approximation and Low Bellman Rank Unresolved cited work
Reference 703
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.