Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T15:44:02.195783Z
Paper Citation Record · LEDGER
As of 31 July 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2605.05791.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-09T15:44:02.195783Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b1ae8dfb-5bb3-45b2-836f-d570da5bf619 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation c1a8919b-3210-4821-bd10-b630a18ad11f · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Near-optimal regret bounds for reinforcement learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation f3091fcb-e355-45fc-9025-b40a1df186cd · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Minimax regret bounds for reinforcement learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 897c28ae-8703-4ca9-94d0-3a7791c9a819 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Mastering the game of no-press diplomacy via human-regularized reinforcement learning and planning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 519f5cf9-f73b-4aba-af13-f9f783e9f9c4 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bartlett, Dylan J
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation fe5d2fbb-c30c-42f0-b650-7da941acc247 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 74e84a5d-9914-4a6f-8784-ed360cbf9e41 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas and Steven E
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation e1ca5152-09ed-4f2a-afb5-e3bb032a6975 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas and John N
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 32012686-8901-42c7-a6fd-bc6521fe02bb · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Gomes, and Kilian Q
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation b2dc5a1e-3e65-452b-a2ae-2830c7cba317 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Discounted dynamic programming
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 743b01f3-01a5-40bb-8a12-5460e851465e · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration R-max - a general polynomial time algorithm for near-optimal reinforcement learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation ccf93b47-2938-4a21-808c-62470bd3fdc8 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 0b485bc1-202e-4432-8b73-6a11b0ebabd1 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Magnetic control of tokamak plasmas through deep reinforcement learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 521cd7da-d148-47e8-b606-7e0c51d893aa · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernel-based reinforcement learning: A finite-time analysis
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation cd6d747d-3355-43d1-bb20-cf1fe3879fc6 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Provable model-based nonlinear bandit and reinforcement learning: Shelve optimism, embrace virtual curvature
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 5de3bf24-f4b3-48ab-aae0-e3d39cd678a3 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bilinear classes: A structural framework for provable generalization in RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 60a2d0d4-885a-4c57-9870-f797946f707f · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Risk bounds and Rademacher complexity in batch reinforcement learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 9e9d45e5-f20d-4fde-af40-c4c9b569229c · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Reinforcement learning with Gaussian processes
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation c68325f5-ed2e-432e-8f6c-b168138d2f8b · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Human-level play in the game of diplomacy by combining language models with strategic reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 440a82da-5ae4-454f-98e1-c858015aad24 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Error propagation for approximate policy and value iteration
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation ba043d90-4c2a-4c08-96df-6dbe61577529 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Average cost Markov decision processes with weakly continuous transition probabilities
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation c1d7bca8-0efd-4e66-84e4-bf20ff7760ce · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration The Statistical Complexity of Interactive Decision Making
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation ef247747-d978-48d4-923f-262f1b6d978f · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d073014f-50a4-4efd-ac03-72d647e5a069 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Addressing function approximation error in actor-critic methods
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 547d289d-815f-4b06-9658-1761e0918c2f · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration A theory of regularized Markov decision processes
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 0a432abb-cf91-4f23-9bfc-7f7c24d6074d · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Spectral normalisation for deep reinforcement learning: An optimisation perspective
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 50907821-f764-4d94-b49a-70addb183c10 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Size-independent sample complexity of neural networks
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 42f288cb-5c6a-4ce7-98ab-28fa68b556e2 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Stable function approximation in dynamic programming
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 76583532-a526-4ac3-bd59-58110d08cdd5 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation af4830cb-a08e-4152-b305-fa1c5c79c06e · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Lasserre
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation a90d1bae-713b-45ec-873c-37019f51dd2d · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Lasserre
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 70f94b68-9692-4930-93be-636ede5e6c9a · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Contextual decision processes with low Bellman rank are PAC -learnable
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 03f46a25-2e7e-4505-af21-432566a8218e · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Provably efficient reinforcement learning with linear function approximation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 6c2599e3-9136-4690-af11-038d7bdfc72e · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 113fb8fb-a674-48c0-9e9a-c517096b1780 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Approximately optimal approximate reinforcement learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 2d5f05e2-c996-4695-baaf-bde1e71e2355 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Near-optimal reinforcement learning in polynomial time
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 6d0aa6fc-f2e3-415a-aa59-8fa841c5e5bb · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Spectral normalization for generative adversarial networks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation b596eccb-cb0b-4bda-b60f-1a912822e294 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Human-level control through deep reinforcement learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d24ea03e-cc84-4a5d-8e64-b4bd769686cd · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Error bounds for approximate policy iteration
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 7b8d8a5a-905c-4e10-a5c4-e895b104da35 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Performance bounds in L ^p -norm for approximate value iteration
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 10e415ad-38d0-42f6-b0aa-8812e5788d82 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Finite-time bounds for fitted value iteration
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation f663f7d3-71a1-4558-a26b-58661a259ff4 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernel-based reinforcement learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 78637ddd-93f6-4705-ba10-1b5f14686acf · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration (more) efficient reinforcement learning via posterior sampling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 25a8eccf-0410-44ce-a06e-3e0620701e08 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Online learning via sequential complexities
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation f42dfdb5-4875-47a0-b53b-0fe68eb0f242 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 94eb0723-8515-47f2-9b54-0aec6e0442d7 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Eluder dimension and the sample complexity of optimistic exploration
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation a7cbdacb-8cf1-4ca9-b6e0-bb3da04c50cd · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration a l. Conditions for optimality in dynamic programming and for the limit of n-stage optimal policies to be optimal. Zeitschrift f \
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation ecc6379e-502f-42af-b0e0-72bb08616b95 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Approximate modified policy iteration and its application to the game of Tetris
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 8bc49cab-d57c-4a2d-bd5a-2882f791dec0 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Devon Hjelm, Aaron Courville, and Philip Bachman
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation c40c3e13-b940-479f-b4b8-910b07a1e96c · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Mastering the game of Go with deep neural networks and tree search
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d0a75709-9d72-4536-8268-918a4f0aefef · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration CURL : Contrastive unsupervised representations for reinforcement learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 6980bb18-66ce-4320-b1fa-35dfff5a9328 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Decoupling representation learning from reinforcement learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 7d458331-e9d1-4369-94aa-c5b959e080e5 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Negative dynamic programming
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation f6f0b7c5-799d-4330-bf46-20ddc45462b6 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Optimistic posterior sampling for reinforcement learning with few samples and tight guarantees
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 4958b48e-bf46-4584-b4b8-f3ff4cf7bf80 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration From Dirichlet to Rubin : Optimistic exploration in RL without bonuses
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation caed532b-19a6-421b-8213-94b18ecd769f · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Analysis of temporal-difference learning with function approximation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 40ff3a06-8940-44aa-b17e-a03171e6164b · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernelized reinforcement learning with order optimal regret bounds
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 39aa3ebf-9a96-4592-bf18-bf2d18999fbe · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Grandmaster level in StarCraft II using multi-agent reinforcement learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 637aad1f-b364-4135-b59e-25595319e3e7 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation d696f753-80d6-421b-93f7-ac88e163b2e1 · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
Observation 12d497a9-3b6f-4ef9-92dd-730aae09ecdd · outbound
A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Offline reinforcement learning with realizability and single-policy concentrability
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.
No inbound Pith citation observations are available.