Pith. sign in

Paper Citation Record · LEDGER

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration

As of 31 July 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2605.05791.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05791 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-09T15:44:02.195783Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy58
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1ae8dfb-5bb3-45b2-836f-d570da5bf619 · outbound

This paper cites Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Optimistic posterior sampling for reinforcement learning: Worst-case regret bounds

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.322555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:1ea0b35dd029ad3f2cfb92f680ace09899e1b9a201462584e261a09e435e51cb

Observation c1a8919b-3210-4821-bd10-b630a18ad11f · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Near-optimal regret bounds for reinforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.325556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:bf25fa39cecd93fa45dbc6f31ba8b24ad3dd6da7b9fe700f1ed2c9255aeff2f2

Observation f3091fcb-e355-45fc-9025-b40a1df186cd · outbound

This paper cites Minimax regret bounds for reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Minimax regret bounds for reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.330641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:edde57298396a3833b4d8a632d9e6b00d83cf87194acf0b411c8da28f2dc40c6

Observation 897c28ae-8703-4ca9-94d0-3a7791c9a819 · outbound

This paper cites Mastering the game of no-press diplomacy via human-regularized reinforcement learning and planning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Mastering the game of no-press diplomacy via human-regularized reinforcement learning and planning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.316031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:6087d762a3d398bc717631fa3a15307166af21f2387f73ea8fa9712468b89a14

Observation 519f5cf9-f73b-4aba-af13-f9f783e9f9c4 · outbound

This paper cites Bartlett, Dylan J.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bartlett, Dylan J

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.311009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:bffcc31efb9d32d1f5cf1b7494761cb64a1ee83ca793cb3c78a90b80c838bcc5

Observation fe5d2fbb-c30c-42f0-b650-7da941acc247 · outbound

This paper cites Bertsekas.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.357083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:dfbcb22345f4302384b8131ef77a6cacc0bd8532ecec762c085a31cb347f0fc1

Observation 74e84a5d-9914-4a6f-8784-ed360cbf9e41 · outbound

This paper cites Bertsekas and Steven E.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas and Steven E

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.229472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:17033e30743db03bbb54c748067d5fc2ac60373603276ba839bcda0a370f6227

Observation e1ca5152-09ed-4f2a-afb5-e3bb032a6975 · outbound

This paper cites Bertsekas and John N.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bertsekas and John N

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.288235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:96a31a78b756bd07de7cab77a5a31d1753549dfd36612234190065493eab0e9a

Observation 32012686-8901-42c7-a6fd-bc6521fe02bb · outbound

This paper cites Gomes, and Kilian Q.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Gomes, and Kilian Q

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.232267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:4446195b22140caae9abb1f7f49f3d932e156b796030c2d4c17811b10ed78eeb

Observation b2dc5a1e-3e65-452b-a2ae-2830c7cba317 · outbound

This paper cites Discounted dynamic programming.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Discounted dynamic programming

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.406621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:34ad381b66382afec6a5f3412311723af86d5303e5edf2d9db3f92e978de62bb

Observation 743b01f3-01a5-40bb-8a12-5460e851465e · outbound

This paper cites R-max - a general polynomial time algorithm for near-optimal reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration R-max - a general polynomial time algorithm for near-optimal reinforcement learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.377490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:058a0a084da5db1a17331d72f44303f436f24466a8e8fa616f8b044eb59ede04

Observation ccf93b47-2938-4a21-808c-62470bd3fdc8 · outbound

This paper cites an unresolved cited work.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:06:26.347822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:01a2931711e00e680008ecc77631f966f17172c242c7b361ba835e3f34e0add4

Observation 0b485bc1-202e-4432-8b73-6a11b0ebabd1 · outbound

This paper cites Magnetic control of tokamak plasmas through deep reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Magnetic control of tokamak plasmas through deep reinforcement learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.235676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:901534d63f92ab17e9cc02b2d2941a1f476dda1ce482e23b3a6433879ec4df06

Observation 521cd7da-d148-47e8-b606-7e0c51d893aa · outbound

This paper cites Kernel-based reinforcement learning: A finite-time analysis.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernel-based reinforcement learning: A finite-time analysis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.350928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:0ecc4fcd04a9e0778fbf97d99f632ede940d21f8c5c6fd57d585a6c3bd1b98ae

Observation cd6d747d-3355-43d1-bb20-cf1fe3879fc6 · outbound

This paper cites Provable model-based nonlinear bandit and reinforcement learning: Shelve optimism, embrace virtual curvature.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Provable model-based nonlinear bandit and reinforcement learning: Shelve optimism, embrace virtual curvature

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.223175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:8e62fb4a048f6bef41cfb8e9de641a3f41bb68bf58f3c471ce890b1cdd5672f9

Observation 5de3bf24-f4b3-48ab-aae0-e3d39cd678a3 · outbound

This paper cites Bilinear classes: A structural framework for provable generalization in RL.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bilinear classes: A structural framework for provable generalization in RL

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.245101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:0795b391b6609c6bf1f9ad163e29134e97d3fa63081aab5492f6ae0b09c321db

Observation 60a2d0d4-885a-4c57-9870-f797946f707f · outbound

This paper cites Risk bounds and Rademacher complexity in batch reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Risk bounds and Rademacher complexity in batch reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.398121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:67ddbafecef569608d359eabf07d8726199d939c6edd12ae0d7760929f28e04d

Observation 9e9d45e5-f20d-4fde-af40-c4c9b569229c · outbound

This paper cites Reinforcement learning with Gaussian processes.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Reinforcement learning with Gaussian processes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.272086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:310d1bbb3b376dad45c931c05cc1db027a5810f701e404afdb15d40753a91f07

Observation c68325f5-ed2e-432e-8f6c-b168138d2f8b · outbound

This paper cites Human-level play in the game of diplomacy by combining language models with strategic reasoning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Human-level play in the game of diplomacy by combining language models with strategic reasoning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.266768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:93e06cf7bcd0c2ac0d8484dd97effe12d04c8658a2baa641ddd02bbec06d49fc

Observation 440a82da-5ae4-454f-98e1-c858015aad24 · outbound

This paper cites Error propagation for approximate policy and value iteration.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Error propagation for approximate policy and value iteration

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.395132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:ff4ba6ab4b81088e1a2a3ce173478afc3d87b33ce7dbd911264d15277a079f95

Observation ba043d90-4c2a-4c08-96df-6dbe61577529 · outbound

This paper cites Average cost Markov decision processes with weakly continuous transition probabilities.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Average cost Markov decision processes with weakly continuous transition probabilities

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.291209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:8bb0891e3658792fb7375451b901bea4a414f6d2a23bee766da1db1b0844d8bc

Observation c1d7bca8-0efd-4e66-84e4-bf20ff7760ce · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration The Statistical Complexity of Interactive Decision Making

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:09.943214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:02df30766885f25cc47e77a1d8beac7b53184fbda4de8e8d2e6aeb2f3adca14b

Observation ef247747-d978-48d4-923f-262f1b6d978f · outbound

This paper cites Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Foster, Akshay Krishnamurthy, David Simchi-Levi, and Yunzong Xu

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.409882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:f1796ee823cdbc6bdb341a6da4c7703db6c7d0693b118c745cc3437a5eb89bae

Observation d073014f-50a4-4efd-ac03-72d647e5a069 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Addressing function approximation error in actor-critic methods

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.380275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:f5a06a902aa88657c6f63852863a774dc9397a40d21fbe336d9ed385f4d1e0bc

Observation 547d289d-815f-4b06-9658-1761e0918c2f · outbound

This paper cites A theory of regularized Markov decision processes.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration A theory of regularized Markov decision processes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.251680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:5bc56f0a7ac7f5843b905571f15e4192fc2b86bd470d87daf0d779029ac51ccf

Observation 0a432abb-cf91-4f23-9bfc-7f7c24d6074d · outbound

This paper cites Spectral normalisation for deep reinforcement learning: An optimisation perspective.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Spectral normalisation for deep reinforcement learning: An optimisation perspective

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.388065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:171961a872ffc01203079ad23038576d5a2a4ced5b9f504d2c3a897d5e942a05

Observation 50907821-f764-4d94-b49a-70addb183c10 · outbound

This paper cites Size-independent sample complexity of neural networks.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Size-independent sample complexity of neural networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.299975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:e266894cff11f44c90a519fd79c9564f076fc454dacdb530c224cfc853e2392e

Observation 42f288cb-5c6a-4ce7-98ab-28fa68b556e2 · outbound

This paper cites Stable function approximation in dynamic programming.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Stable function approximation in dynamic programming

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.293876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:826e34753974c0c99445c1ec396ea1385b21719f308b0e27f8c006a82420289c

Observation 76583532-a526-4ac3-bd59-58110d08cdd5 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.304905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:e2a619c08897efca2c8b4b63d1eb5219a35fbf4004b2701bfa6c547aaf450da5

Observation af4830cb-a08e-4152-b305-fa1c5c79c06e · outbound

This paper cites Lasserre.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Lasserre

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.296505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:e404844d308e5e476a41d56ede44d6d7779eb7397ccf88d9288b9d4c66874e2b

Observation a90d1bae-713b-45ec-873c-37019f51dd2d · outbound

This paper cites Lasserre.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Lasserre

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.307965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:f4870545077d8773ec7e1f50e5fa485553330283362ba3582efad136a975aedd

Observation 70f94b68-9692-4930-93be-636ede5e6c9a · outbound

This paper cites Contextual decision processes with low Bellman rank are PAC -learnable.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Contextual decision processes with low Bellman rank are PAC -learnable

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.277490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:d357230508653de65e2a06f9a4ea2eb024d777247aa98cd23335560df2385e01

Observation 03f46a25-2e7e-4505-af21-432566a8218e · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Provably efficient reinforcement learning with linear function approximation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.383125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:911f2ee52f71de6c55b98f41263dce3138ae590f6c34b7b96292a4adaa315afc

Observation 6c2599e3-9136-4690-af11-038d7bdfc72e · outbound

This paper cites Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bellman eluder dimension: New rich classes of RL problems, and sample-efficient algorithms

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.344765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:dfb80ed956b4f15cf113907039dc7dba02e0483361f5da6c8727b741e8081714

Observation 113fb8fb-a674-48c0-9e9a-c517096b1780 · outbound

This paper cites Approximately optimal approximate reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Approximately optimal approximate reinforcement learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.336767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:4cbdf6a435fa6f9dcb4dd35d5507a6af8b51cc4bade8a6280c7e05e032a32a62

Observation 2d5f05e2-c996-4695-baaf-bde1e71e2355 · outbound

This paper cites Near-optimal reinforcement learning in polynomial time.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Near-optimal reinforcement learning in polynomial time

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.367247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:6042182d578307d65e42fef0ff8f140010d472193fe1794455044377b8b5c1ae

Observation 6d0aa6fc-f2e3-415a-aa59-8fa841c5e5bb · outbound

This paper cites Spectral normalization for generative adversarial networks.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Spectral normalization for generative adversarial networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.285287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:1dd73112197ad7936d318af1c5e961d6252673c00be17fb23920124059cb4b99

Observation b596eccb-cb0b-4bda-b60f-1a912822e294 · outbound

This paper cites Human-level control through deep reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Human-level control through deep reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.319310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:8005fb9b283cfedd0da9ec8e5767dacfc38e66c846452718ec2ff555f1c35e81

Observation d24ea03e-cc84-4a5d-8e64-b4bd769686cd · outbound

This paper cites Error bounds for approximate policy iteration.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Error bounds for approximate policy iteration

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.361141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:4ebce1521195ec9921ae209325b58a80a3c592901b8f12e3c706e2dc0bc5d9cb

Observation 7b8d8a5a-905c-4e10-a5c4-e895b104da35 · outbound

This paper cites Performance bounds in L ^p -norm for approximate value iteration.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Performance bounds in L ^p -norm for approximate value iteration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.263917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:775b7e3cec3f3fd6b29f9bc7099435d1af31fedf121a1a5e56ffa0e76c107714

Observation 10e415ad-38d0-42f6-b0aa-8812e5788d82 · outbound

This paper cites Finite-time bounds for fitted value iteration.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Finite-time bounds for fitted value iteration

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.274771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:f4e9bb35db9b14f3464953dda75baad3b64eff68fd2de6a576ebbe7299fae30c

Observation f663f7d3-71a1-4558-a26b-58661a259ff4 · outbound

This paper cites Kernel-based reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernel-based reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.258253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:3f74cf74ede31eb7d38cbfeca55581f455d569b0267ebdf7afc9e85586b66988

Observation 78637ddd-93f6-4705-ba10-1b5f14686acf · outbound

This paper cites (more) efficient reinforcement learning via posterior sampling.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration (more) efficient reinforcement learning via posterior sampling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.371166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:a28f24df60cf87dfc87525b6e1e2431ca0fd65fb76e7c4ea675b52a3735cffc7

Observation 25a8eccf-0410-44ce-a06e-3e0620701e08 · outbound

This paper cites Online learning via sequential complexities.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Online learning via sequential complexities

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.280440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:195f3324327a5152a557673e27085a2d205a059d75c24b7bbaab9e4be4de57b7

Observation f42dfdb5-4875-47a0-b53b-0fe68eb0f242 · outbound

This paper cites Bridging offline reinforcement learning and imitation learning: A tale of pessimism.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Bridging offline reinforcement learning and imitation learning: A tale of pessimism

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.413708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:0415822ad45282256c17c7b74069b6e11086763bd52445b26e273bc90113d167

Observation 94eb0723-8515-47f2-9b54-0aec6e0442d7 · outbound

This paper cites Eluder dimension and the sample complexity of optimistic exploration.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Eluder dimension and the sample complexity of optimistic exploration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.269513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:cdc1c4972b2c16b21f301b4531cd676180a21dc94e3266a7a0ad9938445aed99

Observation a7cbdacb-8cf1-4ca9-b6e0-bb3da04c50cd · outbound

This paper cites a l. Conditions for optimality in dynamic programming and for the limit of n-stage optimal policies to be optimal. Zeitschrift f \.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration a l. Conditions for optimality in dynamic programming and for the limit of n-stage optimal policies to be optimal. Zeitschrift f \

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.339412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:ada27250b0b87f71b1bc0a8ddefee6543307506f33925f9fc256b1ad453ea582

Observation ecc6379e-502f-42af-b0e0-72bb08616b95 · outbound

This paper cites Approximate modified policy iteration and its application to the game of Tetris.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Approximate modified policy iteration and its application to the game of Tetris

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.334005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:2f2d1b50a8f3d4664f6dba23619e90f1f33e0703205731fb5db2441002b6378f

Observation 8bc49cab-d57c-4a2d-bd5a-2882f791dec0 · outbound

This paper cites Devon Hjelm, Aaron Courville, and Philip Bachman.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Devon Hjelm, Aaron Courville, and Philip Bachman

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.216853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:6b1aef159dbe99aa4a6eda10665db32156b8e40941f83c5713eb15af313aef4f

Observation c40c3e13-b940-479f-b4b8-910b07a1e96c · outbound

This paper cites Mastering the game of Go with deep neural networks and tree search.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Mastering the game of Go with deep neural networks and tree search

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.254965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:1e2324b4421b6bbcddb729ea35bbd90fb19524d35f2764c378756b0925ad31ab

Observation d0a75709-9d72-4536-8268-918a4f0aefef · outbound

This paper cites CURL : Contrastive unsupervised representations for reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration CURL : Contrastive unsupervised representations for reinforcement learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.374282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:e7f0df594cda95ae4dfda4d9d0714775543dcf94b94fc68dcd96fecff6587fb0

Observation 6980bb18-66ce-4320-b1fa-35dfff5a9328 · outbound

This paper cites Decoupling representation learning from reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Decoupling representation learning from reinforcement learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.261040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:777d0af7f1d32aeb0e1733ad83eadc3f30a194a7c3e016d1890fafba1b8cf28f

Observation 7d458331-e9d1-4369-94aa-c5b959e080e5 · outbound

This paper cites Negative dynamic programming.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Negative dynamic programming

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.226634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:072d4cb625f6244480b39f1ff300e4b97b08398ae194bdf368b79ef264beb743

Observation f6f0b7c5-799d-4330-bf46-20ddc45462b6 · outbound

This paper cites Optimistic posterior sampling for reinforcement learning with few samples and tight guarantees.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Optimistic posterior sampling for reinforcement learning with few samples and tight guarantees

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.391856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:3b8bf043dd10591084fea18d15722012ab5b7efe8b4db0d0243dbc76b05bf2c7

Observation 4958b48e-bf46-4584-b4b8-f3ff4cf7bf80 · outbound

This paper cites From Dirichlet to Rubin : Optimistic exploration in RL without bonuses.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration From Dirichlet to Rubin : Optimistic exploration in RL without bonuses

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.353770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:158889e46b2c8ee830c5f81c3d5a463d93cc16c2c0aaf2aacecbd1d7472821d5

Observation caed532b-19a6-421b-8213-94b18ecd769f · outbound

This paper cites Analysis of temporal-difference learning with function approximation.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Analysis of temporal-difference learning with function approximation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.211306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:e4fe0c1c755854b13f46801fda649202c60c388fe894e529ea94df277a556552

Observation 40ff3a06-8940-44aa-b17e-a03171e6164b · outbound

This paper cites Kernelized reinforcement learning with order optimal regret bounds.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Kernelized reinforcement learning with order optimal regret bounds

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.238224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:a9fcaf98cd772611505a30254375df0b5d1f5b563a903a05092c463f2965a39f

Observation 39aa3ebf-9a96-4592-bf18-bf2d18999fbe · outbound

This paper cites Grandmaster level in StarCraft II using multi-agent reinforcement learning.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Grandmaster level in StarCraft II using multi-agent reinforcement learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.241267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:f357a8527b6b9c94e75b7d5dbd48c9ef86c8cb2c4788d15ce7e63b55f237b20c

Observation 637aad1f-b364-4135-b59e-25595319e3e7 · outbound

This paper cites Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Reinforcement learning with general value function approximation: Provably efficient approach via bounded eluder dimension

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.403109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:c3210cf82ec12882e18545ae1d5b48dd568b0615de449bca399bd1c773f6d34e

Observation d696f753-80d6-421b-93f7-ac88e163b2e1 · outbound

This paper cites an unresolved cited work.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-26T01:06:26.248511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:99eb98e47783f6e9071137bb4aa27f27376f319eabd09c08fdc5289c6a3d633e

Observation 12d497a9-3b6f-4ef9-92dd-730aae09ecdd · outbound

This paper cites Offline reinforcement learning with realizability and single-policy concentrability.

A Measure-Theoretic Finite-Sample Theory for Adaptive-Data Fitted Q-Iteration Offline reinforcement learning with realizability and single-policy concentrability

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T01:06:26.214289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=arxiv_source observed=2026-05-09T15:44:02.195783Z digest=sha256:b94a443cc2002e6c0ed7aa94858fc647cc3e882858517bee9f17f34782d20d4e

Pith citing papers

No inbound Pith citation observations are available.