Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:37:48.385697Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2510.03494.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:37:48.385697Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 851be0ef-d974-4bbf-abc2-9b2e0931a8ac · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Learning theory from first principles
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6930e05b-caa4-4e09-abdd-419b03b77900 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Provably efficient exploration in policy optimization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbb3904c-6608-4d62-a8d1-ea626d1bdb0c · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Information-theoretic considerations in batch reinforcement learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47090f8a-d141-4495-95f3-40951e0db31b · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36eca9eb-7f0b-4190-86dd-dab045312fec · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Probability inequalities for sums of bounded random variables
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2075fb25-b8e9-4e80-b20f-abd240864aca · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Random design analysis of ridge regression
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49916fad-ed64-4902-8810-cd297c9ea285 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 565e0fc1-2104-4e50-8ede-a1781cd62515 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline reinforcement learning in large state spaces: Algorithms and guarantees
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f57979a-9ecb-4f3a-87e4-68a6747ccf42 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Provably efficient reinforcement learning with linear function approximation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18238f8f-0713-410c-9592-c9032ac99a42 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Batch policy learning under constraints
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ece97d-2dd6-4469-9d6f-de5a07d6b671 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db25f7a-4f60-42d8-a3c1-9b52724f9e93 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Finite-time bounds for fitted value iteration
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5f4d3b-e3bf-45de-a3ad-2fc950e22232 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Markov decision processes: discrete stochastic dynamic programming
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd1b858-0f0a-40de-ab6c-d0b1bc2bb3eb · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Trajectory data suffices for statistically efficient learning in offline rl with linear qpi-realizability and concentrability
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 492e7b46-22a4-4905-a156-40eac738b4c2 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Minimum-volume ellipsoids: Theory and algorithms
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c93dcc-69ed-427b-9172-539d2db59dca · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability High-dimensional probability: An introduction with applications in data science, volume 47
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad74ca0a-a1c0-4f24-b752-a3918b30bff1 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Online RL in Linearly $q^\pi$-Realizable MDPs Is as Easy as in Linear MDPs If You Learn What to Ignore
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c4022cb-3c3a-4401-a892-d7974d5cb265 · outbound
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Learning near optimal policies with low inherent bellman error
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.