Pith. sign in

Paper Citation Record · LEDGER

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability

As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2510.03494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.03494 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:37:48.385697Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 851be0ef-d974-4bbf-abc2-9b2e0931a8ac · outbound

This paper cites Learning theory from first principles.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Learning theory from first principles

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.505697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.505697Z digest=sha256:123f681bd5d2ecc1e1d410a10dfb0301560ec40bfc0345273b40e493704b32bb

Observation 6930e05b-caa4-4e09-abdd-419b03b77900 · outbound

This paper cites Provably efficient exploration in policy optimization.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Provably efficient exploration in policy optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.544687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.544687Z digest=sha256:274b318248308f4e33e724cec645b15484e3989b6db795f949b8c18c44f234ac

Observation cbb3904c-6608-4d62-a8d1-ea626d1bdb0c · outbound

This paper cites Information-theoretic considerations in batch reinforcement learning.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Information-theoretic considerations in batch reinforcement learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.591936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.591936Z digest=sha256:52a50c7acb45dbd35ae825e7661395592662dc57ee7dd1ecc90ba96777c0aaca

Observation 47090f8a-d141-4495-95f3-40951e0db31b · outbound

This paper cites Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline Reinforcement Learning: Fundamental Barriers for Value Function Approximation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.612684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.612684Z digest=sha256:0ae7427d987eb8757493d72dcd0045a737dd4e526ab7569b206fa126e3c091fe

Observation 36eca9eb-7f0b-4190-86dd-dab045312fec · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Probability inequalities for sums of bounded random variables

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.658083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.658083Z digest=sha256:c465465b12a7417b3b4dd5d3918de6b98e66a3f77787bc6aa47e3b1c05b3ee24

Observation 2075fb25-b8e9-4e80-b20f-abd240864aca · outbound

This paper cites Random design analysis of ridge regression.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Random design analysis of ridge regression

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.702880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.702880Z digest=sha256:240d1d6a8c9ff03805b6082507e267389557fa4840378bb781ac45cb414c6c07

Observation 49916fad-ed64-4902-8810-cd297c9ea285 · outbound

This paper cites Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline Reinforcement Learning: Role of State Aggregation and Trajectory Data

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.745008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.745008Z digest=sha256:49a55bf73cbbf57d7cf70552e99616a556cf61f8f2c1d22fa0a59e7913b807c7

Observation 565e0fc1-2104-4e50-8ede-a1781cd62515 · outbound

This paper cites Offline reinforcement learning in large state spaces: Algorithms and guarantees.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Offline reinforcement learning in large state spaces: Algorithms and guarantees

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.788708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.788708Z digest=sha256:d96a99d4660a96f16446114a625abdcfd8cc4809e6a9f4d44c48bc93c62c7d8d

Observation 2f57979a-9ecb-4f3a-87e4-68a6747ccf42 · outbound

This paper cites Provably efficient reinforcement learning with linear function approximation.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Provably efficient reinforcement learning with linear function approximation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.833062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.833062Z digest=sha256:cd82a5a5b5485ca321e3616a4f75dcc91a55b963538009da35648cf44b8ef451

Observation 18238f8f-0713-410c-9592-c9032ac99a42 · outbound

This paper cites Batch policy learning under constraints.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Batch policy learning under constraints

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.906821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.906821Z digest=sha256:34cddee405ec74738dd8d11a0a4430a9dba93ad88fdee9dd8023f55457abb2cf

Observation 96ece97d-2dd6-4469-9d6f-de5a07d6b671 · outbound

This paper cites Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Sample and Oracle Efficient Reinforcement Learning for MDPs with Linearly-Realizable Value Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:47.941595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:47.941595Z digest=sha256:86a839dee7e96d3ae477003680f8967ea9840997f0dc83018226006e97548d6d

Observation 3db25f7a-4f60-42d8-a3c1-9b52724f9e93 · outbound

This paper cites Finite-time bounds for fitted value iteration.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Finite-time bounds for fitted value iteration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.007752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.007752Z digest=sha256:d22e015c75133c010b8dd61584e4276aa46a6389b4f8d1a9237c7353c983f5d4

Observation fd5f4d3b-e3bf-45de-a3ad-2fc950e22232 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Markov decision processes: discrete stochastic dynamic programming

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.113390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.113390Z digest=sha256:a1ac21de4b3ff1bde6d07f1367176e5c85079dfbc4cacf12f8eee415a9c01b65

Observation 2cd1b858-0f0a-40de-ab6c-d0b1bc2bb3eb · outbound

This paper cites Trajectory data suffices for statistically efficient learning in offline rl with linear qpi-realizability and concentrability.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Trajectory data suffices for statistically efficient learning in offline rl with linear qpi-realizability and concentrability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.192979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.192979Z digest=sha256:29c4674ec76c22abeb487ea6c0175c5cf479b149e563f8009bbe7b3ffa675148

Observation 492e7b46-22a4-4905-a156-40eac738b4c2 · outbound

This paper cites Minimum-volume ellipsoids: Theory and algorithms.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Minimum-volume ellipsoids: Theory and algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.256135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.256135Z digest=sha256:90550d2ca8f3f3a1b69857fa51cf98947fd4fc8fde99550fe756674bc391b340

Observation c2c93dcc-69ed-427b-9172-539d2db59dca · outbound

This paper cites High-dimensional probability: An introduction with applications in data science, volume 47.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability High-dimensional probability: An introduction with applications in data science, volume 47

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.289220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.289220Z digest=sha256:45baddcf5745bc57f8c62e4687ea2e0edaba5ddc532493cfd91d52674b7f284f

Observation ad74ca0a-a1c0-4f24-b752-a3918b30bff1 · outbound

This paper cites Online RL in Linearly $q^\pi$-Realizable MDPs Is as Easy as in Linear MDPs If You Learn What to Ignore.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Online RL in Linearly $q^\pi$-Realizable MDPs Is as Easy as in Linear MDPs If You Learn What to Ignore

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.336588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.336588Z digest=sha256:3f2d34769569d4a8c10a7b8a843024bbd13c81ea6a472a8a338b800071a1272c

Observation 5c4022cb-3c3a-4401-a892-d7974d5cb265 · outbound

This paper cites Learning near optimal policies with low inherent bellman error.

Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^\pi$-Realizability and Concentrability Learning near optimal policies with low inherent bellman error

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:37:48.385697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:37:48.385697Z digest=sha256:59cef77228eda7314f7aba966bc15da292d91569c69a71e8705c49be7fdddecb

Pith citing papers

No inbound Pith citation observations are available.