Pith. sign in

Paper Citation Record · LEDGER

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures

As of 11 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 0 inbound Pith citation observations for arXiv:2501.02089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02089 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:16.651877Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 109 outbound references displayed

  • verified exact0
  • verified fuzzy61
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cd67070-287d-4075-be8a-2fa3de722aa9 · outbound

This paper cites Im- proved algorithms for linear stochastic bandits.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Im- proved algorithms for linear stochastic bandits

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.208948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.208948Z digest=sha256:1d40b6790570a82507ddb3cb7d8c885f48a6dcc7b721650eec4832c9738d01b6

Observation 1529a580-e3fc-49a1-886e-e55ef4ae0448 · outbound

This paper cites Model-based reinforcement learning with a generative model is minimax op- timal.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Model-based reinforcement learning with a generative model is minimax op- timal

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.214445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.214445Z digest=sha256:2340d09c2b4bc1279d54f7d1aa885b7dbc2e5a983d0d98b427124effcf941ade

Observation d79c0843-8f33-40dc-b94f-c4a8d025be67 · outbound

This paper cites Degenerate nonlinear programming with a quadratic growth condition.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Degenerate nonlinear programming with a quadratic growth condition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.219509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.219509Z digest=sha256:740c5d640e229c1d7f15032acebc548c26a12562dbb9594e7e5f33377c663456

Observation c4dce0c3-ce49-4107-bfd8-61e49fb6134b · outbound

This paper cites Fitted q- iteration in continuous action-space mdps.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Fitted q- iteration in continuous action-space mdps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.225342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.225342Z digest=sha256:c12f41c4033fa5e0ccf253af0df395ccb40562dd8d594f009db4c32c82f3b455

Observation 1b85e5d2-4428-4f56-bda2-9ae3cc6a4408 · outbound

This paper cites Learning the target network in function space.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Learning the target network in function space

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.230182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.230182Z digest=sha256:9eaa5520230788378ac9278824483e62c22cc626c760b3b9ae03992828c09a3b

Observation 95f88b01-9202-49e2-ad3c-3b1126b78f35 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem, 2002.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite-time analysis of the multiarmed bandit problem, 2002

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.234126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.234126Z digest=sha256:ac7da8aa321f97dd311fab5eb6b89d2773e40dc51ba0055c36d5b2d120e5ced9

Observation 4bfac92d-a9d1-4350-9b10-538a891de260 · outbound

This paper cites Minimax regret bounds for reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax regret bounds for reinforcement learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.238833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.238833Z digest=sha256:e36f951b9a6d4d0462981a321097eb0e94cebee53b34caa857c519d358f63da6

Observation 612dbc69-42f1-4026-aac5-f9cb2c4542aa · outbound

This paper cites Prov- ably efficient q-learning with low switching cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Prov- ably efficient q-learning with low switching cost

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.242675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.242675Z digest=sha256:a521a491f364ba87fceaea2de7c5fa620c3f6720afda166b4b70635c7134a4f4

Observation 391d16ad-2c33-4164-bce9-9a283ade1ad1 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.247682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.247682Z digest=sha256:39b5f9208dc2baa07f25da119a6367ba036e3478fb8f24259ee325d14c5d9c22

Observation 6fed46ad-9f26-4721-a0fd-d6a4296b182e · outbound

This paper cites Dynamic programming.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dynamic programming

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.252535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.252535Z digest=sha256:420702811212150c2fd9918987acd2215f1b0845e3b9e18b8574264b3fe4fdb0

Observation 2d5bcf53-36f9-43ae-95c6-59bda6936c10 · outbound

This paper cites Online learning with switching costs and other adaptive adversaries.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Online learning with switching costs and other adaptive adversaries

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.257046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.257046Z digest=sha256:5521260bf9072e53492c568ba4786c4899df4e76a315701a4e21a0359fda246d

Observation 579c1e47-5e27-49f6-b5c0-15309c726438 · outbound

This paper cites Information-theoretic considera- tions in batch reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Information-theoretic considera- tions in batch reinforcement learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.261458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.261458Z digest=sha256:80fd67fdf58ba263b2a59d76c449840bb61f48a7a0c5fcc3f658dda85d18fdf8

Observation 6e30b8fe-222f-4a3d-8460-5f0773f20b1e · outbound

This paper cites Deep reinforcement learning from human preferences.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Deep reinforcement learning from human preferences

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.265753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.265753Z digest=sha256:eb6c513b55599ad71ed41a565e0ffccd7397879b76c06726b9edd2763d8fb53c

Observation b77f5264-9dea-4525-966d-53eb0a2c6bb8 · outbound

This paper cites Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Pessimistic Nonlinear Least-Squares Value Iteration for Offline Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.269952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.269952Z digest=sha256:bf67cdc23cefad857a4bcd8d8fc216d5f1a8e72743576e78ab0d4b013dc4f2b6

Observation 30728537-2956-4a7d-9cae-85becbeb2052 · outbound

This paper cites Minimax-optimal off- policy evaluation with linear function approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax-optimal off- policy evaluation with linear function approximation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.274588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.274588Z digest=sha256:d119b32653b4150fa1bf53e41213c4e74a9d8ddc5871d77c0ca9b6b6dd34df92

Observation 2f664f1a-61d3-4d8b-b260-dafc2f2b3c11 · outbound

This paper cites Doubly Robust Policy Evaluation and Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Doubly Robust Policy Evaluation and Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.279147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.279147Z digest=sha256:7252f547812cbef47c1fb0f6ff8b3e4c4234e8e93a6015a7da9d3843d166130a

Observation 9ff4366c-7ce2-4045-b441-874bd0c7ad4b · outbound

This paper cites Tree-based batch mode reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Tree-based batch mode reinforcement learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.284580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.284580Z digest=sha256:68ea20be43166f0f0e7bc059084907bda909e148110466f9bb10feef8765841d

Observation 3e421dd0-e1f5-4858-aff6-0396b456a2f3 · outbound

This paper cites Discovering faster matrix multipli- cation algorithms with reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Discovering faster matrix multipli- cation algorithms with reinforcement learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.289795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.289795Z digest=sha256:b3c8edf30b1024e7203b5522ae65a9ef3432863a09cffa2937d415ad5bdc5c8b

Observation 1547a81a-a1db-4e9e-8690-9deba09ac06b · outbound

This paper cites Theory of statistical estimation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Theory of statistical estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.294226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.294226Z digest=sha256:8f77e3c029e5c773f92055fb4b05e04576f51a3372c6f31b62889b1f58d6b4f2

Observation b23adf15-01b8-408c-ba93-c0a3705e9685 · outbound

This paper cites A Provably Efficient Algorithm for Linear Markov Decision Process with Low Switching Cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures A Provably Efficient Algorithm for Linear Markov Decision Process with Low Switching Cost

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.298033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.298033Z digest=sha256:ade86a427d563f3fd1c963c6a8dd0d9ba53a599706692adf9b9473e90f9ce023

Observation 6db73a91-af48-4e56-80e4-676867e272e1 · outbound

This paper cites Batched multi-armed bandits problem.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batched multi-armed bandits problem

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.302003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.302003Z digest=sha256:fa84eb28bcec2606f1f425200792f4c46949d7253c5f8f1e54e059f973a56226

Observation cac3d414-6306-424a-b5bb-2427cbdfb052 · outbound

This paper cites Off-policy deep rein- forcement learning by bootstrapping the covariate shift.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Off-policy deep rein- forcement learning by bootstrapping the covariate shift

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.305987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.305987Z digest=sha256:74e82cf2086e615d54ad1f5f84f3fb685466ceb7009180c7983b10abbb8cd3bb

Observation 961f3efe-270f-40b7-9217-70dd5c1cf63b · outbound

This paper cites Minimax pac bounds on the sample complexity of rein- forcement learning with a generative model.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax pac bounds on the sample complexity of rein- forcement learning with a generative model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.311486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.311486Z digest=sha256:55f0fbb4a84e6dba0b01b5091a46043f706b2e8bea4426ea039564ccf5c5560d

Observation 689b4757-97cd-4604-b07c-f3f151f94402 · outbound

This paper cites Maxmin expected utility with non-unique prior.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Maxmin expected utility with non-unique prior

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.317103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.317103Z digest=sha256:5268a5604ad42b16e9d6d6631b60a32e3b44ff6cf5300c23a2db42a5713663f8

Observation 1b6a0a7d-db16-4d7a-98c6-b9ab21630074 · outbound

This paper cites Approximate solutions to Markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Approximate solutions to Markov decision processes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.321994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.321994Z digest=sha256:bfffbcabf6d6f685183cb2d5f3278fc074e35487fc7671128946b698f4b04f04

Observation 2fe80bfc-4112-4f5f-b37b-55194084191d · outbound

This paper cites Networkgym: Reinforcement learn- ing environments for multi-access traffic management in net- work simulation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Networkgym: Reinforcement learn- ing environments for multi-access traffic management in net- work simulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.326694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.326694Z digest=sha256:c9bb11925b460511ff3db3d8c9f816da9ce1745a29e2e65d57e85829d508088d

Observation 60ca303f-2507-4a8c-a827-34d6d588f2ca · outbound

This paper cites Consistent on-line off-policy evaluation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Consistent on-line off-policy evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.331632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.331632Z digest=sha256:e3f716372338c12b88c6b85d23174a190d12b896f54f62a92bd84774530af882

Observation 05eb0f54-0dd3-451b-b852-3dffe9027cc4 · outbound

This paper cites Bootstrapping fitted q-evaluation for off- policy inference.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bootstrapping fitted q-evaluation for off- policy inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.336214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.336214Z digest=sha256:783b6725d11f5c5d162d60144f02ee343075c919a6ca328d63a061e8d4728548

Observation 667d8a66-1ef5-4361-8463-37b978db587a · outbound

This paper cites Effi- cient estimation of average treatment effects using the estimated propensity score.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Effi- cient estimation of average treatment effects using the estimated propensity score

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.340340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.340340Z digest=sha256:ba8639e59fbbac7f95943deaf365584c96cd1b8a77a209d7fc8a82d524a480b3

Observation d26e45f8-545a-4681-901f-2b34b8f84dab · outbound

This paper cites A generalization of sampling without replacement from a finite universe.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures A generalization of sampling without replacement from a finite universe

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.344523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.344523Z digest=sha256:d91ae6d38f8f21ea1720de759814ae4267c87d2da61176144892b60ab8cad0c2

Observation f0b9d2f9-6be9-4d6c-b1c6-7370f8df9449 · outbound

This paper cites Towards deployment-efficient reinforcement learning: Lower bound and optimality.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards deployment-efficient reinforcement learning: Lower bound and optimality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.348936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.348936Z digest=sha256:9b314e75548c2c1e25803fabfebb2dbd2232bccdb1dfd1381bf8ec772ab04fbd

Observation 043356c3-92bf-4056-b959-a07c8de46a47 · outbound

This paper cites Aleatoric and epis- temic uncertainty in machine learning: An introduction to con- cepts and methods.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Aleatoric and epis- temic uncertainty in machine learning: An introduction to con- cepts and methods

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.353131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.353131Z digest=sha256:35834ec6ec056c9df7d338fb7af8852a04c1797488b892e5127b8600a0aba83d

Observation 4c4b3a65-6f0c-404d-aa1f-1ff225bc51bd · outbound

This paper cites Doubly robust off-policy value eval- uation for reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Doubly robust off-policy value eval- uation for reinforcement learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.788348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.357051Z digest=sha256:2762da9f3647a22670602410fd4b563eab3d1acb573dff9941d72a6985fb57ca

Observation e6085e59-902e-41c4-91d9-e09f3fc8ff07 · outbound

This paper cites Offline reinforcement learning in large state spaces: Algorithms and guarantees.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning in large state spaces: Algorithms and guarantees

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.775309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.360774Z digest=sha256:5c946716dd7966917069f44e3da378c1473c1416decb5099cf68a085c1e16781

Observation f52c8254-9bed-41b8-9dff-eb535d10df04 · outbound

This paper cites Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Is pessimism provably efficient for offline rl? In International Conference on Machine Learning, pages 5084–5096

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.761998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.365059Z digest=sha256:3fc0eb8c5e8deac6f97f5624a9262444f3cd5932426d3adbfd29f45e93fcbec7

Observation b8a3ab84-a748-4bcd-b827-85ff45cac308 · outbound

This paper cites Double reinforcement learning for efficient off-policy evaluation in markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Double reinforcement learning for efficient off-policy evaluation in markov decision processes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.748834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.369711Z digest=sha256:5d43daf8563925a9380c52e73c844afe58618a8030ceb2f48cd50b71aa3ad708

Observation 5a83a642-146b-4251-9a99-64b8b41d7be2 · outbound

This paper cites Efficiently breaking the curse of horizon in off-policy evaluation with double reinforce- ment learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Efficiently breaking the curse of horizon in off-policy evaluation with double reinforce- ment learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.735715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.373992Z digest=sha256:8cc591c845dfbe11a7f9dc19c87a4288e9b180808c52c0d53573b5527ec3cd57

Observation bafe5139-e632-4f6d-8cb9-5edc5238010c · outbound

This paper cites Near-optimal reinforce- ment learning in polynomial time.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal reinforce- ment learning in polynomial time

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.722332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.378249Z digest=sha256:30bc93dd3b4cb3547479b66f901a52b79f9a605277ea33a9009bf62d0e7a3bd3

Observation 6226483e-26dc-4d44-9f62-557c512b6647 · outbound

This paper cites Introduction to empirical processes and semiparametric inference, volume 61.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Introduction to empirical processes and semiparametric inference, volume 61

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.710137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.383340Z digest=sha256:c752b6edbcbfa3c563fa4178734346730eaecaa5679c0b93048496f32a5fc509

Observation 8641d3d2-57e5-4f99-8e8a-5d4477510b90 · outbound

This paper cites Offline reinforcement learning with fisher divergence critic regularization.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning with fisher divergence critic regularization

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.698074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.387687Z digest=sha256:3b93e43a4f246ce9f00bb9890a1b19c8f82e381d19df0bd87d6c144a672cbc25

Observation 0663e18c-acf2-4cba-9ed9-b16a229433a0 · outbound

This paper cites Conservative q-learning for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Conservative q-learning for offline reinforcement learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.685186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.391784Z digest=sha256:0873a6e518a94305cfe70ae7563c54ad02925945c96acf14adaa3f9024c3d14e

Observation 3aa17862-1c9b-4384-92c8-b34eecdbfe0a · outbound

This paper cites Minimax theory.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax theory

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.671509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.395913Z digest=sha256:1d959021aa9f8106b1636efa320fb22f85004d04fc7dc531a956a45b277b6849

Observation 9a2ba5fa-fec8-4916-8dc0-a6bd09c3fd44 · outbound

This paper cites Batch policy learning under constraints.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batch policy learning under constraints

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.658549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.400745Z digest=sha256:962e9972a0b78b7b8c8f7073c8365fef1b193b9f866eca47ebf06d4c47e12019

Observation 53fbf41f-b39e-4273-baa9-63ed9b715275 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.405537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.405537Z digest=sha256:96c4795f12f740e504b900dfe3e3b80dd0c3e52920353c5cdf5807877670401a

Observation 215ce0da-c7fa-4b64-8547-77caf4e57389 · outbound

This paper cites Breaking the sample size barrier in model-based reinforcement learning with a generative model.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Breaking the sample size barrier in model-based reinforcement learning with a generative model

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.645299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.410690Z digest=sha256:1d5dab53aadcef4164dbfafeb2102975d26a5b49549b5e9ec8339f85df14e3dd

Observation 49255547-06f2-4497-9765-454caa65547e · outbound

This paper cites Offline reinforcement learning with closed-form policy improvement operators.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Offline reinforcement learning with closed-form policy improvement operators

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.632786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.415116Z digest=sha256:c01b9a88e3747850c78dea9d73e8ef48bdbee598cda5837e6d5c96b76a6ff078

Observation f09321f7-cc91-43c8-804d-e6d4873c4936 · outbound

This paper cites Unbi- ased offline evaluation of contextual-bandit-based news article recommendation algorithms.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Unbi- ased offline evaluation of contextual-bandit-based news article recommendation algorithms

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.620627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.419615Z digest=sha256:f2690ad894b1f4b58e9249f3b827b2b25e52bb89e4706943d1948a2123f4675f

Observation 9a16b704-7ff1-4654-80e9-82d0632306c6 · outbound

This paper cites Monte Carlo strategies in scientific computing, volume 10.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Monte Carlo strategies in scientific computing, volume 10

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.608305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.423566Z digest=sha256:97a6683fa6c957194ceff7113e7fbfe60d670f9c58490bb272a45cc1cebf34d8

Observation 78954958-2c3a-4b48-8a42-23141e5ff526 · outbound

This paper cites Breaking the curse of horizon: Infinite-horizon off-policy es- timation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Breaking the curse of horizon: Infinite-horizon off-policy es- timation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.595029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.427799Z digest=sha256:221c3eca9fa550591545c6898836e33a88257a55f6d519a850d64ee2b0af112b

Observation 57a75e49-66aa-434d-9959-297a5db9a7af · outbound

This paper cites Off-policy policy gradient with stationary distribution cor- rection.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Off-policy policy gradient with stationary distribution cor- rection

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.581997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.432172Z digest=sha256:69baf571a6e83210d7cdcda2de8f3278bbe5ed46a8c6af0897b14634e0deee9f

Observation ae97e546-50b3-4c9b-bf78-01671a0aa7ee · outbound

This paper cites Provably good batch off-policy reinforcement learning without great exploration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Provably good batch off-policy reinforcement learning without great exploration

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.567624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.436590Z digest=sha256:18b79172ac39cef613a028ce478299cf12c01ce33d9c4d8d21d9fafac7fd3df2

Observation 405675f4-ba25-4549-8375-532ad9945af0 · outbound

This paper cites Mildly conservative q-learning for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mildly conservative q-learning for offline reinforcement learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.554612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.441947Z digest=sha256:99da91715b9f0994764d99c0a6549183cd16b6ad59995655f5f06dbda0358839

Observation 9e7b0ac0-cd9f-41d8-acd5-3af09958e87e · outbound

This paper cites the distribution-norm to the res- cue.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures the distribution-norm to the res- cue

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.541828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.446115Z digest=sha256:41eb337295315d90bd4ac05dffa8c1604135ca2e81a14acdb8c515e36363213a

Observation 54ed687c-d2c2-46b3-a236-996f74795263 · outbound

This paper cites Faster sorting algorithms discovered using deep reinforcement learn- ing.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Faster sorting algorithms discovered using deep reinforcement learn- ing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.529395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.450397Z digest=sha256:f0a1f631933a03ce4a997658cc75ea66b91a51284eebdafa003c8498b4f20877

Observation 1e0d4ece-f616-4e4f-81d8-bdb7858b2412 · outbound

This paper cites Neural adaptive video streaming with pensieve.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Neural adaptive video streaming with pensieve

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.516072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.455122Z digest=sha256:ea3d262d65f95400c95afe5cb33b253446b6f3334247ba1a533b9765abe6a0ca

Observation 8b2e305d-de76-487c-835f-3755f4a0be3a · outbound

This paper cites Deployment-efficient reinforce- ment learning via model-based offline optimization.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Deployment-efficient reinforce- ment learning via model-based offline optimization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.502046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.459496Z digest=sha256:38452deb41e1ef50304aa281be2a15214ba4e8be15b88b4b70b16c4a7cda70e4

Observation cb40f6e4-1f27-4e42-bff5-f758f1b319b9 · outbound

This paper cites Dependent central limit theorems and in- variance principles.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dependent central limit theorems and in- variance principles

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.487751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.464018Z digest=sha256:dcf568dd69f952de9bf1a1e88e4b874c8133912a7ed0c9bf0be853ca6f30ffa6

Observation 58ccbf41-73d4-412d-89f5-14b877a01c8b · outbound

This paper cites Variance-aware off-policy evaluation with linear function ap- proximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Variance-aware off-policy evaluation with linear function ap- proximation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.475680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.468042Z digest=sha256:c25cdce04d912ffc3f289bed989f00857c95a2c2446838413ed61dc2a6820490

Observation e34d18aa-ef27-4d2a-9fb0-74054f8e5149 · outbound

This paper cites Human-level control through deep reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Human-level control through deep reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.462963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.471469Z digest=sha256:bec59f9f81019fb1d6e84f9adaceb05e26784bc276aa1a67613c6daa90be70c1

Observation dd6bb0c5-0cfa-427b-b9ca-0fed1a73afff · outbound

This paper cites Bootstrapping: A nonparametric approach to statistical infer- ence.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bootstrapping: A nonparametric approach to statistical infer- ence

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.450599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.475970Z digest=sha256:4eba8d56f3aba4d309d985a5b2772f7daa1134b51825f4a283a6831afd29fb6f

Observation e59f5fa6-37ca-4640-8f65-b706af237290 · outbound

This paper cites Finite-time bounds for fit- ted value iteration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite-time bounds for fit- ted value iteration

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.437575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.479912Z digest=sha256:98c6d56674c53693682f0065880f874978749b52d63a31a2529fde3668f025ca

Observation 193c6ec1-b31e-4457-95f6-cac07abce26a · outbound

This paper cites Marginal mean models for dynamic regimes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Marginal mean models for dynamic regimes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.422478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.483698Z digest=sha256:12e599a1fc6f1503ddf9760c730286adf58aa21d307c4980eb982c6ac94664d9

Observation 3ccf2f67-d769-44e9-9bce-ee8a2068faaf · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribu- tion corrections.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Dualdice: Behavior-agnostic estimation of discounted stationary distribu- tion corrections

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.409007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.487956Z digest=sha256:6822be4336a23c9c099cdaef29a556c751ffb10238201ff59bba65ad45241fb3

Observation 95de1a03-1199-4c28-ad5b-2985258fee37 · outbound

This paper cites AlgaeDICE: Policy Gradient from Arbitrary Experience.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures AlgaeDICE: Policy Gradient from Arbitrary Experience

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.491915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.491915Z digest=sha256:dbd2101252adf55b17b7578302fe53f40b179d68f771a137adf703b4248e241f

Observation b0d4c04b-e228-4175-ae87-a8383b436721 · outbound

This paper cites Optimal medication dosing from suboptimal clinical ex- amples: A deep reinforcement learning approach.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Optimal medication dosing from suboptimal clinical ex- amples: A deep reinforcement learning approach

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.394785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.497573Z digest=sha256:612c4716ea640656ff12cef4236307ebbcfec2749fcea55b0651194ce6ba6880

Observation c3585e0a-0abb-469f-9e0b-71f14dfd013e · outbound

This paper cites On sample-efficient of- fline reinforcement learning: Data diversity, posterior sampling and beyond.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On sample-efficient of- fline reinforcement learning: Data diversity, posterior sampling and beyond

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.381978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.502392Z digest=sha256:2696d1b83951f3ca2abd3999fc525135a5862806e45134c6a48ce94aca986394

Observation 19b323a4-cbe9-4848-895f-91a27aa5631a · outbound

This paper cites On instance-dependent bounds for offline reinforcement learning with linear function approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On instance-dependent bounds for offline reinforcement learning with linear function approximation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.368016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.507044Z digest=sha256:0c0c095311e973c86c396ba043b8ae8e46862cf77c5122c94df1a2693dab6d58

Observation 2d4b2c17-a3f3-47bd-bd3d-27744fb6f1b5 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.354576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.511379Z digest=sha256:5f192506bf4ba00752ed569147f4093e1621b0c9c03290af7fd12a4248784834

Observation 2e04c9c0-c068-49f1-81b0-e36439b6e035 · outbound

This paper cites Batched bandit problems.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Batched bandit problems

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.341189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.516827Z digest=sha256:b516fb8a29018a4d500d9b4d03485b8e6aeb9ad436b9b0903f71a0dd04641d5c

Observation 4c44f74f-5eb9-4085-b8d9-9d5dd9f2e6d4 · outbound

This paper cites Approximate Dynamic Programming: Solv- ing the curses of dimensionality , volume 703.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Approximate Dynamic Programming: Solv- ing the curses of dimensionality , volume 703

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.326945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.521801Z digest=sha256:021d9e8d9a49439d3bb43438d092ea7a7b212778105d9d02361a19dc6ac569ec

Observation e5b586e5-8f70-4316-b061-bf6d2bf8990a · outbound

This paper cites Eligibility traces for off-policy policy evalua- tion.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Eligibility traces for off-policy policy evalua- tion

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.311762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.525885Z digest=sha256:201436c2e903865be440a15b73350d49f5e8e759b25a45d8a428d3eb04adcc2a

Observation 45d70d9d-3624-4ccc-8f27-5dd679f15df5 · outbound

This paper cites Markov decision processes.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Markov decision processes

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.298253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.529327Z digest=sha256:3ed4ccbc9d92ec7348bd2c2140937c2dc781d385f7a0ad9939e3f069c3d5e28a

Observation 563ed0ae-d2b1-41ed-91ff-088985f8c4ae · outbound

This paper cites Near-optimal deployment effi- ciency in reward-free reinforcement learning with linear func- tion approximation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal deployment effi- ciency in reward-free reinforcement learning with linear func- tion approximation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.285372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.533749Z digest=sha256:7326a4fd6c3b29e43142926c61a5f98728d22200c298421231f21f0c44248748

Observation 5cc730c1-b9ae-4717-bff6-d5aed1d48b13 · outbound

This paper cites Sample- efficient reinforcement learning with loglog (t) switching cost.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Sample- efficient reinforcement learning with loglog (t) switching cost

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.270273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.539089Z digest=sha256:540bae17c2083042e1d68ff4c7208d9528fdc688eb8286c864d8d2d202623f08

Observation 90471ee9-5843-4fbd-be07-5c0ef964fc56 · outbound

This paper cites Logarithmic switch- ing cost in reinforcement learning beyond linear mdps.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Logarithmic switch- ing cost in reinforcement learning beyond linear mdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.255721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.543674Z digest=sha256:64134d7bdf48a50c0d43308d73d832604ae4bcf74dccf7f962108f382d9f15e5

Observation d74fba3d-d6a5-40ca-bb2e-ac5dc47a8672 · outbound

This paper cites Bridging offline reinforcement learning and im- itation learning: A tale of pessimism.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bridging offline reinforcement learning and im- itation learning: A tale of pessimism

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.242121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.548582Z digest=sha256:4ad93039bfdb7b0faba099875dba2408c35e696fd0515337800f2a2d85903350

Observation 08b9b6dd-f4fb-4a25-bcee-9eba6b35c3e6 · outbound

This paper cites Nearly horizon-free offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Nearly horizon-free offline reinforcement learning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.228950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.553237Z digest=sha256:be13bfb1237146e40bd4175a3963b3942871a2555a6e2fb5b1843a440cd1ae67

Observation 73f5d59e-d8e3-46ff-8735-08a530c69867 · outbound

This paper cites Mastering the game of go with deep neural networks and tree search.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mastering the game of go with deep neural networks and tree search

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.215211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.558147Z digest=sha256:378a47284c06028d51f9fe0a9c98dd9870780c8c622b215181032bb38d3fb5ce

Observation 60a06f68-e025-467a-9587-758852db3281 · outbound

This paper cites Mastering the game of go without human knowledge.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Mastering the game of go without human knowledge

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.563181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.563181Z digest=sha256:45f71079342ef618d60d119467ca173c679d92ddb7d14fa05c4a73ef5418f203

Observation aa648aef-6fb2-4b3c-9e69-3af10e8a56ca · outbound

This paper cites Learning to summarize with human feed- back.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Learning to summarize with human feed- back

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.192685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.567487Z digest=sha256:87af9800fae7b640b2dd3d9a6ae1c89b5563b738c4b065f21a045cc872ed5668

Observation d0e89996-56c5-4de9-b783-0482330b2bde · outbound

This paper cites Reinforcement learn- ing: An introduction.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Reinforcement learn- ing: An introduction

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.179346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.571125Z digest=sha256:89559ae53faa6f0aa4e162204b154cee94c836b8164db0393ec40cbc4d0a08c5

Observation 5c4d9419-5927-4a22-adc4-ac8db1fc9fd2 · outbound

This paper cites Finite time bounds for sampling based fitted value iteration.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Finite time bounds for sampling based fitted value iteration

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.166147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.574853Z digest=sha256:7ab96bbba65b82eb72d2310b75b1b9a5f278a935955cc26fb4a826aba4406e22

Observation 1f620bd3-35e8-4b44-a8e0-c92594300be9 · outbound

This paper cites Semiparametric theory and missing data, volume 4.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Semiparametric theory and missing data, volume 4

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.578347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.578347Z digest=sha256:10283703d53ffa9dcd130ddac200e49dc2509867c81b1063d594916f14864e57

Observation 5253eba6-d85a-49eb-a3e9-63f1629d1595 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Minimax weight and q-function learning for off-policy evaluation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.143639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.584284Z digest=sha256:7d6503416efbc97f16400b90a3198a7abc2da813a646481c45f7840133848a41

Observation fdef9586-60e1-47ab-aaa9-a545b885a559 · outbound

This paper cites Asymptotic statistics, volume 3.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Asymptotic statistics, volume 3

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.129716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.588943Z digest=sha256:e9a90ce8f0933ee099e10f0ee9f4709f8e5a65c648f5bbd45edb15401ca3465a

Observation f87c33c9-fd40-4ca5-8ba6-c586cc08db69 · outbound

This paper cites High-dimensional statistics: A non- asymptotic viewpoint, volume 48.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures High-dimensional statistics: A non- asymptotic viewpoint, volume 48

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.116170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.593223Z digest=sha256:84cc82c55cb4d31366a3a5f2fb614b0cdeaa5ff197c10e908bbf1e0e8e518cfc

Observation 7b684ee2-762a-4cab-bb9b-37310552c199 · outbound

This paper cites Provably efficient reinforcement learning with linear function approxi- mation under adaptivity constraints.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Provably efficient reinforcement learning with linear function approxi- mation under adaptivity constraints

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.102768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.598331Z digest=sha256:cd92279186bfb2fd6f56074026abdae891f608631067a77b32248ecd4bf2d598

Observation d95b0957-9ce6-4423-818d-b7197348b8e4 · outbound

This paper cites On gap-dependent bounds for offline reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On gap-dependent bounds for offline reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.088796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.602873Z digest=sha256:e4d6bca53dbcebe82229df46261c87b8ae098d5dfbe2c2f8c3b53352fe272337

Observation 0646dc31-11d4-48ce-b901-1af3c63df95b · outbound

This paper cites Opti- mal and adaptive off-policy evaluation in contextual bandits.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Opti- mal and adaptive off-policy evaluation in contextual bandits

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.075476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.607201Z digest=sha256:64c3db6bb1ce9911b50e7f98f9b8b661312d712703caf380efe5f16473d5b29f

Observation 17268b18-fa42-4ba1-9a94-47bdce98e4c1 · outbound

This paper cites Behavior Regularized Offline Reinforcement Learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Behavior Regularized Offline Reinforcement Learning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.611265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.611265Z digest=sha256:31b5aa66aa1cd78d955282c65c8f32f57e6e69d9fa7bc47ed3d3a427d9878050

Observation 4c6f5143-4834-4b7f-b907-f772f2121fc5 · outbound

This paper cites On the optimality of batch policy optimization algorithms.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures On the optimality of batch policy optimization algorithms

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.061734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.615198Z digest=sha256:219a04c594d806480a4487cd26450af518ce8cd09b91766c0db3da9d49f35a6c

Observation 9757f4e0-0dbe-4062-bba3-b3cfcc744b21 · outbound

This paper cites Q* approximation schemes for batch reinforcement learning: A theoretical comparison.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Q* approximation schemes for batch reinforcement learning: A theoretical comparison

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.047894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.619015Z digest=sha256:cee3b047c11bdd224378067dcfd33bc107858c8de7e6a0b9094ebe606029399f

Observation 189aff4c-fe89-4510-a5a7-af8d297a211c · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginal- ized importance sampling.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards optimal off-policy evaluation for reinforcement learning with marginal- ized importance sampling

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.034006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.622611Z digest=sha256:1bcaf53ba8603481bbc1d44897a114a9aee6cb6fd7e2789a4c0c96814cfcbd24

Observation da2895fe-2f71-49aa-bfa3-6016f79f0d28 · outbound

This paper cites Bellman-consistent pessimism for offline rein- forcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Bellman-consistent pessimism for offline rein- forcement learning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:17.020810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.626325Z digest=sha256:8f8201460843567130ec1c3e4c750ddfa700605d21b2a6c047e847b6e1df535b

Observation 19817067-7e7a-4a91-9f41-05d340030dd3 · outbound

This paper cites Policy finetuning: Bridging sample-efficient offline and online reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Policy finetuning: Bridging sample-efficient offline and online reinforcement learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.630025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.630025Z digest=sha256:c431b8f9a36d8cb3c5dc9466d5e8d697f48d7ea03c96220968acd2902fc5ff13

Observation 6064c2ed-6b28-489a-8a2b-74a6bfba6e3b · outbound

This paper cites Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Nearly Minimax Optimal Offline Reinforcement Learning with Linear Function Approximation: Single-Agent MDP and Markov Game

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:16.633815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:16.633815Z digest=sha256:c2494b1305c7e31947370046dab04c2ba24711f52132d35391691e48b3dfb46b

Observation 6dfea6f6-3eb1-4951-bbdd-83da7727c7d5 · outbound

This paper cites Asymptotically efficient off- policy evaluation for tabular reinforcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Asymptotically efficient off- policy evaluation for tabular reinforcement learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.998146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.637934Z digest=sha256:ed57fa5bac387db9d5f4f894bd6c1c3f9955ef41f96da8f4be1d8ea42f5d2273

Observation 80472f83-d677-47e9-aeff-1e7a3590bdcf · outbound

This paper cites Towards instance-optimal of- fline reinforcement learning with pessimism.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Towards instance-optimal of- fline reinforcement learning with pessimism

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.984804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.643262Z digest=sha256:73611ed8340864ee514ca92433bf096b1090d2453f88a7f1142f88232620aced

Observation cc4bef7c-c292-402e-8207-a764985f2dcf · outbound

This paper cites Near-optimal prov- able uniform convergence in offline policy evaluation for rein- forcement learning.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal prov- able uniform convergence in offline policy evaluation for rein- forcement learning

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.970026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.647430Z digest=sha256:8dcb19632bbe75cd07bd3e0e5311f742a75bd7bcc11bbbe2900f4e9993d0f66f

Observation 5397d07a-78d9-4e51-baa1-2aa64fa26568 · outbound

This paper cites Near-optimal of- fline reinforcement learning via double variance reduction.

On the Statistical Complexity for Offline and Low-Adaptive Reinforcement Learning with Structures Near-optimal of- fline reinforcement learning via double variance reduction

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:16.955788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:19:16.651877Z digest=sha256:b65008a44fef996dea1af7e92120230ace79bf60438b2a43568387f544ad7792

Pith citing papers

No inbound Pith citation observations are available.