Fitted occupancy-ratio evaluation (FORE) contracts in KL divergence under only occupancy-ratio realizability, enabling offline policy evaluation without Bellman completeness.
Finite-Time Bounds for Fitted Value Iteration
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
stat.ML 2years
2026 2representative citing papers
Finite-sample risk bounds for DQN with ReLU networks are extended to τ-mixing data, showing an extra dimensionality penalty in the convergence rate due to dependence.
citing papers explorer
-
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Fitted occupancy-ratio evaluation (FORE) contracts in KL divergence under only occupancy-ratio realizability, enabling offline policy evaluation without Bellman completeness.
-
Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $\tau$-Mixing
Finite-sample risk bounds for DQN with ReLU networks are extended to τ-mixing data, showing an extra dimensionality penalty in the convergence rate due to dependence.