Value-based deep RL shows TD-overfitting: large batches degrade Q-function accuracy for small networks but not large ones, enabling compute-optimal splits between model size and update frequency.
Updated 2023 Mar
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Compute-Optimal Scaling for Value-Based Deep RL
Value-based deep RL shows TD-overfitting: large batches degrade Q-function accuracy for small networks but not large ones, enabling compute-optimal splits between model size and update frequency.