REVIEW 3 cited by
Observational Overfitting in Reinforcement Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A major component of overfitting in model-free reinforcement learning (RL) involves the case where the agent may mistakenly correlate reward with certain spurious features from the observations generated by the Markov Decision Process (MDP). We provide a general framework for analyzing this scenario, which we use to design multiple synthetic benchmarks from only modifying the observation space of an MDP. When an agent overfits to different observation spaces even if the underlying MDP dynamics is fixed, we term this observational overfitting. Our experiments expose intriguing properties especially with regards to implicit regularization, and also corroborate results from previous works in RL generalization and supervised learning (SL).
Forward citations
Cited by 3 Pith papers
-
Online Training and Pruning of Deep Reinforcement Learning Networks
A method that prunes OFENet-based reinforcement learning networks during training, reducing them to a fraction of their original size with minimal performance loss.
-
Efficient Reinforcement Learning Through Adaptively Pretrained Visual Encoder
APE pretrains a ResNet18 encoder with adaptively selected augmentations and freezes its early layers during policy learning, improving sample efficiency of DreamerV3 and DrQ-v2 on several visual RL benchmarks.
-
MIGT: Memory Instance Gated Transformer Framework for Financial Portfolio Management
MIGT, a PPO-based portfolio agent using a Gated Instance Attention transformer, reports higher backtest returns and risk-adjusted ratios than 15 strategies on DJIA data for 2019-2021, but with weak statistical evidence.
Discussion (0). Continue with ORCID to comment.