Pith. sign in

REVIEW 1 cited by

Observational Overfitting in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.02975 v2 pith:YBXREELN submitted 2019-12-06 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords learningoverfittingagentobservationobservationalreinforcementanalyzingbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A major component of overfitting in model-free reinforcement learning (RL) involves the case where the agent may mistakenly correlate reward with certain spurious features from the observations generated by the Markov Decision Process (MDP). We provide a general framework for analyzing this scenario, which we use to design multiple synthetic benchmarks from only modifying the observation space of an MDP. When an agent overfits to different observation spaces even if the underlying MDP dynamics is fixed, we term this observational overfitting. Our experiments expose intriguing properties especially with regards to implicit regularization, and also corroborate results from previous works in RL generalization and supervised learning (SL).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Training and Pruning of Deep Reinforcement Learning Networks

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A method that prunes OFENet-based reinforcement learning networks during training, reducing them to a fraction of their original size with minimal performance loss.

Pith tools