Pith. sign in

REVIEW 1 cited by

Efficient decorrelation of features using Gramian in Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.08610 v1 pith:HX7EL5UZ submitted 2019-11-19 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords featureslearningsettingapproachcomplexitydemonstrategamesgramian
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning good representations is a long standing problem in reinforcement learning (RL). One of the conventional ways to achieve this goal in the supervised setting is through regularization of the parameters. Extending some of these ideas to the RL setting has not yielded similar improvements in learning. In this paper, we develop an online regularization framework for decorrelating features in RL and demonstrate its utility in several test environments. We prove that the proposed algorithm converges in the linear function approximation setting and does not change the main objective of maximizing cumulative reward. We demonstrate how to scale the approach to deep RL using the Gramian of the features achieving linear computational complexity in the number of features and squared complexity in size of the batch. We conduct an extensive empirical study of the new approach on Atari 2600 games and show a significant improvement in sample efficiency in 40 out of 49 games.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Decorrelated Soft Actor-Critic for Efficient Deep Reinforcement Learning

    cs.LG 2025-01 conditional novelty 5.0 of 10

    Decorrelated Soft Actor-Critic (DSAC) adds layerwise input decorrelation to discrete SAC and reports faster wall-clock training in 5 of 7 Atari games and better reward in 2, though the gains are partly confounded by p...

Pith tools