Pith. sign in

REVIEW 1 cited by

t-DGR: A Trajectory-Based Deep Generative Replay Method for Continual Learning in Decision Making

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.02576 v2 pith:HSCSJ2LB submitted 2024-01-04 cs.LG cs.AIcs.NE

classification cs.LGcs.AIcs.NE
keywords continuallearninggenerativeapproachdeepmethodreplaytasks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep generative replay has emerged as a promising approach for continual learning in decision-making tasks. This approach addresses the problem of catastrophic forgetting by leveraging the generation of trajectories from previously encountered tasks to augment the current dataset. However, existing deep generative replay methods for continual learning rely on autoregressive models, which suffer from compounding errors in the generated trajectories. In this paper, we propose a simple, scalable, and non-autoregressive method for continual learning in decision-making tasks using a generative model that generates task samples conditioned on the trajectory timestep. We evaluate our method on Continual World benchmarks and find that our approach achieves state-of-the-art performance on the average success rate metric among continual learning methods. Code is available at https://github.com/WilliamYue37/t-DGR.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continual Task Learning through Adaptive Policy Self-Composition

    cs.LG 2024-11 conditional novelty 5.0 of 10

    CompoFormer adaptively composes prior task policies in a Decision Transformer to improve stability and plasticity in continual offline RL, and the paper introduces the OCW benchmark.

Pith tools