Pith. sign in

REVIEW 1 cited by

SF-DQN: Provable Knowledge Transfer using Successor Feature for Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.15920 v2 pith:TPWMDCL5 submitted 2024-05-24 cs.LG stat.ML

classification cs.LGstat.ML
keywords sf-dqntransferdeeplearningprovableq-functionrewardsuccessor
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper studies the transfer reinforcement learning (RL) problem where multiple RL problems have different reward functions but share the same underlying transition dynamics. In this setting, the Q-function of each RL problem (task) can be decomposed into a successor feature (SF) and a reward mapping: the former characterizes the transition dynamics, and the latter characterizes the task-specific reward function. This Q-function decomposition, coupled with a policy improvement operator known as generalized policy improvement (GPI), reduces the sample complexity of finding the optimal Q-function, and thus the SF \& GPI framework exhibits promising empirical performance compared to traditional RL methods like Q-learning. However, its theoretical foundations remain largely unestablished, especially when learning the successor features using deep neural networks (SF-DQN). This paper studies the provable knowledge transfer using SFs-DQN in transfer RL problems. We establish the first convergence analysis with provable generalization guarantees for SF-DQN with GPI. The theory reveals that SF-DQN with GPI outperforms conventional RL approaches, such as deep Q-network, in terms of both faster convergence rate and better generalization. Numerical experiments on real and synthetic RL tasks support the superior performance of SF-DQN \& GPI, aligning with our theoretical findings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Performance of Spike-based Deep Q-Learning using Ternary Neurons

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An asymmetric ternary spiking neuron with a trainable negative threshold improves deep spiking Q-network scores on six of seven Atari games, but the theoretical explanation and the headline performance metric are not ...

Pith tools