Pith. sign in

REVIEW

Sample-efficient Deep Reinforcement Learning for Dialog Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1612.06000 v1 pith:L6QY7UIT submitted 2016-12-18 cs.AI cs.LGstat.ML

classification cs.AIcs.LGstat.ML
keywords policydialoglearningmethodsdialogsgradientnumberreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Representing a dialog policy as a recurrent neural network (RNN) is attractive because it handles partial observability, infers a latent representation of state, and can be optimized with supervised learning (SL) or reinforcement learning (RL). For RL, a policy gradient approach is natural, but is sample inefficient. In this paper, we present 3 methods for reducing the number of dialogs required to optimize an RNN-based dialog policy with RL. The key idea is to maintain a second RNN which predicts the value of the current policy, and to apply experience replay to both networks. On two tasks, these methods reduce the number of dialogs/episodes required by about a third, vs. standard policy gradient methods.

Discussion (0). Continue with ORCID to comment.

Pith tools