Pith. sign in

REVIEW 1 cited by

Towards End-to-End Learning for Dialog State Tracking and Management using Deep Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1606.02560 v2 pith:L6SZKYK2 submitted 2016-06-08 cs.AI cs.CLcs.LG

classification cs.AIcs.CLcs.LG
keywords learningdialogdeepend-to-endgamemodelproposedreinforcement
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents an end-to-end framework for task-oriented dialog systems using a variant of Deep Recurrent Q-Networks (DRQN). The model is able to interface with a relational database and jointly learn policies for both language understanding and dialog strategy. Moreover, we propose a hybrid algorithm that combines the strength of reinforcement learning and supervised learning to achieve faster learning speed. We evaluated the proposed model on a 20 Question Game conversational game simulator. Results show that the proposed method outperforms the modular-based baseline and learns a distributed representation of the latent dialog state.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing Conversational Product Recommendation via Reinforcement Learning

    cs.IR 2025-06 reject novelty 1.0 of 10

    A position paper sketching how RL (DQN, PPO, RLHF) could optimize conversational product recommendation, without any validation.

Pith tools