REVIEW 1 cited by
Situated Mapping of Sequential Instructions to Actions with Single-step Reward Observation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose a learning approach for mapping context-dependent sequential instructions to actions. We address the problem of discourse and state dependencies with an attention-based model that considers both the history of the interaction and the state of the world. To train from start and goal states without access to demonstrations, we propose SESTRA, a learning algorithm that takes advantage of single-step reward observations and immediate expected reward maximization. We evaluate on the SCONE domains, and show absolute accuracy improvements of 9.8%-25.3% across the domains over approaches that use high-level logical representations.
Forward citations
Cited by 1 Pith paper
-
FlowDelta: Modeling Flow Information Gain in Reasoning for Conversational Machine Comprehension
Modeling the difference between consecutive reasoning states, called FlowDelta, improves conversational machine comprehension accuracy across FlowQA and BERT on CoQA, QuAC, and SCONE.
Discussion (0). Continue with ORCID to comment.