REVIEW 3 cited by
Episodic Transformer for Vision-and-Language Navigation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Interaction and navigation defined by natural language instructions in dynamic environments pose significant challenges for neural agents. This paper focuses on addressing two challenges: handling long sequence of subtasks, and understanding complex human instructions. We propose Episodic Transformer (E.T.), a multimodal transformer that encodes language inputs and the full episode history of visual observations and actions. To improve training, we leverage synthetic instructions as an intermediate representation that decouples understanding the visual appearance of an environment from the variations of natural language instructions. We demonstrate that encoding the history with a transformer is critical to solve compositional tasks, and that pretraining and joint training with synthetic instructions further improve the performance. Our approach sets a new state of the art on the challenging ALFRED benchmark, achieving 38.4% and 8.5% task success rates on seen and unseen test splits.
Forward citations
Cited by 3 Pith papers
-
CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.
-
World-Consistent Data Generation for Vision-and-Language Navigation
A 3D-guided data augmentation method generates world-consistent panoramic training data that improves a VLN agent's performance on unseen environments over prior augmentation baselines.
-
Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning
A simulated mobile robot learns multi-step household instructions better when training is staged from short sub-goals to full instructions, but the supporting experiments lack quantitative comparison.
Discussion (0). Continue with ORCID to comment.