Pith. sign in

REVIEW 3 cited by

Episodic Transformer for Vision-and-Language Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.06453 v2 pith:HTGPYT3F submitted 2021-05-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords instructionstransformerlanguagechallengesepisodichistoryimprovenatural
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Interaction and navigation defined by natural language instructions in dynamic environments pose significant challenges for neural agents. This paper focuses on addressing two challenges: handling long sequence of subtasks, and understanding complex human instructions. We propose Episodic Transformer (E.T.), a multimodal transformer that encodes language inputs and the full episode history of visual observations and actions. To improve training, we leverage synthetic instructions as an intermediate representation that decouples understanding the visual appearance of an environment from the variations of natural language instructions. We demonstrate that encoding the history with a transformer is critical to solve compositional tasks, and that pretraining and joint training with synthetic instructions further improve the performance. Our approach sets a new state of the art on the challenging ALFRED benchmark, achieving 38.4% and 8.5% task success rates on seen and unseen test splits.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models

    cs.RO 2025-08 conditional novelty 6.0 of 10

    Counterfactual language-action relabeling raises instruction-following success from about 26% to 53% in real-world navigation tests.

  2. World-Consistent Data Generation for Vision-and-Language Navigation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A 3D-guided data augmentation method generates world-consistent panoramic training data that improves a VLN agent's performance on unseen environments over prior augmentation baselines.

  3. Mobile Robots through Task-Based Human Instructions using Incremental Curriculum Learning

    cs.RO 2024-12 conditional novelty 3.0 of 10

    A simulated mobile robot learns multi-step household instructions better when training is staged from short sub-goals to full instructions, but the supporting experiments lack quantitative comparison.

Pith tools