Pith. sign in

REVIEW 16 cited by

Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.07954 v1 pith:5W4G55OH submitted 2020-10-15 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords multilingualinstructionlanguagenavigationpathsroom-across-roomvision-and-languageaddressing
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets. It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities. Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators. We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations. We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation

    cs.CV 2026-02 conditional novelty 7.0 of 10

    LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.

  2. Goal-oriented Navigation Instruction Generation with Tour Video Priors

    cs.CV 2026-08 conditional novelty 6.0 of 10

    VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.

  3. Joint On-and-Off Policy Learning for Vision-and-Language Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    JOP-VLN combines DAgger imitation learning with GRPO reinforcement learning, using high-entropy trajectory filtering and error-correction prioritization, achieving 69.9% SR on R2R Val-Unseen.

  4. Structured Observation Language for Efficient and Generalizable Vision-Language Navigation

    cs.CV 2026-03 reject novelty 6.0 of 10

    SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...

  5. Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.

  6. GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation

    cs.RO 2025-09 conditional novelty 6.0 of 10

    GC-VLN decomposes a navigation instruction into a graph of spatial constraints, solves the constraints with an optimizer, and beats prior zero-shot methods on VLN-CE benchmarks without any training.

  7. SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps

    cs.RO 2025-08 unverdicted novelty 6.0 of 10

    A navigation agent can follow abstract hand-drawn sketch maps to reach goals in unseen indoor environments, backed by a new 54k-pair dataset and a model with a 105 percent relative SPL gain.

  8. Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A VLN agent that recursively imagines future views and layouts in a fixed-size neural grid, and adaptively aligns instruction parts to grid cells, achieves state-of-the-art success rates on R2R-CE and ObjectNav.

  9. StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.

  10. Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A whitebox adversarial attack that repaints a single 3D object can redirect or stop a pretrained Vision-and-Language Navigation agent on unseen instructions.

  11. SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments

    cs.RO 2025-07 conditional novelty 5.0 of 10

    An LLM-and-NMPC drone navigation framework that reports 42.4% success on unseen AVDN test data, versus 16.6% for NavGPT, using spatial verbalization and a path memory graph.

  12. Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.

  13. Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation

    cs.CV 2025-04 conditional novelty 5.0 of 10

    MFRA combines a hierarchical DIRformer-style fusion backbone with instruction-guided attention and a GRU history encoder, reporting improved VLN benchmark scores.

  14. Efficient and Generalizable Environmental Understanding for Visual Navigation

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Adding an auxiliary next-state prediction loss to EmbCLIP substantially improves object and point navigation in RoboTHOR and Habitat and boosts supervised vision-and-language navigation baselines.

  15. Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Arabic instructions on the R2R navigation task preserve much of GPT-4o mini's success rate but push Phi-3 and Jais to zero, indicating model capability rather than language is the decisive factor.

  16. Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.

Pith tools