REVIEW 16 cited by
Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We introduce Room-Across-Room (RxR), a new Vision-and-Language Navigation (VLN) dataset. RxR is multilingual (English, Hindi, and Telugu) and larger (more paths and instructions) than other VLN datasets. It emphasizes the role of language in VLN by addressing known biases in paths and eliciting more references to visible entities. Furthermore, each word in an instruction is time-aligned to the virtual poses of instruction creators and validators. We establish baseline scores for monolingual and multilingual settings and multitask learning when including Room-to-Room annotations. We also provide results for a model that learns from synchronized pose traces by focusing only on portions of the panorama attended to in human demonstrations. The size, scope and detail of RxR dramatically expands the frontier for research on embodied language agents in simulated, photo-realistic environments.
Forward citations
Cited by 16 Pith papers
-
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
LangMap is a human-verified navigation benchmark with 18K tasks spanning scene-, room-, region-, and instance-level goals in real-world 3D scans, covering 414 object categories.
-
Goal-oriented Navigation Instruction Generation with Tour Video Priors
VideoNIG tests whether multimodal models can turn tour videos into executable navigation instructions, and a two-stage curriculum with preference optimization improves their spatial reasoning.
-
Joint On-and-Off Policy Learning for Vision-and-Language Navigation
JOP-VLN combines DAgger imitation learning with GRPO reinforcement learning, using high-entropy trajectory filtering and error-correction prioritization, achieving 69.9% SR on R2R Val-Unseen.
-
Structured Observation Language for Efficient and Generalizable Vision-Language Navigation
SOL-Nav encodes RGB-D observations as grid-organized text and uses a 0.6B text-embedding model with four classification heads to predict navigation action blocks, reporting SOTA/comparable R2R-CE/RxR-CE results with 1...
-
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.
-
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation
GC-VLN decomposes a navigation instruction into a graph of spatial constraints, solves the constraints with an optimizer, and beats prior zero-shot methods on VLN-CE benchmarks without any training.
-
SkeNa: Learning to Navigate Unseen Environments Based on Abstract Hand-Drawn Maps
A navigation agent can follow abstract hand-drawn sketch maps to reach goals in unseen indoor environments, backed by a new 54k-pair dataset and a model with a 105 percent relative SPL gain.
-
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
A VLN agent that recursively imagines future views and layouts in a fixed-size neural grid, and adaptively aligns instruction parts to grid cells, achieves state-of-the-art success rates on R2R-CE and ObjectNav.
-
StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling
A streaming navigation framework combining a sliding-window KV cache with depth-based token pruning achieves state-of-the-art results on VLN-CE benchmarks with bounded context and low latency.
-
Hijacking Vision-and-Language Navigation Agents with Adversarial Environmental Attacks
A whitebox adversarial attack that repaints a single 3D object can redirect or stop a pretrained Vision-and-Language Navigation agent on unseen instructions.
-
SkyVLN: Vision-and-Language Navigation and NMPC Control for UAVs in Urban Environments
An LLM-and-NMPC drone navigation framework that reports 42.4% success on unseen AVDN test data, versus 16.6% for NavGPT, using spatial verbalization and a path memory graph.
-
Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
A dual-branch text-imagination system, with one LLM branch for state estimation and one for candidate-direction description, improves R2R navigation success over prior LLM-based VLN methods.
-
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
MFRA combines a hierarchical DIRformer-style fusion backbone with instruction-guided attention and a GRU history encoder, reporting improved VLN benchmark scores.
-
Efficient and Generalizable Environmental Understanding for Visual Navigation
Adding an auxiliary next-state prediction loss to EmbCLIP substantially improves object and point navigation in RoboTHOR and Habitat and boosts supervised vision-and-language navigation baselines.
-
Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models
Arabic instructions on the R2R navigation task preserve much of GPT-4o mini's success rate but push Phi-3 and Jais to zero, indicating model capability rather than language is the decisive factor.
-
Modality-Aware Feature Matching in Visual and Vision-Language Applications: A Comprehensive Survey
A survey organizing feature matching research by modality, from SIFT to transformer-based dense matchers and vision-language models.
Discussion (0). Continue with ORCID to comment.