Pith. sign in

REVIEW 2 cited by

UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.16053 v2 pith:JKVXT7IY submitted 2024-11-25 cs.CV cs.AI

UnitedVLN: Generalizable Gaussian Splatting for Continuous Vision-Language Navigation

classification cs.CV cs.AI
keywords navigationenvironmentssemanticunitedvlnfuturevisualagentcontinuous
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Vision-and-Language Navigation (VLN), where an agent follows instructions to reach a target destination, has recently seen significant advancements. In contrast to navigation in discrete environments with predefined trajectories, VLN in Continuous Environments (VLN-CE) presents greater challenges, as the agent is free to navigate any unobstructed location and is more vulnerable to visual occlusions or blind spots. Recent approaches have attempted to address this by imagining future environments, either through predicted future visual images or semantic features, rather than relying solely on current observations. However, these RGB-based and feature-based methods lack intuitive appearance-level information or high-level semantic complexity crucial for effective navigation. To overcome these limitations, we introduce a novel, generalizable 3DGS-based pre-training paradigm, called UnitedVLN, which enables agents to better explore future environments by unitedly rendering high-fidelity 360 visual images and semantic features. UnitedVLN employs two key schemes: search-then-query sampling and separate-then-united rendering, which facilitate efficient exploitation of neural primitives, helping to integrate both appearance and semantic information for more robust navigation. Extensive experiments demonstrate that UnitedVLN outperforms state-of-the-art methods on existing VLN-CE benchmarks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Instance-Enriched Semantic Maps for Visual Language Navigation

    cs.RO 2026-07 conditional novelty 6.0

    Instance-enriched 2.5D maps with LLM expert-fusion retrieval improve zero-shot object retrieval and navigation over HOV-SG while cutting map storage by ~96%.

  2. Instance-Enriched Semantic Maps for Visual Language Navigation

    cs.RO 2026-07 unverdicted novelty 5.0

    Instance-enriched 2.5D semantic maps plus LLM expert fusion improve VLN object retrieval by >17% and success by >23% while cutting storage ~96% versus 3D baselines.