Pith. sign in

REVIEW 2 minor 100 references

An agent builds a Semantic Gaussian Map from panoramic views and folds geometric, semantic, and appearance uncertainties into a 3D Value Map to guide reliable vision-language navigation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-29 18:11 UTC pith:BWQDY5CB

load-bearing objection This paper adds three explicit uncertainty channels to a Semantic Gaussian Map for VLN, which is a coherent extension but still needs the results to show whether it actually helps navigation.

arxiv 2605.26503 v1 pith:BWQDY5CB submitted 2026-05-26 cs.CV

Uncertainty-Aware Gaussian Map for Vision-Language Navigation

classification cs.CV
keywords vision-language navigationgaussian mapperceptual uncertaintysemantic mapping3d value mapembodied navigationuncertainty estimation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to show that vision-language navigation agents can make better decisions by explicitly modeling perceptual uncertainty rather than ignoring it. It constructs a Semantic Gaussian Map using differentiable 3D Gaussian primitives that capture both the geometry and semantics of the scene from panoramic observations. Three uncertainty types are then estimated and added to turn the map into a unified 3D Value Map whose values act as affordances and constraints during action selection. A sympathetic reader would care because agents that overlook uncertainty often choose unreliable moves when evidence is weak or spatial cues are ambiguous.

Core claim

The paper claims that constructing a Semantic Gaussian Map composed of differentiable 3D Gaussian primitives initialized from panoramic observations, estimating geometric uncertainty through variational perturbations of position and scale, semantic uncertainty by perturbing semantic attributes, and appearance uncertainty via Fisher Information, then incorporating all three into a unified 3D Value Map, grounds the uncertainties as affordances and constraints that support reliable navigation.

What carries the argument

The Semantic Gaussian Map (SGM) of differentiable 3D Gaussian primitives that encodes geometry and semantics, extended by uncertainty estimates into a unified 3D Value Map that supplies navigation affordances and constraints.

Load-bearing premise

The three uncertainty estimates can be computed reliably from the Gaussian primitives and will improve action selection without introducing new failure modes or excessive computational cost.

What would settle it

An ablation study on standard VLN benchmarks that removes the uncertainty integration and finds no drop in success rate or increase in failure modes would show the uncertainties are not functioning as claimed affordances.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Geometric uncertainty estimated via variational perturbations of Gaussian position and scale reveals structural reliability for action choices.
  • Semantic uncertainty obtained by perturbing Gaussian semantic attributes exposes ambiguous interpretations of the scene.
  • Appearance uncertainty measured by Fisher Information quantifies how sensitive rendered observations are to Gaussian-level changes.
  • The resulting 3D Value Map supplies grounded affordances and constraints that improve navigation performance across multiple VLN benchmarks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same uncertainty-grounding approach could be tested in other embodied tasks such as instruction following for manipulation where perceptual doubts also affect planning.
  • Measuring whether the Value Map specifically reduces failures in low-evidence regions would provide a finer test than overall benchmark scores.
  • If the Gaussian representation stays compact, the method might transfer to longer-horizon navigation without recomputing the entire map at every step.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The paper proposes modeling three forms of perceptual uncertainty (geometric, semantic, and appearance) for Vision-Language Navigation (VLN) agents. It constructs a Semantic Gaussian Map (SGM) from panoramic observations using differentiable 3D Gaussian primitives that encode geometry and semantics. Geometric uncertainty is estimated via variational perturbations of Gaussian position/scale; semantic uncertainty via perturbations of semantic attributes; and appearance uncertainty via Fisher Information on rendered observations. These are fused into a unified 3D Value Map that treats uncertainties as affordances and constraints for action selection. The approach is evaluated on multiple VLN benchmarks.

Significance. If the empirical results and ablations hold, the work provides a concrete mechanism for VLN agents to reason about perceptual uncertainty rather than ignoring it, which is a common source of failure in instruction following. The explicit construction of SGM and the three uncertainty estimators (variational on geometry/semantics, Fisher on appearance) offers a unified 3D representation that could improve robustness in ambiguous or partially observed environments. The integration of recent 3D Gaussian primitives with uncertainty quantification is a timely extension of Gaussian-map ideas to the VLN setting.

minor comments (2)
  1. [Abstract] The abstract states that uncertainties are 'incorporated into SGM, extending it into a unified 3D Value Map' and 'ground uncertainties as affordances and constraints,' but the precise mechanism by which the Value Map is queried during policy inference (e.g., how uncertainty values modulate action logits or cost functions) is not previewed; a one-sentence description of the downstream use would strengthen the claim.
  2. [Abstract] The description of Fisher Information for appearance uncertainty refers to 'sensitivity of rendered observations to Gaussian-level variations,' but does not indicate whether this is computed analytically or via sampling; a brief clarification of the computational procedure would aid reproducibility.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive assessment of our work and the recommendation for minor revision. The provided summary accurately captures the core contributions regarding uncertainty modeling in VLN via the Semantic Gaussian Map and the three uncertainty types. No specific major comments were enumerated in the report.

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper's pipeline starts from panoramic observations to initialize differentiable 3D Gaussian primitives forming the Semantic Gaussian Map, then applies standard external techniques (variational perturbations on position/scale/semantic attributes and Fisher Information on rendered observations) to compute the three uncertainty types before fusing them into the 3D Value Map. None of these steps reduce by construction to the target navigation output or to a self-citation chain; the uncertainty estimators are defined independently of the final action selection, and no fitted parameter is relabeled as a prediction. The derivation therefore remains self-contained against external benchmarks and standard probabilistic methods.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only input supplies no concrete free parameters, axioms, or invented entities; all fields left empty.

pith-pipeline@v0.9.1-grok · 5766 in / 1044 out tokens · 26075 ms · 2026-06-29T18:11:54.160022+00:00 · methodology

0 comments
read the original abstract

Vision-Language Navigation (VLN) requires an agent to navigate 3D environments following natural language instructions. During navigation, existing agents commonly encounter perceptual uncertainty, such as insufficient evidence for reliable grounding or ambiguity in interpreting spatial cues, yet they typically ignore such information when predicting actions. In this work, we explicitly model three forms of perceptual uncertainty (i.e., geometric, semantic, and appearance uncertainty) and integrate them into the agent's observation space to enable informed decision-making. Concretely, our agent first constructs a Semantic Gaussian Map (SGM), composed of differentiable 3D Gaussian primitives initialized from panoramic observations, that encodes both the geometric structure and semantic content of the environment. On top of SGM, geometric uncertainty is estimated through variational perturbations of Gaussian position and scale to assess structural reliability; semantic uncertainty is captured by perturbing Gaussian semantic attributes to reveal ambiguous interpretations; and appearance uncertainty is characterized by Fisher Information, which measures the sensitivity of rendered observations to Gaussian-level variations. These uncertainties are incorporated into SGM, extending it into a unified 3D Value Map, which grounds them as affordances and constraints that support reliable navigation. Comprehensive evaluations across multiple VLN benchmarks show the effectiveness of our agent.

Figures

Figures reproduced from arXiv: 2605.26503 by Jianzhe Gao, Rui Liu, Sida Peng, Tongtong Cao, Wenguan Wang, Yingxue Zhang, Yi Yang, Yuxuan Xu, Zhanguang Zhang.

Figure 1
Figure 1. Figure 1: Motivation. Previous VLN agents typically ignore perceptual uncertainty when making decisions. As a result, they often confuse visually similar structures (e.g., multiple doors) due to lim￾ited interior evidence ( ) and struggle when occlusions obscure spatial cues, leaving traversability ambiguous and causing unsafe or suboptimal paths ( ). In contrast, our agent explicitly models and leverages such uncer… view at source ↗
Figure 2
Figure 2. Figure 2: Pipeline overview. At each step, our agent constructs a Semantic Gaussian Map (§3.1) from its panoramic observation O = {I, D}. On top of this map, it estimates geometric U g , seman￾tic U s , and appearance U a uncertainties (§3.2) and embeds them back to obtain a unified 3D Value Map (§3.3) that grounds affordances and constraints. Finally, Gaussian representations F g derived from the value map are conc… view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative results on R2R [1]. (a) Under the instruction “straight towards the windows”, VER [17] misinterprets the layout and stops early, whereas our agent correctly follows the path and reaches the landmark. (b) Our agent bypasses the obstacle and enters the designated region, while VER halts at the “table” without completing the task. See §4.3 for more details. 4 1 2 1 2 3 4 Walk down the red-carpeted… view at source ↗
Figure 4
Figure 4. Figure 4: Representative visual results on R2R [1]. At each step, we show the constructed SGM, the rendered observations, and the aggregated uncertainty map. While SGM captures the geometry and semantic layout, the uncertainty emphasizes ambiguous regions such as reflective surfaces and repetitive structures, offering complementary cues for reliable grounding. See §4.3 for more details [PITH_FULL_IMAGE:figures/full… view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of diverse perceptual forms. From left to right: current observation, SGM, rendered observation, geometric uncertainty map, semantic uncertainty map, appearance uncertainty map. Brighter colors indicate higher uncertainty. See §4.3 for more details. ometric, semantic, and appearance components). SGM captures the geometry and semantic layout, while the uncertainty highlights visually ambiguous… view at source ↗
Figure 6
Figure 6. Figure 6: Ground Truth vs Rendered Observations. The renderings closely match the ground truth, supporting the Fisher-based appearance uncertainty proxy. We provide visual comparisons between the rendered observations and the ground truth [PITH_FULL_IMAGE:figures/full_fig_p020_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Failure Cases. (a) Our agent stops once “the sofa” comes into view, as the current observation already provides sufficient evidence of the target, creating confusion about whether further steps are required. (b) Our agent halts at the doorway instead of reaching “the gomoku board” near the bed, since the board lies inside the room and cannot be observed from the entrance, leaving the agent uncertain and le… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

100 extracted references · 2 canonical work pages · 2 internal anchors

  1. [1]

    Reid, Stephen Gould, and Anton van den Hengel

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S ¨underhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: Interpreting visually- grounded navigation instructions in real environments. InCVPR, 2018. 1, 2, 3, 6, 7, 8, 9, 10, 17, 18, 21, 23

  2. [2]

    Adaptive zone-aware hierarchical planner for vision-language navigation

    Chen Gao, Xingyu Peng, Mi Yan, He Wang, Lirong Yang, Haibing Ren, Hongsheng Li, and Si Liu. Adaptive zone-aware hierarchical planner for vision-language navigation. InCVPR, 2023

  3. [3]

    Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE TPAMI, 2024

    Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, and Liang Wang. Etpnav: Evolving topological planning for vision-language navigation in continuous environments.IEEE TPAMI, 2024

  4. [4]

    Vision-language navigation with energy-based policy

    Rui Liu, Wenguan Wang, and Yi Yang. Vision-language navigation with energy-based policy. InNeurIPS, 2024

  5. [5]

    3d gaussian map with open-set semantic grouping for vision- language navigation

    Jianzhe Gao, Rui Liu, and Wenguan Wang. 3d gaussian map with open-set semantic grouping for vision- language navigation. InICCV, 2025. 1, 2, 7, 8

  6. [6]

    Reinforcement learning with neural radiance fields

    Danny Driess, Ingmar Schubert, Pete Florence, Yunzhu Li, and Marc Toussaint. Reinforcement learning with neural radiance fields. InNeurIPS, 2022. 1

  7. [7]

    Towards versatile embodied navigation

    Hanqing Wang, Wei Liang, Luc V Gool, and Wenguan Wang. Towards versatile embodied navigation. In NeurIPS, 2022

  8. [8]

    Counterfactual cycle- consistent learning for instruction following and generation in vision-language navigation

    Hanqing Wang, Wei Liang, Jianbing Shen, Luc Van Gool, and Wenguan Wang. Counterfactual cycle- consistent learning for instruction following and generation in vision-language navigation. InCVPR, 2022

  9. [9]

    History-enhanced two-stage trans- former for aerial vision-and-language navigation

    Xichen Ding, Jianzhe Gao, Cong Pan, Wenguan Wang, and Jie Qin. History-enhanced two-stage trans- former for aerial vision-and-language navigation. InAAAI, 2026. 1

  10. [10]

    Learning to navigate unseen environments: Back translation with environmental dropout

    Hao Tan, Licheng Yu, and Mohit Bansal. Learning to navigate unseen environments: Back translation with environmental dropout. InNAACL, 2019. 1, 2

  11. [11]

    Think global, act local: Dual-scale graph transformer for vision-and-language navigation

    Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Think global, act local: Dual-scale graph transformer for vision-and-language navigation. InCVPR, 2022. 1, 2, 6, 7, 8, 9, 17, 18, 19, 21

  12. [12]

    Learning navigational visual representations with semantic map supervision

    Yicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt, Trung Bui, Stephen Gould, and Hao Tan. Learning navigational visual representations with semantic map supervision. InCVPR, 2023. 1 10 Published as a conference paper at ICLR 2026

  13. [13]

    Active perception for visual-language navigation.IJCV, 131(3):607–625, 2023

    Hanqing Wang, Wenguan Wang, Wei Liang, Steven CH Hoi, Jianbing Shen, and Luc Van Gool. Active perception for visual-language navigation.IJCV, 131(3):607–625, 2023

  14. [14]

    Room-object entity prompting and reasoning for embodied referring expression.IEEE TPAMI, 46(2):994–1010, 2023

    Chen Gao, Si Liu, Jinyu Chen, Luting Wang, Qi Wu, Bo Li, and Qi Tian. Room-object entity prompting and reasoning for embodied referring expression.IEEE TPAMI, 46(2):994–1010, 2023. 1

  15. [15]

    Bevbert: Multimodal map pre-training for language-guided navigation

    Dong An, Yuankai Qi, Yangguang Li, Yan Huang, Liang Wang, Tieniu Tan, and Jing Shao. Bevbert: Multimodal map pre-training for language-guided navigation. InICCV, 2023. 1, 7, 8, 18, 19, 21

  16. [16]

    Gridmm: Grid memory map for vision-and-language navigation

    Zihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu, and Shuqiang Jiang. Gridmm: Grid memory map for vision-and-language navigation. InICCV, 2023. 1, 7, 8

  17. [17]

    V olumetric environment representation for vision-language navi- gation

    Rui Liu, Wenguan Wang, and Yi Yang. V olumetric environment representation for vision-language navi- gation. InCVPR, 2024. 1, 6, 7, 8, 17, 19, 21

  18. [18]

    Efficient training of artificial neural networks for autonomous navigation.Neural computation, 3(1), 1991

    Dean A Pomerleau. Efficient training of artificial neural networks for autonomous navigation.Neural computation, 3(1), 1991. 1

  19. [19]

    Apprenticeship learning via inverse reinforcement learning

    Pieter Abbeel and Andrew Y Ng. Apprenticeship learning via inverse reinforcement learning. InICML,

  20. [20]

    Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation

    Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation. InCVPR, 2019. 1, 2, 7

  21. [21]

    Vln bert: A recurrent vision-and-language bert for navigation

    Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez-Opazo, and Stephen Gould. Vln bert: A recurrent vision-and-language bert for navigation. InCVPR, 2021. 6, 7

  22. [22]

    Active visual information gathering for vision-language navigation

    Hanqing Wang, Wenguan Wang, Tianmin Shu, Wei Liang, and Jianbing Shen. Active visual information gathering for vision-language navigation. InECCV, 2020

  23. [23]

    Target-driven structured transformer planner for vision-language navigation

    Yusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang, Lirong Yang, Haibing Ren, Huaxia Xia, and Si Liu. Target-driven structured transformer planner for vision-language navigation. InACM MM, 2022. 1

  24. [24]

    Pathdreamer: A world model for indoor navigation

    Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson. Pathdreamer: A world model for indoor navigation. InICCV, 2021. 1, 2, 23

  25. [25]

    Dreamwalker: Mental planning for continuous vision-language navigation

    Hanqing Wang, Wei Liang, Luc Van Gool, and Wenguan Wang. Dreamwalker: Mental planning for continuous vision-language navigation. InICCV, 2023. 2

  26. [26]

    Navigation world models

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation world models. In CVPR, 2025. 1, 2

  27. [27]

    Room-across-room: Multi- lingual vision-and-language navigation with dense spatiotemporal grounding

    Alexander Ku, Peter Anderson, Roma Patel, Eugene Ie, and Jason Baldridge. Room-across-room: Multi- lingual vision-and-language navigation with dense spatiotemporal grounding. InEMNLP, 2020. 2, 6, 7, 8, 17, 23

  28. [28]

    Reverie: Remote embodied visual referring expression in real indoor environments

    Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expression in real indoor environments. InCVPR,

  29. [29]

    2, 3, 6, 7, 9, 10, 17, 23

  30. [30]

    Speaker-follower models for vision-and-language navigation

    Daniel Fried, Ronghang Hu, V olkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language navigation. InNeurIPS, 2018. 2

  31. [31]

    Bird’s-eye-view scene graph for vision-language navigation

    Rui Liu, Xiaohan Wang, Wenguan Wang, and Yi Yang. Bird’s-eye-view scene graph for vision-language navigation. InICCV, 2023. 2, 6, 7, 8

  32. [32]

    Evolving graphical planner: Contextual global planning for vision-and-language navigation

    Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. Evolving graphical planner: Contextual global planning for vision-and-language navigation. InNeurIPS, 2020. 2

  33. [33]

    Topological and semantic map generation for mobile robot indoor navigation

    Yujing Chen, Jinmin Zhang, and Yunjiang Lou. Topological and semantic map generation for mobile robot indoor navigation. InICIRA, 2021

  34. [34]

    Structured scene mem- ory for vision-language navigation

    Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. Structured scene mem- ory for vision-language navigation. InCVPR, 2021. 2

  35. [35]

    Language and visual entity relationship graph for agent navigation

    Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. Language and visual entity relationship graph for agent navigation. InNeurIPS, 2020. 2 11 Published as a conference paper at ICLR 2026

  36. [36]

    Migc: Multi-instance generation controller for text-to-image synthesis

    Dewei Zhou, You Li, Fan Ma, Xiaoting Zhang, and Yi Yang. Migc: Multi-instance generation controller for text-to-image synthesis. InCVPR, 2024. 2

  37. [37]

    Controllable navigation instruction generation with chain of thought prompting

    Xianghao Kong, Jinyu Chen, Wenguan Wang, Hang Su, Xiaolin Hu, Yi Yang, and Si Liu. Controllable navigation instruction generation with chain of thought prompting. InECCV, 2024

  38. [38]

    Navigation instruction generation with bev perception and large language models

    Sheng Fan, Rui Liu, Wenguan Wang, and Yi Yang. Navigation instruction generation with bev perception and large language models. InECCV, 2024

  39. [39]

    Scene map-based prompt tuning for navigation instruction generation

    Sheng Fan, Rui Liu, Wenguan Wang, and Yi Yang. Scene map-based prompt tuning for navigation instruction generation. InCVPR, 2025

  40. [40]

    Bootstrapping language-guided navigation learning with self-refining data flywheel

    Zun Wang, Jialu Li, Yicong Hong, Songze Li, Kunchang Li, Shoubin Yu, Yi Wang, Yu Qiao, Yali Wang, Mohit Bansal, et al. Bootstrapping language-guided navigation learning with self-refining data flywheel. InICLR, 2025. 2

  41. [41]

    Do visual imaginations improve vision-and-language navigation agents? InCVPR, 2025

    Akhil Perincherry, Jacob Krantz, and Stefan Lee. Do visual imaginations improve vision-and-language navigation agents? InCVPR, 2025. 2, 7, 8

  42. [42]

    Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

    Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S ¨underhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. InCVPR, 2018. 2

  43. [43]

    Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation

    Xin Wang, Wenhan Xiong, Hongmin Wang, and William Yang Wang. Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation. InECCV, 2018. 2

  44. [44]

    Cosmo: Com- bination of selective memorization for low-cost vision-and-language navigation

    Siqi Zhang, Yanyuan Qiao, Qunbo Wang, Zike Yan, Qi Wu, Zhihua Wei, and Jing Liu. Cosmo: Com- bination of selective memorization for low-cost vision-and-language navigation. InICCV, 2025. 2, 7, 8

  45. [45]

    3d gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4), 2023

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4), 2023. 2, 3

  46. [46]

    Uncertainty-aware 3d reconstruction for dynamic underwater scenes

    Rui Liu, Zhibo Duan, Jianzhe Gao, Yi Yang, and Wenguan Wang. Uncertainty-aware 3d reconstruction for dynamic underwater scenes. InICLR, 2026

  47. [47]

    Open-set semantic gaussian splatting slam with expandable representation

    Yucheng Yan, Chen Liang, Wenguan Wang, and Yi Yang. Open-set semantic gaussian splatting slam with expandable representation. InICLR, 2026. 2

  48. [48]

    Gaussnav: Gaussian splatting for visual navigation.IEEE TPAMI, 2025

    Xiaohan Lei, Min Wang, Wengang Zhou, and Houqiang Li. Gaussnav: Gaussian splatting for visual navigation.IEEE TPAMI, 2025. 2

  49. [49]

    Gaussian-based world model: Gaussian priors for voxel-based occupancy prediction and future motion prediction

    Tuo Feng, Wenguan Wang, and Yi Yang. Gaussian-based world model: Gaussian priors for voxel-based occupancy prediction and future motion prediction. InICCV, 2025

  50. [50]

    Underwater visual slam with depth uncertainty and medium modeling

    Rui Liu, Sheng Fan, Wenguan Wang, and Yi Yang. Underwater visual slam with depth uncertainty and medium modeling. InICCV, 2025. 2

  51. [51]

    Llm as copilot for coarse-grained vision- and-language navigation

    Yanyuan Qiao, Qianyi Liu, Jiajun Liu, Jing Liu, and Qi Wu. Llm as copilot for coarse-grained vision- and-language navigation. InECCV, 2024. 2

  52. [52]

    Benchmarking uncertainty disentanglement: Specialized uncertainties for specialized tasks

    B ´alint Mucs ´anyi, Michael Kirchhof, and Seong Joon Oh. Benchmarking uncertainty disentanglement: Specialized uncertainties for specialized tasks. InNeurIPS, 2024. 3

  53. [53]

    Springer Science & Business Media, 2012

    Radford M Neal.Bayesian learning for neural networks, volume 118. Springer Science & Business Media, 2012. 3

  54. [54]

    Stochastic neural radiance fields: Quantifying uncertainty in implicit 3d representations

    Jianxiong Shen, Adria Ruiz, Antonio Agudo, and Francesc Moreno-Noguer. Stochastic neural radiance fields: Quantifying uncertainty in implicit 3d representations. In3DV, 2021

  55. [55]

    Conditional-flow nerf: Ac- curate 3d modelling with reliable uncertainty quantification

    Jianxiong Shen, Antonio Agudo, Francesc Moreno-Noguer, and Adria Ruiz. Conditional-flow nerf: Ac- curate 3d modelling with reliable uncertainty quantification. InECCV, 2022. 3

  56. [56]

    Simple and scalable predictive un- certainty estimation using deep ensembles

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive un- certainty estimation using deep ensembles. InNeurIPS, 2017. 3

  57. [57]

    Accurate uncertainty estimation and decomposition in ensemble learning

    Jeremiah Liu, John Paisley, Marianthi-Anna Kioumourtzoglou, and Brent Coull. Accurate uncertainty estimation and decomposition in ensemble learning. InNeurIPS, 2019. 12 Published as a conference paper at ICLR 2026

  58. [58]

    Pitfalls of in-domain un- certainty estimation and ensembling in deep learning

    Arsenii Ashukha, Alexander Lyzhov, Dmitry Molchanov, and Dmitry Vetrov. Pitfalls of in-domain un- certainty estimation and ensembling in deep learning. InICLR, 2020

  59. [59]

    Density-aware nerf ensembles: Quantifying predictive uncertainty in neural radiance fields

    Niko S ¨underhauf, Jad Abou-Chakra, and Dimity Miller. Density-aware nerf ensembles: Quantifying predictive uncertainty in neural radiance fields. InICRA, 2023

  60. [60]

    Implicit variational infer- ence for high-dimensional posteriors

    Anshuk Uppal, Kristoffer Stensbo-Smidt, Wouter Boomsma, and Jes Frellsen. Implicit variational infer- ence for high-dimensional posteriors. InNeurIPS, 2023. 3

  61. [61]

    Sampling with riemannian hamiltonian monte carlo in a constrained space

    Yunbum Kook, Yin-Tat Lee, Ruoqi Shen, and Santosh Vempala. Sampling with riemannian hamiltonian monte carlo in a constrained space. InNeurIPS, 2022. 3

  62. [62]

    Dropout as a bayesian approximation: Representing model uncer- tainty in deep learning

    Yarin Gal and Zoubin Ghahramani. Dropout as a bayesian approximation: Representing model uncer- tainty in deep learning. InICML, 2016. 3

  63. [63]

    Variational dropout and the local reparameterization trick

    Durk P Kingma, Tim Salimans, and Max Welling. Variational dropout and the local reparameterization trick. InNeurIPS, 2015. 3

  64. [64]

    Bayes’ rays: Uncer- tainty quantification for neural radiance fields

    Lily Goli, Cody Reading, Silvia Sell ´an, Alec Jacobson, and Andrea Tagliasacchi. Bayes’ rays: Uncer- tainty quantification for neural radiance fields. InCVPR, 2024. 3, 5

  65. [65]

    Sketched lanczos uncertainty score: a low-memory summary of the fisher information

    Marco Miani, Lorenzo Beretta, and Søren Hauberg. Sketched lanczos uncertainty score: a low-memory summary of the fisher information. InNeurIPS, 2024. 5

  66. [66]

    Bounding the invertibility of privacy- preserving instance encoding using fisher information

    Kiwan Maeng, Chuan Guo, Sanjay Kariyappa, and G Edward Suh. Bounding the invertibility of privacy- preserving instance encoding using fisher information. InNeurIPS, 2023. 3, 5

  67. [67]

    Variational multi-scale representation for estimating uncertainty in 3d gaussian splatting

    Ruiqi Li and Yiu-ming Cheung. Variational multi-scale representation for estimating uncertainty in 3d gaussian splatting. InNeurIPS, 2024. 3, 4

  68. [68]

    Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting

    Alex Hanson, Allen Tu, Vasu Singla, Mayuka Jayawardhana, Matthias Zwicker, and Tom Goldstein. Pup 3d-gs: Principled uncertainty pruning for 3d gaussian splatting. InCVPR, 2025. 3, 5

  69. [69]

    SAM 2: Segment Anything in Images and Videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R ¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024. 4, 7, 18

  70. [70]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021. 4

  71. [71]

    An interactive navigation method with effect-oriented affordance

    Xiaohan Wang, Yuehu Liu, Xinhang Song, Yuyi Liu, Sixian Zhang, and Shuqiang Jiang. An interactive navigation method with effect-oriented affordance. InCVPR, 2024. 5

  72. [72]

    V oxposer: Compos- able 3d value maps for robotic manipulation with language models

    Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. V oxposer: Compos- able 3d value maps for robotic manipulation with language models. InCoRL, 2023. 5

  73. [73]

    Learning semantics-aware locomotion skills from human demonstration

    Yuxiang Yang, Xiangyun Meng, Wenhao Yu, Tingnan Zhang, Jie Tan, and Byron Boots. Learning semantics-aware locomotion skills from human demonstration. InCoRL, 2023. 5

  74. [74]

    Open-fusion: Real-time open-vocabulary 3d mapping and queryable scene representation

    Kashu Yamazaki, Taisei Hanyu, Khoa V o, Thang Pham, Minh Tran, Gianfranco Doretto, Anh Nguyen, and Ngan Le. Open-fusion: Real-time open-vocabulary 3d mapping and queryable scene representation. InICRA, 2024. 5

  75. [75]

    Image quality assessment: from error visibility to structural similarity.IEEE TIP, 13(4):600–612, 2004

    Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE TIP, 13(4):600–612, 2004. 6

  76. [76]

    Bert: Pre-training of deep bidirec- tional transformers for language understanding

    Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirec- tional transformers for language understanding. InNAACL, 2019. 6

  77. [77]

    History aware multimodal trans- former for vision-and-language navigation

    Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. History aware multimodal trans- former for vision-and-language navigation. InNeurIPS, 2021. 6, 7, 8

  78. [78]

    Scene-intuitive agent for remote embodied visual grounding

    Xiangru Lin, Guanbin Li, and Yizhou Yu. Scene-intuitive agent for remote embodied visual grounding. InCVPR, 2021. 6

  79. [79]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In Yoshua Bengio and Yann LeCun, editors,ICLR, 2015. 6 13 Published as a conference paper at ICLR 2026

  80. [80]

    Gordon, and Drew Bagnell

    St ´ephane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. InAISTATS, 2011. 7, 17

Showing first 80 references.