REVIEW 2 major objections 90 references
3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation
T0 review · 2 major / 0 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read A 3D Gaussian map initialized from point clouds and grouped by open-set semantics enables agents to navigate from language instructions in unseen spaces.
desk verdict The 3D Gaussian map with open-set grouping is a reasonable integration for VLN but the abstract supplies zero numbers or baselines so the gains cannot be checked. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The 3D Gaussian Map with Open-Set Semantic Grouping, which converts sparse point clouds into semantically clustered differentiable primitives that carry both geometry and open-world labels for downstream action prediction.
What would settle it
Ablating the open-set semantic grouping step on the REVERIE benchmark and observing no gain in success rate or SPL over a baseline that uses only the raw Gaussian primitives would falsify the claim that the grouping step is what enables reliable multi-level prediction.
Extended reading notes
Core claim
The central claim is that an Egocentric Scene Map of 3D Gaussians, initialized from pseudo-lidar and enriched by Open-Set Semantic Grouping into instance and category memberships, yields a unified 3D Gaussian Map that supports Multi-Level Action Prediction combining spatial-semantic cues at multiple granularities, thereby improving agent decision-making in complex, unseen 3D environments for vision-language navigation.
Load-bearing premise
Sparse pseudo-lidar point clouds supply enough geometric structure for the open-set grouping step to produce a map that remains reliable for action prediction in environments never seen during training.
Editorial extensions
If this is right
- Agents obtain spatial-semantic cues at multiple granularities for each decision step.
- The same map representation supports navigation on R2R, R4R, and REVERIE without task-specific retraining.
- Open-set grouping removes the need for closed-world object vocabularies when encountering novel items.
- Online map construction from egocentric views allows continuous updating during traversal.
Reading between the lines
- The grouped Gaussian primitives could serve as input for other embodied tasks such as object rearrangement or question answering that also require 3D instance awareness.
- Because the grouping operates in an open-set manner, the map might transfer to environments whose object categories were never labeled in the original training data.
- Multi-level prediction could be extended to handle instructions of varying linguistic complexity by weighting granularity levels dynamically.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript proposes a 3D Gaussian Map representation for vision-language navigation. It constructs an egocentric scene map online by initializing differentiable 3D Gaussians from sparse pseudo-lidar point clouds to supply geometric priors, applies an Open-Set Semantic Grouping operation to assign instance-level or category semantics to each Gaussian, and introduces a Multi-Level Action Prediction strategy that fuses spatial-semantic cues at multiple granularities for decision making. The approach is claimed to be validated through extensive experiments on the R2R, R4R, and REVERIE benchmarks.
Significance. If the performance gains are shown to be robust under standard VLN evaluation protocols, the work would offer a unified differentiable 3D representation that jointly encodes geometry and open-set semantics, addressing a recognized limitation of prior map-based VLN methods that either ignore fine-grained 3D structure or rely on closed-set labels. The engineering integration of 3D Gaussians with open-set grouping could serve as a reusable scene representation for other embodied tasks.
major comments (2)
- [Abstract] Abstract: the central claim that the method 'validate[s] the effectiveness' on R2R/R4R/REVERIE val-unseen splits rests on the assumption that sparse pseudo-lidar initialization supplies sufficiently dense and accurate geometric priors for open-set grouping to recover reliable instance semantics without dense reconstruction. No quantitative characterization of pseudo-lidar density, depth error distribution, or failure modes in textureless regions is supplied, which is load-bearing for the generalization argument.
- [Abstract] Abstract: the reported validation provides no information on baseline comparisons, error bars, data splits, or ablation controls. Without these, it is impossible to determine whether the multi-level prediction gains are attributable to the proposed map or to other design choices.
Simulated Author's Rebuttal
We thank the referee for the thoughtful comments on our manuscript. We address each major comment point-by-point below.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central claim that the method 'validate[s] the effectiveness' on R2R/R4R/REVERIE val-unseen splits rests on the assumption that sparse pseudo-lidar initialization supplies sufficiently dense and accurate geometric priors for open-set grouping to recover reliable instance semantics without dense reconstruction. No quantitative characterization of pseudo-lidar density, depth error distribution, or failure modes in textureless regions is supplied, which is load-bearing for the generalization argument.
Authors: We agree that a quantitative characterization of the pseudo-lidar properties would strengthen the generalization claims. In the revised version we will add a dedicated analysis (new subsection or appendix) reporting point density statistics, depth error distributions, and failure cases in textureless regions across the R2R/R4R/REVERIE environments. revision: yes
-
Referee: [Abstract] Abstract: the reported validation provides no information on baseline comparisons, error bars, data splits, or ablation controls. Without these, it is impossible to determine whether the multi-level prediction gains are attributable to the proposed map or to other design choices.
Authors: The abstract is intentionally concise. The full manuscript already contains the requested information: baseline comparisons appear in Tables 1–3, standard val-unseen splits are used throughout, ablation controls are reported in Section 4.3, and error bars are shown for key metrics. These elements allow readers to attribute performance gains to the proposed components. revision: no
Circularity Check
No circularity: engineering pipeline with external validation
full rationale
The paper describes an online construction pipeline: 3D Gaussians initialized from sparse pseudo-lidar point clouds, enriched by open-set semantic grouping into a unified map, then used for multi-level action prediction. No equations, fitted parameters, or self-citations are presented that reduce the claimed performance gains on R2R/R4R/REVERIE to quantities defined by the method's own inputs. The central claims rest on empirical results from public benchmarks rather than any self-referential derivation or uniqueness theorem imported from prior author work. This is the expected non-finding for an applied CV integration paper.
Assumptions & free parameters
Cite this review
Pith. "Pith review of 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation." pith.science (2026). https://pith.science/paper/XBVRRRLH
@misc{pith2026260526500,
author = {Pith},
title = {Pith review of: 3D Gaussian Map with Open-Set Semantic Grouping for Vision-Language Navigation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XBVRRRLH}},
note = {Machine review of arXiv:2605.26500}
}
read the original abstract
Vision-language navigation (VLN) requires an agent to traverse complex 3D environments based on natural language instructions, necessitating a thorough scene understanding. While existing works equip agents with various scene representations to enhance spatial awareness, they often neglect the complex 3D geometry and rich semantics in VLN scenarios, limiting the ability to generalize across diverse and unseen environments. To address these challenges, this work proposes a 3D Gaussian Map that represents the environment as a set of differentiable 3D Gaussians and accordingly develops a navigation strategy for VLN. Specifically, Egocentric Scene Map is constructed online by initializing 3D Gaussians from sparse pseudo-lidar point clouds, providing informative geometric priors for scene understanding. Each Gaussian primitive is further enriched through Open-Set Semantic Grouping operation, which groups 3D Gaussians based on their membership in object instances or stuff categories within the open world, resulting in a unified 3D Gaussian Map. Building on this map, Multi-Level Action Prediction strategy, which combines spatial-semantic cues at multiple granularities, is designed to assist agents in decision-making. Extensive experiments conducted on three public benchmarks (i.e., R2R, R4R, and REVERIE) validate the effectiveness of our method.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Bevbert: Multimodal map pre-training for language-guided navigation
Dong An, Yuankai Qi, Yangguang Li, Yan Huang, Liang Wang, Tieniu Tan, and Jing Shao. Bevbert: Multimodal map pre-training for language-guided navigation. InICCV, 2023. 1, 2, 6, 7
2023
-
[2]
Etpnav: Evolving topo- logical planning for vision-language navigation in continu- ous environments.IEEE TPAMI, 2024
Dong An, Hanqing Wang, Wenguan Wang, Zun Wang, Yan Huang, Keji He, and Liang Wang. Etpnav: Evolving topo- logical planning for vision-language navigation in continu- ous environments.IEEE TPAMI, 2024. 1
2024
-
[3]
Reid, Stephen Gould, and Anton van den Hengel
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko S¨underhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. Vision-and-language navigation: In- terpreting visually-grounded navigation instructions in real environments. InCVPR, 2018. 1, 2, 3, 5, 6, 7, 8
2018
-
[4]
3d semantic parsing of large-scale indoor spaces
Iro Armeni, Ozan Sener, Amir R Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese. 3d semantic parsing of large-scale indoor spaces. InCVPR, 2016. 2
2016
-
[5]
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Hal- ber, Matthias Niebner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang. Matterport3d: Learning from rgb-d data in indoor environments. In3DV, 2017. 3
2017
-
[6]
Object goal naviga- tion using goal-oriented semantic exploration
Devendra Singh Chaplot, Dhiraj Prakashchand Gandhi, Ab- hinav Gupta, and Russ R Salakhutdinov. Object goal naviga- tion using goal-oriented semantic exploration. InNeurIPS,
-
[7]
Neural topological slam for vi- sual navigation
Devendra Singh Chaplot, Ruslan Salakhutdinov, Abhinav Gupta, and Saurabh Gupta. Neural topological slam for vi- sual navigation. InCVPR, 2020. 1, 2
2020
-
[8]
pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction
David Charatan, Sizhe Lester Li, Andrea Tagliasacchi, and Vincent Sitzmann. pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. InCVPR,
Show all 90 references
-
[9]
A survey on 3d gaussian splatting.arXiv preprint arXiv:2401.03890, 2024
Guikun Chen and Wenguan Wang. A survey on 3d gaussian splatting.arXiv preprint arXiv:2401.03890, 2024. 3
2024 arXiv
-
[10]
Learning active camera for multi-object navigation
Peihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu, Wen- bing Huang, Thomas Li, Mingkui Tan, and Chuang Gan. Learning active camera for multi-object navigation. In NeurIPS, 2022. 2
2022
-
[11]
Weakly- supervised multi-granularity map learning for vision-and- language navigation
Peihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng, Thomas Li, Mingkui Tan, and Chuang Gan. Weakly- supervised multi-granularity map learning for vision-and- language navigation. InNeurIPS, 2022. 1
2022
-
[12]
History aware multimodal transformer for vision-and-language navigation
Shizhe Chen, Pierre-Louis Guhur, Cordelia Schmid, and Ivan Laptev. History aware multimodal transformer for vision-and-language navigation. InNeurIPS, 2021. 5, 6
2021
-
[13]
Think global, act local: Dual-scale graph transformer for vision-and-language navi- gation
Shizhe Chen, Pierre-Louis Guhur, Makarand Tapaswi, Cordelia Schmid, and Ivan Laptev. Think global, act local: Dual-scale graph transformer for vision-and-language navi- gation. InCVPR, 2022. 1, 3, 5, 6
2022
-
[14]
Text-to-3d using gaussian splatting
Zilong Chen, Feng Wang, Yikai Wang, and Huaping Liu. Text-to-3d using gaussian splatting. InCVPR, 2024. 3
2024
-
[15]
Scannet: Richly-annotated 3d reconstructions of indoor scenes
Angela Dai, Angel X Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In CVPR, 2017. 2
2017
-
[16]
Evolving graphical planner: Contextual global planning for vision-and-language navigation
Zhiwei Deng, Karthik Narasimhan, and Olga Russakovsky. Evolving graphical planner: Contextual global planning for vision-and-language navigation. InNeurIPS, 2020. 1, 2, 6
2020
-
[17]
Unconstrained scene generation with locally conditioned radiance fields
Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W Taylor, and Joshua M Susskind. Unconstrained scene generation with locally conditioned radiance fields. In ICCV, 2021. 2
2021
-
[18]
Reinforcement learning with neural ra- diance fields
Danny Driess, Ingmar Schubert, Pete Florence, Yunzhu Li, and Marc Toussaint. Reinforcement learning with neural ra- diance fields. InNeurIPS, 2022. 1
2022
-
[19]
Evidential active recognition: Intelligent and prudent open-world embodied perception
Lei Fan, Mingfu Liang, Yunxuan Li, Gang Hua, and Ying Wu. Evidential active recognition: Intelligent and prudent open-world embodied perception. InCVPR, 2024. 2
2024
-
[20]
Navi- gation instruction generation with bev perception and large language models
Sheng Fan, Rui Liu, Wenguan Wang, and Yi Yang. Navi- gation instruction generation with bev perception and large language models. InECCV, 2024. 2
2024
-
[21]
Scene map-based prompt tuning for navigation instruction genera- tion
Sheng Fan, Rui Liu, Wenguan Wang, and Yi Yang. Scene map-based prompt tuning for navigation instruction genera- tion. InCVPR, 2025. 2
2025
-
[22]
Speaker-follower models for vision-and-language naviga- tion
Daniel Fried, Ronghang Hu, V olkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg- Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. Speaker-follower models for vision-and-language naviga- tion. InNeurIPS, 2018. 1, 2, 6
2018
-
[23]
Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation
Xiao Fu, Shangzhan Zhang, Tianrun Chen, Yichong Lu, Lanyun Zhu, Xiaowei Zhou, Andreas Geiger, and Yiyi Liao. Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation. In3DV, 2022. 1
2022
-
[24]
Dynamic view synthesis from dynamic monocular video
Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. Dynamic view synthesis from dynamic monocular video. In ICCV, 2021. 1, 3
2021
-
[25]
Room-object entity prompting and reasoning for embodied referring expression.IEEE TPAMI, 46(2):994– 1010, 2023
Chen Gao, Si Liu, Jinyu Chen, Luting Wang, Qi Wu, Bo Li, and Qi Tian. Room-object entity prompting and reasoning for embodied referring expression.IEEE TPAMI, 46(2):994– 1010, 2023. 1
2023
-
[26]
Cross-modal map learning for vision and language navigation
Georgios Georgakis, Karl Schmeckpeper, Karan Wanchoo, Soham Dan, Eleni Miltsakaki, Dan Roth, and Kostas Dani- ilidis. Cross-modal map learning for vision and language navigation. InCVPR, 2022. 1, 2
2022
-
[27]
Airbert: In-domain pretrain- ing for vision-and-language navigation
Pierre-Louis Guhur, Makarand Tapaswi, Shizhe Chen, Ivan Laptev, and Cordelia Schmid. Airbert: In-domain pretrain- ing for vision-and-language navigation. InICCV, 2021. 6
2021
-
[28]
Multi-view reconstruction via sfm-guided monocular depth estimation
Haoyu Guo, He Zhu, Sida Peng, Haotong Lin, Yunzhi Yan, Tao Xie, Wenguan Wang, Xiaowei Zhou, and Hujun Bao. Multi-view reconstruction via sfm-guided monocular depth estimation. InCVPR, 2025. 3
2025
-
[29]
Language and visual entity relationship graph for agent navigation
Yicong Hong, Cristian Rodriguez, Yuankai Qi, Qi Wu, and Stephen Gould. Language and visual entity relationship graph for agent navigation. InNeurIPS, 2020. 2, 6
2020
-
[30]
Vln bert: A recurrent vision- and-language bert for navigation
Yicong Hong, Qi Wu, Yuankai Qi, Cristian Rodriguez- Opazo, and Stephen Gould. Vln bert: A recurrent vision- and-language bert for navigation. InCVPR, 2021. 5, 6
2021
-
[31]
Learning navigational visual representations with semantic map super- vision
Yicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernon- court, Trung Bui, Stephen Gould, and Hao Tan. Learning navigational visual representations with semantic map super- vision. InCVPR, 2023. 1
2023
-
[32]
Stay on the path: Instruction fidelity in vision-and-language navigation
Vihan Jain, Gabriel Magalhaes, Alexander Ku, Ashish Vaswani, Eugene Ie, and Jason Baldridge. Stay on the path: Instruction fidelity in vision-and-language navigation. In ACL, 2019. 2, 5, 6
2019
-
[33]
Hifi4g: High-fidelity human performance rendering via compact gaussian splatting
Yuheng Jiang, Zhehao Shen, Penghao Wang, Zhuo Su, Yu Hong, Yingliang Zhang, Jingyi Yu, and Lan Xu. Hifi4g: High-fidelity human performance rendering via compact gaussian splatting. InCVPR, 2024. 3
2024
-
[34]
Deformation and correspondence aware un- supervised synthetic-to-real scene flow estimation for point clouds
Zhao Jin, Yinjie Lei, Naveed Akhtar, Haifeng Li, and Mu- nawar Hayat. Deformation and correspondence aware un- supervised synthetic-to-real scene flow estimation for point clouds. InCVPR, 2022. 2
2022
-
[35]
Splatam: Splat track & map 3d gaussians for dense rgb-d slam
Nikhil Keetha, Jay Karhade, Krishna Murthy Jatavallabhula, Gengshan Yang, Sebastian Scherer, Deva Ramanan, and Jonathon Luiten. Splatam: Splat track & map 3d gaussians for dense rgb-d slam. InCVPR, 2024. 3
2024
-
[36]
Bert: Pre-training of deep bidirectional trans- formers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. Bert: Pre-training of deep bidirectional trans- formers for language understanding. InNAACL, 2019. 5
2019
-
[37]
3d gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4), 2023
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering.ACM TOG, 42(4), 2023. 2, 3, 4
2023
-
[38]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. InICLR, 2015. 5
2015
-
[39]
Controllable navigation in- struction generation with chain of thought prompting
Xianghao Kong, Jinyu Chen, Wenguan Wang, Hang Su, Xi- aolin Hu, Yi Yang, and Si Liu. Controllable navigation in- struction generation with chain of thought prompting. In ECCV, 2024. 2
2024
-
[40]
Renderable neural radiance map for visual navigation
Obin Kwon, Jeongho Park, and Songhwai Oh. Renderable neural radiance map for visual navigation. InCVPR, 2023. 1, 2
2023
-
[41]
Envedit: Environment editing for vision-and-language navigation
Jialu Li, Hao Tan, and Mohit Bansal. Envedit: Environment editing for vision-and-language navigation. InCVPR, 2022. 2
2022
-
[42]
3d neural scene representations for visuomotor control
Yunzhu Li, Shuang Li, Vincent Sitzmann, Pulkit Agrawal, and Antonio Torralba. 3d neural scene representations for visuomotor control. InCoRL, 2022. 1
2022
-
[43]
V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion
Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao, Jose M Alvarez, Sanja Fidler, Chen Feng, and Anima Anand- kumar. V oxformer: Sparse voxel transformer for camera- based 3d semantic scene completion. InCVPR, 2023. 2
2023
-
[44]
Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching
Yixun Liang, Xin Yang, Jiantao Lin, Haodong Li, Xiao- gang Xu, and Yingcong Chen. Luciddreamer: Towards high- fidelity text-to-3d generation via interval score matching. In CVPR, 2024. 3
2024
-
[45]
Scene-intuitive agent for remote embodied visual grounding
Xiangru Lin, Guanbin Li, and Yizhou Yu. Scene-intuitive agent for remote embodied visual grounding. InCVPR,
-
[46]
Vision-language naviga- tion with random environmental mixup
Chong Liu, Fengda Zhu, Xiaojun Chang, Xiaodan Liang, Zongyuan Ge, and Yi-Dong Shen. Vision-language naviga- tion with random environmental mixup. InICCV, 2021. 2, 6
2021
-
[47]
Bird’s-eye-view scene graph for vision-language navigation
Rui Liu, Xiaohan Wang, Wenguan Wang, and Yi Yang. Bird’s-eye-view scene graph for vision-language navigation. InICCV, 2023. 1, 2, 6
2023
-
[48]
Vision-language nav- igation with energy-based policy
Rui Liu, Wenguan Wang, and Yi Yang. Vision-language nav- igation with energy-based policy. InNeurIPS, 2024. 1
2024
-
[49]
V olumetric envi- ronment representation for vision-language navigation
Rui Liu, Wenguan Wang, and Yi Yang. V olumetric envi- ronment representation for vision-language navigation. In CVPR, 2024. 1, 2, 5
2024
-
[50]
Editing conditional radiance fields
Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. Editing conditional radiance fields. InICCV, 2021. 1, 3
2021
-
[51]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. InECCV, 2020. 1, 2, 3
2020
-
[52]
Soat: A scene-and object-aware transformer for vision-and-language navigation
Abhinav Moudgil, Arjun Majumdar, Harsh Agrawal, Ste- fan Lee, and Dhruv Batra. Soat: A scene-and object-aware transformer for vision-and-language navigation. InNeurIPS,
-
[53]
Seeing the un-scene: Learning amodal semantic maps for room navigation
Medhini Narasimhan, Erik Wijmans, Xinlei Chen, Trevor Darrell, Dhruv Batra, Devi Parikh, and Amanpreet Singh. Seeing the un-scene: Learning amodal semantic maps for room navigation. InECCV, 2020. 2
2020
-
[54]
Neural map: Structured memory for deep reinforcement learning
Emilio Parisotto and Ruslan Salakhutdinov. Neural map: Structured memory for deep reinforcement learning. In ICLR, 2018. 1
2018
-
[55]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. InCVPR, 2021. 3
2021
-
[56]
Reverie: Remote embodied visual referring expres- sion in real indoor environments
Yuankai Qi, Qi Wu, Peter Anderson, Xin Wang, William Yang Wang, Chunhua Shen, and Anton van den Hengel. Reverie: Remote embodied visual referring expres- sion in real indoor environments. InCVPR, 2020. 2, 3, 5, 6, 8
2020
-
[57]
Hop: history-and-order aware pre- training for vision-and-language navigation
Yanyuan Qiao, Yuankai Qi, Yicong Hong, Zheng Yu, Peng Wang, and Qi Wu. Hop: history-and-order aware pre- training for vision-and-language navigation. InCVPR, 2022. 6
2022
-
[58]
Holis- tic lstm for pedestrian trajectory prediction.IEEE TIP, 30: 3229–3239, 2021
Ruijie Quan, Linchao Zhu, Yu Wu, and Yi Yang. Holis- tic lstm for pedestrian trajectory prediction.IEEE TIP, 30: 3229–3239, 2021. 2
2021
-
[59]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. InICML, 2021. 3, 4
2021
-
[60]
Occupancy anticipation for efficient exploration and navigation
Santhosh K Ramakrishnan, Ziad Al-Halah, and Kristen Grauman. Occupancy anticipation for efficient exploration and navigation. InECCV, 2020. 2
2020
-
[61]
Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. Sam 2: Segment anything in images and videos.arXiv preprint arXiv:2408.00714, 2024. 3, 4
2024 arXiv
-
[62]
Gordon, and Drew Bagnell
St ´ephane Ross, Geoffrey J. Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. InAISTATS, 2011. 5
2011
-
[63]
Toward open set recogni- tion.IEEE TPAMI, 35(7):1757–1772, 2012
Walter J Scheirer, Anderson de Rezende Rocha, Archana Sapkota, and Terrance E Boult. Toward open set recogni- tion.IEEE TPAMI, 35(7):1757–1772, 2012. 2
2012
-
[64]
Language embedded 3d gaussians for open- vocabulary scene understanding
Jin-Chuan Shi, Miao Wang, Hao-Bin Duan, and Shao- Hua Guan. Language embedded 3d gaussians for open- vocabulary scene understanding. InCVPR, 2024. 3
2024
-
[65]
Snerl: Semantic-aware neural radiance fields for reinforcement learning
Dongseok Shim, Seungjae Lee, and H Jin Kim. Snerl: Semantic-aware neural radiance fields for reinforcement learning. InICML, 2023. 2
2023
-
[66]
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. Sequence to sequence learning with neural networks. InNeurIPS, 2014. 1
2014
-
[67]
Learning to nav- igate unseen environments: Back translation with environ- mental dropout
Hao Tan, Licheng Yu, and Mohit Bansal. Learning to nav- igate unseen environments: Back translation with environ- mental dropout. InNAACL, 2019. 1, 2, 6
2019
-
[68]
Active visual information gathering for vision-language navigation
Hanqing Wang, Wenguan Wang, Tianmin Shu, Wei Liang, and Jianbing Shen. Active visual information gathering for vision-language navigation. InECCV, 2020. 2, 6
2020
-
[69]
Structured scene memory for vision- language navigation
Hanqing Wang, Wenguan Wang, Wei Liang, Caiming Xiong, and Jianbing Shen. Structured scene memory for vision- language navigation. InCVPR, 2021. 1, 2, 3, 6
2021
-
[70]
Towards versatile embodied navigation
Hanqing Wang, Wei Liang, Luc V Gool, and Wenguan Wang. Towards versatile embodied navigation. InNeurIPS,
-
[71]
Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation
Hanqing Wang, Wei Liang, Jianbing Shen, Luc Van Gool, and Wenguan Wang. Counterfactual cycle-consistent learn- ing for instruction following and generation in vision- language navigation. InCVPR, 2022. 6
2022
-
[72]
Dreamwalker: Mental planning for continuous vision-language navigation
Hanqing Wang, Wei Liang, Luc Van Gool, and Wenguan Wang. Dreamwalker: Mental planning for continuous vision-language navigation. InICCV, 2023. 2
2023
-
[73]
Active perception for visual-language navigation.IJCV, 131(3):607–625, 2023
Hanqing Wang, Wenguan Wang, Wei Liang, Steven CH Hoi, Jianbing Shen, and Luc Van Gool. Active perception for visual-language navigation.IJCV, 131(3):607–625, 2023. 1
2023
-
[74]
Reinforced cross-modal matching and self- supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. Reinforced cross-modal matching and self- supervised imitation learning for vision-language navigation. InCVPR, 2019. 2, 6
2019
-
[75]
Lana: A language-capable navigator for instruction follow- ing and generation
Xiaohan Wang, Wenguan Wang, Jiayi Shao, and Yi Yang. Lana: A language-capable navigator for instruction follow- ing and generation. InCVPR, 2023. 6
2023
-
[76]
Image quality assessment: from error visibility to structural similarity.IEEE TIP, 13(4):600–612, 2004
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity.IEEE TIP, 13(4):600–612, 2004. 5
2004
-
[77]
Gridmm: Grid memory map for vision-and- language navigation
Zihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu, and Shuqiang Jiang. Gridmm: Grid memory map for vision-and- language navigation. InICCV, 2023. 1, 2, 6
2023
-
[78]
Lookahead exploration with neural radiance representation for continuous vision- language navigation
Zihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu, Junjie Hu, Ming Jiang, and Shuqiang Jiang. Lookahead exploration with neural radiance representation for continuous vision- language navigation. InCVPR, 2024. 1, 2
2024
-
[79]
Vector-decomposed disentanglement for domain- invariant object detection
Aming Wu, Rui Liu, Yahong Han, Linchao Zhu, and Yi Yang. Vector-decomposed disentanglement for domain- invariant object detection. InICCV, 2021. 2
2021
-
[80]
4k4d: Real-time 4d view synthesis at 4k resolution
Zhen Xu, Sida Peng, Haotong Lin, Guangzhao He, Jiaming Sun, Yujun Shen, Hujun Bao, and Xiaowei Zhou. 4k4d: Real-time 4d view synthesis at 4k resolution. InCVPR, 2024. 3
2024
-
[81]
Gaussian grouping: Segment and edit anything in 3d scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu, and Lei Ke. Gaussian grouping: Segment and edit anything in 3d scenes. InECCV, 2024. 3
2024
-
[82]
Compositional scene representation learning via reconstruc- tion: A survey.IEEE TPAMI, 45(10):11540–11560, 2023
Jinyang Yuan, Tonglin Chen, Bin Li, and Xiangyang Xue. Compositional scene representation learning via reconstruc- tion: A survey.IEEE TPAMI, 45(10):11540–11560, 2023. 2
2023
-
[83]
Target- driven structured transformer planner for vision-language navigation
Yusheng Zhao, Jinyu Chen, Chen Gao, Wenguan Wang, Lirong Yang, Haibing Ren, Huaxia Xia, and Si Liu. Target- driven structured transformer planner for vision-language navigation. InACM MM, 2022. 6
2022
-
[84]
In-place scene labelling and understanding with implicit scene representation
Shuaifeng Zhi, Tristan Laidlow, Stefan Leutenegger, and An- drew J Davison. In-place scene labelling and understanding with implicit scene representation. InICCV, 2021. 1, 3
2021
-
[85]
Empowering embodied visual tracking with visual foundation models and offline rl
Fangwei Zhong, Kui Wu, Hai Ci, Churan Wang, and Hao Chen. Empowering embodied visual tracking with visual foundation models and offline rl. InECCV, 2024. 2
2024
-
[86]
Unrealzoo: Enriching photo- realistic virtual worlds for embodied ai
Fangwei Zhong, Kui Wu, Churan Wang, Hao Chen, Hai Ci, Zhoujun Li, and Yizhou Wang. Unrealzoo: Enriching photo- realistic virtual worlds for embodied ai. InICCV, 2025. 1
2025
-
[87]
Migc: Multi-instance generation controller for text-to-image synthesis
Dewei Zhou, You Li, Fan Ma, Xiaoting Zhang, and Yi Yang. Migc: Multi-instance generation controller for text-to-image synthesis. InCVPR, 2024. 2
2024
-
[88]
Hugs: Holistic urban 3d scene understanding via gaus- sian splatting
Hongyu Zhou, Jiahao Shao, Lu Xu, Dongfeng Bai, Weichao Qiu, Bingbing Liu, Yue Wang, Andreas Geiger, and Yiyi Liao. Hugs: Holistic urban 3d scene understanding via gaus- sian splatting. InCVPR, 2024. 3
2024
-
[89]
Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Shijie Zhou, Haoran Chang, Sicheng Jiang, Zhiwen Fan, Ze- hao Zhu, Dejia Xu, Pradyumna Chari, Suya You, Zhangyang Wang, and Achuta Kadambi. Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields. In CVPR, 2024. 3
2024
-
[90]
Vision-language navigation with self-supervised auxiliary reasoning tasks
Fengda Zhu, Yi Zhu, Xiaojun Chang, and Xiaodan Liang. Vision-language navigation with self-supervised auxiliary reasoning tasks. InCVPR, 2020. 6
2020
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.