REVIEW 5 major objections 5 minor 1 cited by
RoboRetriever: Single-Camera Robot Object Retrieval via Active and Interactive Perception with Dynamic Scene Graph
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read RoboRetriever claims that a robot with only one wrist-mounted RGB-D camera can retrieve hidden, occluded, or semantically specified objects by actively choosing viewpoints and physically interacting with the scene, reporting 70–90%…
desk verdict A real single-camera retrieval system with a clever prompting scheme, but the headline numbers rest on undisclosed trial counts and the closed-model dependency makes the evaluation hard to trust as reported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the dynamic hierarchical scene graph, defined as $G=(V,E)$, where each node carries a cropped image history, an accumulated point cloud, semantic attributes (fine-grained name, movable/static flag, confidence, occluded flag, partial-view flag, and a free-form description), and edges encode five relations: behind, belong, inside, on, under. Relations pointing into unexplored regions trigger 'Unknown' nodes that hypothesize objects behind occluders or inside containers, which is what lets the supervisor formulate exploration goals. The action-driving mechanism is the visual prompting scheme: the system samples candidate camera directions on a virtual sphere centered on the target object, renders canonical front/left/right views of the point cloud, and has GPT-o3 choose the best direction and then the best pose, keeping the camera aimed at the object center; this turns pose selection into a grounded multiple-choice decision rather than free-form coordinate generation.
What would settle it
A concrete test: replicate the Hidden Inside and Compositional Reasoning tasks while logging every GPT-o3 decision, and count (a) camera poses that collide or point away from the scene, (b) object merges that conflate distinct instances, and (c) rollouts that need human approval or correction; if any of these occur in more than a small fraction of trials, the claimed single-camera autonomy is not established.
Extended reading notes
Core claim
The central claim is that active and interactive perception can be unified under a single-camera constraint: instead of assuming a fixed or multi-camera setup with full scene visibility, the robot builds a dynamic hierarchical scene graph from its wrist-camera observations, continuously updates that graph as it moves and interacts, and lets a reasoning vision-language model decide both which object to pursue and which action to take next. The paper's reported results show that this integration succeeds on tasks where each ingredient alone fails: active perception alone cannot open a closed drawer, interactive perception alone cannot choose where to look, and a VLM without grounded scene memory hallucinates camera poses. The framework is presented as class-agnostic and task-agnostic, adapting to new instructions and environments without task-specific priors.
Load-bearing premise
The whole pipeline trusts closed-source vision-language models GPT-o3 and GPT-4o to make every semantic decision—matching objects across views, updating attributes, inferring relations, adding unknown nodes, choosing what to do next, and picking the next camera pose—so if those models hallucinate, drift, or become unavailable, the robot's autonomy collapses.
Editorial extensions
If this is right
- Retrieval no longer requires full scene visibility: a single moving camera plus physical interaction can expose objects that no fixed view captures.
- New instructions do not require changing perception code: the same supervisor prompt consumes the current scene graph and produces a plan for any object named in natural language.
- Memory of past actions and objects pays off in sequential tasks: later instructions can reuse the accumulated graph instead of re-exploring the scene.
- The reported gap over the GPT-o3 baseline suggests that grounding viewpoint choice in rendered point-cloud views substantially reduces hallucinated camera poses.
- The combination of active and interactive perception handles occlusion modes that either alone cannot: the robot can look behind, under, and inside, and can move obstructions when vision alone is insufficient.
Reading between the lines
- Inference: because pose selection is a discrimination task over sampled candidates rather than generative coordinate regression, the scheme may work with any VLM strong enough to compare rendered views; swapping GPT-o3 for an open-weight model would test this directly.
- Inference: the 'Unknown' nodes function as explicit exploration frontiers, so graph-edit distance may serve as a measure of information gain; a next-action policy could optimize expected graph change instead of relying on the VLM's free choice.
- Inference: the supervisor's decision loop is not tied to tabletop arms; the same scene-graph-plus-prompt pattern could drive mobile manipulators or dual-arm systems, where each camera move is still a 6-DoF viewpoint selection.
- Inference: the reported numbers average over successes; a stricter evaluation would separate perception accuracy from planning accuracy by replaying logged VLM decisions against ground-truth scene states, which the paper does not report.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RoboRetriever, a framework for object retrieval with a single wrist-mounted RGB-D camera and free-form natural-language instructions. The system builds and incrementally updates a dynamic hierarchical scene graph through a grounding module (GPT-4o, DINO-X, SAM), a memory module that includes hypothetical Unknown nodes, a supervisor module based on GPT-o3, and an action module that coordinates active perception, interactive perception, and manipulation. Active perception is driven by a novel visual prompting scheme in which GPT-o3 selects among sampled camera directions and poses. The method is evaluated on six real-world tabletop task categories against RoboEXP, AP-VLM, and a GPT-o3 baseline, reporting success rates of 70--90% and object discovery rates up to 95%, along with ablations and a supplementary video of human-intervention scenarios.
Significance. If the reported results hold, RoboRetriever would be a meaningful advance for single-camera, language-driven object retrieval in partially observable scenes, and the integration of active and interactive perception within one modular framework is genuinely interesting. The dynamic scene graph with Unknown nodes and the task-aware visual prompting scheme for 6-DoF camera control are plausible contributions. However, the current evidence is insufficient to support the strong generalization and robustness claims: the central quantitative results lack trial counts and uncertainty measures, the closest language-grounded baselines are absent, and the heavy reliance on undisclosed closed-model prompts makes independent verification impossible. The paper is promising but needs substantial empirical and transparency revisions.
major comments (5)
- [Experiments (Table 2)] The central generalization claim rests on Table 2, but the manuscript never states the number of rollouts per variation or per task category. The setup says each category has three variations, so if only one rollout was executed per variation, the reported 90% success rate could correspond to 3/3 successes with a binomial 95% confidence interval of roughly 29--100%, and the 70% Compositional Reasoning result could be 2/3 with an even wider interval. Please report the exact number of trials per variation and per category, the per-variation outcomes, and confidence intervals, and if the trial counts are this small, temper the headline success-rate claims accordingly.
- [Experiment Setup (Baselines)] The evaluation omits the closest language-grounded interactive search baselines, such as the language-grounded dynamic scene graph approach of Honerkamp et al. and CuriousBot, both of which are cited in the paper. In addition, the GPT-o3 baseline is asymmetric: it receives ground-truth interactive perception and manipulation actions and is not given the memory module, which makes the large performance margin less informative. Please add at least one recent language-grounded interactive search baseline, or explicitly justify why those methods are not comparable, and equalize the information available to baselines where possible.
- [Action module / Grounding module] The entire pipeline depends on closed commercial models (GPT-4o and GPT-o3) for instance matching, semantic attribute updates, relation inference, Unknown-node placement, supervisor decisions, and camera pose selection, but the paper provides no exact prompts, API versions, sampling parameters, or temperature settings, and it does not audit VLM failures. Because no code or data are released, these omissions prevent independent reproduction and make the results vulnerable to model-version drift. Please include the full prompts and model configuration in an appendix, and add a failure analysis with representative examples of VLM errors and their effect on task outcomes.
- [Action module (Active perception)] The claim that the method adapts 'without task-specific priors' is strained by several unreported tunable parameters in the active perception module: the sphere radius and distance factor, the number of sampled candidate directions N, and the number of sampled candidate poses M. None of these values is given, and no sensitivity analysis is provided. Please report these parameters and show that the results do not depend critically on their choice.
- [Ablation Study (Human intervention)] The claimed robustness under human interventions is supported only by a statement that 'more details can be found in the supplementary video,' with no quantitative success rate, trial count, or failure description in the paper. Similarly, the ablation results in Figure 6D and 6E are presented without error bars or explicit numbers of rollouts. Please provide quantitative results with trial counts for the ablation conditions and for human-intervention trials.
minor comments (5)
- [Comparison with baselines] There is a typo in the phrase 'generative rather than discrimitive approach'; it should be 'discriminative'.
- [Table 2] The table uses '/' for GED entries of AP-VLM and GPT-o3; please explain whether GED was not computed for these baselines and clarify the exact computation of the ground-truth and predicted scene graphs.
- [Figure 5] The visual legend for the scene graph is difficult to follow because omitted nodes are indicated with ellipses and several node types overlap; please enlarge the figure and provide a clearer key.
- [Experiment Setup] The paper should state explicitly that all trials share the same tabletop, robot arm, camera, and gripper, and discuss the resulting limitations for generalization to other embodiments and environments.
- [Metrics] The Object Discovery Rate is defined as 'the percentage of discovered objects out of all objects in the environment,' but the denominator is ambiguous for objects that are never visible from any viewpoint; please specify how the ground-truth object inventory was obtained.
Circularity Check
No significant circularity: RoboRetriever's claims are supported by real-robot experiments, not by fitted inputs or self-cited theorems.
full rationale
The paper makes no formal derivation from inputs to outputs; its central claims are supported by physical rollouts in author-designed task categories. No parameter is fitted to a subset of data and then renamed as a prediction: the GPT-o3/GPT-4o modules are used as fixed black-box components, and the metrics (success rate, ODR, GED) are measured against ground-truth environment objects, designated placements, and ground-truth graphs. The GPT-o3 baseline is an honest external comparison, and the system's advantage is attributed to memory, grounding, and scene-graph scaffolding rather than to the VLM itself. The only same-author citation (Wang et al. 2023) appears in the Related Work as background and is not load-bearing for any claim. No equation reduces to its own input, no cited uniqueness theorem is imported, and no known result is merely renamed. Therefore the derivation chain is self-contained and the empirical evaluation is not circular. Concerns about undisclosed rollout counts or closed-source model reliability are evidentiary or robustness issues, not circularity.
Assumptions & free parameters
free parameters (3)
- Active perception sphere radius and distance factor =
Not reported
- Number of sampled candidate directions N =
Not reported
- Number of sampled candidate poses M =
Not reported
assumptions (5)
- domain assumption GPT-o3 and GPT-4o outputs (semantic labels, instance matching, relations, action decisions, camera pose choices) are sufficiently reliable for autonomous operation.
- domain assumption DINO-X, SAM, and ZeroMatch provide accurate open-vocabulary detection, segmentation, and point cloud registration.
- domain assumption Hand-eye calibration and robot pose reports are accurate enough for point cloud transformation into the base frame.
- domain assumption The implemented action primitives (Open, Close, Pick&Place, Rotate) succeed often enough that the perception loop advances.
- domain assumption The author-designed task categories and environments are representative of general object retrieval scenarios.
invented entities (1)
-
Unknown node in the dynamic hierarchical scene graph
Cite this review
Pith. "Pith review of RoboRetriever: Single-Camera Robot Object Retrieval via Active and Interactive Perception with Dynamic Scene Graph." pith.science (2026). https://pith.science/paper/RDFOKN2V
@misc{pith2026250812916,
author = {Pith},
title = {Pith review of: RoboRetriever: Single-Camera Robot Object Retrieval via Active and Interactive Perception with Dynamic Scene Graph},
year = {2026},
howpublished = {\url{https://pith.science/paper/RDFOKN2V}},
note = {Machine review of arXiv:2508.12916}
}
read the original abstract
Humans effortlessly retrieve objects in cluttered, partially observable environments by combining visual reasoning, active viewpoint adjustment, and physical interaction-with only a single pair of eyes. In contrast, most existing robotic systems rely on carefully positioned fixed or multi-camera setups with complete scene visibility, which limits adaptability and incurs high hardware costs. We present \textbf{RoboRetriever}, a novel framework for real-world object retrieval that operates using only a \textbf{single} wrist-mounted RGB-D camera and free-form natural language instructions. RoboRetriever grounds visual observations to build and update a \textbf{dynamic hierarchical scene graph} that encodes object semantics, geometry, and inter-object relations over time. The supervisor module reasons over this memory and task instruction to infer the target object and coordinate an integrated action module combining \textbf{active perception}, \textbf{interactive perception}, and \textbf{manipulation}. To enable task-aware scene-grounded active perception, we introduce a novel visual prompting scheme that leverages large reasoning vision-language models to determine 6-DoF camera poses aligned with the semantic task goal and geometry scene context. We evaluate RoboRetriever on diverse real-world object retrieval tasks, including scenarios with human intervention, demonstrating strong adaptability and robustness in cluttered scenes with only one RGB-D camera.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
FUSE: Active Functional Affordance Grounding through Adaptive Semantic-Geometric Evidence Acquisition
FUSE, an entropy-gated planner that combines amortized viewpoint prediction with explicit semantic-geometric exploration, achieves the best non-oracle active functional grounding results on a new Habitat benchmark whi...
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
An, B.; Geng, Y.; Chen, K.; Li, X.; Dou, Q.; and Dong, H. 2024. RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose Estimation. In IEEE International Conference on Robotics and Automation, ICRA 2024, Yokohama, Japan, May 13-17, 2024 , 7748--7755. IEEE
work page 2024
-
[4]
P.; Yang, Y.; Siva, R.; Milan, D.; Topcu, U.; and Wang, Z
Bhatt, N. P.; Yang, Y.; Siva, R.; Milan, D.; Topcu, U.; and Wang, Z. 2024. Know Where You're Uncertain When Planning with Multimodal Foundation Models: A Formal Framework. CoRR, abs/2411.01639
arXiv 2024
-
[5]
Buchanan, R.; R \" o fer, A.; Moura, J.; Valada, A.; and Vijayakumar, S. 2024. Online Estimation of Articulated Objects with Factor Graphs using Vision and Proprioceptive Sensing. In IEEE International Conference on Robotics and Automation, ICRA 2024, Yokohama, Japan, May 13-17, 2024 , 16111--16117. IEEE
work page 2024
-
[6]
Chen, L.; Song, Y.; Bao, H.; and Zhou, X. 2023. Perceiving Unseen 3D Objects by Poking the Objects. In IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023 , 4834--4841. IEEE
work page 2023
-
[7]
Dai, Q.; Zhu, Y.; Geng, Y.; Ruan, C.; Zhang, J.; and Wang, H. 2023. GraspNeRF: Multiview-based 6-DoF Grasp Detection for Transparent and Specular Objects Using Generalizable NeRF. In IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023 , 1757--1763. IEEE
work page 2023
-
[8]
Dass, S.; Hu, J.; Abbatematteo, B.; Stone, P.; and Mart \' n - Mart \' n, R. 2024. Learning to Look: Seeking Information for Decision Making via Policy Factorization. In Agrawal, P.; Kroemer, O.; and Burgard, W., eds., Conference on Robot Learning, 6-9 November 2024, Munich, Germany, volume 270 of Proceedings of Machine Learning Research, 4425--4445. PMLR
work page 2024
Show all 39 references
-
[9]
Dengler, N.; M \ A z cke, J.; Menon, R.; and Bennewitz, M. 2025. Efficient Manipulation-Enhanced Semantic Mapping With Uncertainty-Informed Action Selection. arXiv preprint arXiv:2506.02286
2025 arXiv
-
[10]
Ding, W.; Majcherczyk, N.; Deshpande, M.; Qi, X.; Zhao, D.; Madhivanan, R.; and Sen, A. 2023. Learning to View: Decision Transformers for Active Object Detection. In IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023 , 7140--...
2023
-
[11]
Fang, H.; Wang, C.; Fang, H.; Gou, M.; Liu, J.; Yan, H.; Liu, W.; Xie, Y.; and Lu, C. 2023. AnyGrasp: Robust and Efficient Grasp Perception in Spatial and Temporal Domains. IEEE Trans. Robotics , 39(5): 3929--3945
2023
-
[12]
Y.; Wortsman, M.; Ilharco, G.; Schmidt, L.; and Song, S
Gadre, S. Y.; Wortsman, M.; Ilharco, G.; Schmidt, L.; and Song, S. 2023. CoWs on Pasture: Baselines and Benchmarks for Language-Driven Zero-Shot Object Navigation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, ...
2023
-
[13]
Geng, H.; Xu, H.; Zhao, C.; Xu, C.; Yi, L.; Huang, S.; and Wang, H. 2023. Gapartnet: Cross-category domain-generalizable object perception and manipulation via generalizable and actionable parts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2023
-
[14]
K.; Modayil, J.; Mirowski, P
Grimes, M. K.; Modayil, J.; Mirowski, P. W.; Rao, D.; and Hadsell, R. 2023. Learning to Look by Self-Prediction. Trans. Mach. Learn. Res., 2023
2023
-
[15]
Honerkamp, D.; B \"u chner, M.; Despinoy, F.; Welschehold, T.; and Valada, A. 2024. Language-grounded dynamic scene graphs for interactive object search with mobile manipulation. IEEE Robotics and Automation Letters
2024
-
[16]
Hsu, C.; Jiang, Z.; and Zhu, Y. 2023. Ditto in the House: Building Articulation Models of Indoor Scenes through Interactive Perception. In IEEE International Conference on Robotics and Automation, ICRA 2023, London, UK, May 29 - June 2, 2023 , 3933--3939. IEEE
2023
-
[17]
Huang, W.; Wang, C.; Li, Y.; Zhang, R.; and Fei - Fei, L. 2024. ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation. In Agrawal, P.; Kroemer, O.; and Burgard, W., eds., Conference on Robot Learning, 6-9 November 2024, Munich, Germany, v...
2024
-
[18]
Jiang, H.; Huang, B.; Wu, R.; Li, Z.; Garg, S.; Nayyeri, H.; Wang, S.; and Li, Y. 2024. RoboEXP: Action-Conditioned Scene Graph via Interactive Exploration for Robotic Manipulation. In Agrawal, P.; Kroemer, O.; and Burgard, W., eds., Conference on Robot Learning, 6-9 November ...
2024
-
[19]
Jiang, H.; Xie, J.; Yang, J.; Yu, L.; and Zheng, J. 2025. Zero-shot RGB-D Point Cloud Registration with Pre-trained Large Vision Model. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA, June 11-15, 2025 , 16943--16952. Computer V...
2025
-
[20]
Jin, L.; Chen, X.; R \" u ckin, J.; and Popovic, M. 2023. NeU-NBV: Next Best View Planning Using Uncertainty Estimation in Image-Based Neural Rendering. In IROS , 11305--11312
2023
-
[21]
M.; Yi, B.; Bonnen, T.; Goldberg, K.; and Kanazawa, A
Kerr, J.; Hari, K.; Weber, E.; Kim, C. M.; Yi, B.; Bonnen, T.; Goldberg, K.; and Kanazawa, A. 2025. Eye, Robot: Learning to Look to Act with a BC-RL Perception-Action Loop. arXiv preprint arXiv:2506.10968
2025
-
[22]
C.; Lo, W.-Y.; Doll \'a r, P.; and Girshick, R
Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A. C.; Lo, W.-Y.; Doll \'a r, P.; and Girshick, R. 2023. Segment Anything. arXiv:2304.02643
2023 arXiv
-
[23]
Krátký, V.; Silano, G.; Vrba, M.; Papaioannidis, C.; Mademlis, I.; Pěnička, R.; Pitas, I.; and Saska, M. 2025. Gesture-Controlled Aerial Robot Formation for Human-Swarm Interaction in Safety Monitoring Applications. IEEE Robotics and Automation Letters, 10(8): 8244--8251
2025
-
[24]
Leusmann, J.; Belardinelli, A.; Haliburton, L.; Hasler, S.; Schmidt, A.; Mayer, S.; Gienger, M.; and Wang, C. 2025. Investigating LLM-Driven Curiosity in Human-Robot Interaction. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1--16
2025
-
[25]
K.; Nowe, A.; and Vanderborght, B
Liu, G.; De Winter, J.; Steckelmacher, D.; Hota, R. K.; Nowe, A.; and Vanderborght, B. 2023. Synergistic task and motion planning with reinforcement learning-based non-prehensile actions. IEEE Robotics and Automation Letters, 8(5): 2764--2771
2023
-
[26]
Ma, H.; Shi, M.; Gao, B.; and Huang, D. 2024. Active Perception for Grasp Detection via Neural Graspness Field. In Globersons, A.; Mackey, L.; Belgrave, D.; Fan, A.; Paquet, U.; Tomczak, J. M.; and Zhang, C., eds., Advances in Neural Information Processing Systems 38: Annual C...
2024
-
[27]
Y.; Ehsani, K.; and Song, S
Nie, N.; Gadre, S. Y.; Ehsani, K.; and Song, S. 2023. Structure from action: Learning interactions for 3d articulated object structure discovery. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 1222--1229. IEEE
2023
-
[28]
Pan, M.; Zhang, J.; Wu, T.; Zhao, Y.; Gao, W.; and Dong, H. 2025. OmniManip: Towards General Robotic Manipulation via Object-Centric Interaction Primitives as Spatial Constraints. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, TN, USA,...
2025
-
[29]
Ren, T.; Chen, Y.; Jiang, Q.; Zeng, Z.; Xiong, Y.; Liu, W.; Ma, Z.; Shen, J.; Gao, Y.; Jiang, X.; Chen, X.; Song, Z.; Zhang, Y.; Huang, H.; Gao, H.; Liu, S.; Zhang, H.; Li, F.; Yu, K.; and Zhang, L. 2024. DINO-X: A Unified Vision Model for Open-World Object Detection and Under...
2024 arXiv
-
[30]
Shang, J.; and Ryoo, M. S. 2023. Active Vision Reinforcement Learning under Limited Visual Observability. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information Processing Systems 36: Annual Conference on Neural Infor...
2023
-
[31]
Shi, Y.; Wen, D.; Chen, G.; Welte, E.; Liu, S.; Peng, K.; Stiefelhagen, R.; and Rayyes, R. 2025. VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility. arXiv preprint arXiv:2503.12609
2025 arXiv
-
[32]
Sripada, V.; Carter, S.; Guerin, F.; and Ghalamzan, A. 2024. AP-VLM: Active Perception Enabled by Vision-Language Models. CoRR, abs/2409.17641
2024 arXiv
-
[33]
Uppal, S.; Agarwal, A.; Xiong, H.; Shaw, K.; and Pathak, D. 2024. SPIN: Simultaneous Perception, Interaction and Navigation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , 18133--18142. IEEE
2024
-
[34]
Wang, G.; Li, H.; Zhang, S.; Guo, D.; Liu, Y.; and Liu, H. 2025 a . Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation. IEEE Robotics Autom. Lett. , 10(4): 3422--3429
2025
-
[35]
Wang, H.; Qi, L.; Fang, B.; and Sun, Y. 2023. Hierarchical visual policy learning for long-horizon robot manipulation in densely cluttered scenes. arXiv preprint arXiv:2312.02697
2023 arXiv
-
[36]
Wang, Y.; Fermoselle, L.; Kelestemur, T.; Wang, J.; and Li, Y. 2025 b . CuriousBot: Interactive Mobile Exploration via Actionable 3D Relational Object Graph. arXiv preprint arXiv:2501.13338
2025
-
[37]
Werby, A.; Huang, C.; B \" u chner, M.; Valada, A.; and Burgard, W. 2024. Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation. In Kulic, D.; Venture, G.; Bekris, K. E.; and Coronado, E., eds., Robotics: Science and Systems XX, Delft, The Netherl...
2024
-
[38]
Yan, Z.; Li, S.; Wang, Z.; Wu, L.; Wang, H.; Zhu, J.; Chen, L.; and Liu, J. 2025. Dynamic Open-Vocabulary 3D Scene Graphs for Long-Term Language-Guided Mobile Manipulation. IEEE Robotics Autom. Lett. , 10(5): 4252--4259
2025
-
[39]
Zhang, X.; Wang, D.; Han, S.; Li, W.; Zhao, B.; Wang, Z.; Duan, X.; Fang, C.; Li, X.; and He, J. 2023. Affordance-Driven Next-Best-View Planning for Robotic Grasping. In Tan, J.; Toussaint, M.; and Darvish, K., eds., Conference on Robot Learning, CoRL 2023, 6-9 November 2023, ...
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.