Pith. sign in

REVIEW 18 cited by

OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.12202 v2 pith:43QNJ7CL submitted 2024-01-22 cs.RO cs.AIcs.CVcs.LG

OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics

classification cs.RO cs.AIcs.CVcs.LG
keywords ok-robotmodelsroboticsgraspingnavigationopenperformancesystems
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Remarkable progress has been made in recent years in the fields of vision, language, and robotics. We now have vision models capable of recognizing objects based on language queries, navigation systems that can effectively control mobile systems, and grasping models that can handle a wide range of objects. Despite these advancements, general-purpose applications of robotics still lag behind, even though they rely on these fundamental capabilities of recognition, navigation, and grasping. In this paper, we adopt a systems-first approach to develop a new Open Knowledge-based robotics framework called OK-Robot. By combining Vision-Language Models (VLMs) for object detection, navigation primitives for movement, and grasping primitives for object manipulation, OK-Robot offers a integrated solution for pick-and-drop operations without requiring any training. To evaluate its performance, we run OK-Robot in 10 real-world home environments. The results demonstrate that OK-Robot achieves a 58.5% success rate in open-ended pick-and-drop tasks, representing a new state-of-the-art in Open Vocabulary Mobile Manipulation (OVMM) with nearly 1.8x the performance of prior work. On cleaner, uncluttered environments, OK-Robot's performance increases to 82%. However, the most important insight gained from OK-Robot is the critical role of nuanced details when combining Open Knowledge systems like VLMs with robotic modules. Videos of our experiments and code are available on our website: https://ok-robot.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control

    cs.RO 2026-06 unverdicted novelty 7.0

    HALO distills VLM priors via question-answering objectives and applies sparse attention to enable reliable memory retrieval from up to eight minutes of history in imitation-learned visuomotor policies.

  2. eMEM: A Hybrid Spatio-Temporal Memory System For Embodied Agents

    cs.RO 2026-06 unverdicted novelty 7.0

    eMEM is a multi-index memory architecture with tiered consolidation and ten recall tools for embodied agents, scoring 80.8 weighted mean on eMEM-Bench covering eight cognitive psychology paradigms and outperforming a ...

  3. OmniRobotHome: A Multi-Camera Platform for Real-Time Multiadic Human-Robot Interaction

    cs.RO 2026-04 unverdicted novelty 7.0

    A 48-camera residential platform delivers real-time occlusion-robust 3D perception and coordinated actuation for multi-human multi-robot interaction in a shared home workspace.

  4. HyperDCM: Dynamic Cluster Memory Replay in Hyperbolic Space for Continual Robotic Navigation Across Scenes

    cs.RO 2026-07 reject novelty 6.0

    A scene-graph-plus-hyperbolic replay memory reduces reported performance drops in continual diffusion navigation, but the headline Drop metric is defined across different scenes and does not clearly measure forgetting.

  5. GraspGen-X: Cross-Embodiment 6-DOF Diffusion-based Grasping

    cs.RO 2026-05 unverdicted novelty 6.0

    GraspGen-X extends diffusion 6-DOF grasping to cross-embodiment via swept-volume gripper encoding, trained on procedural grippers and 2B grasps, claiming best zero-shot generalization to novel grippers in sim and real tests.

  6. $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

    cs.LG 2025-04 unverdicted novelty 6.0

    π_{0.5} is a VLA model that achieves long-horizon dexterous manipulation in entirely new homes through co-training on heterogeneous tasks and multi-source data including web and semantic predictions.

  7. Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models

    cs.RO 2025-02 unverdicted novelty 6.0

    A hierarchical VLA architecture lets robots follow complex instructions and situated feedback by separating high-level reasoning from low-level control.

  8. Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks

    cs.RO 2024-12 unverdicted novelty 6.0

    Uni-NaVid unifies diverse embodied navigation tasks into one video-based vision-language-action model trained on 3.6 million samples from four sub-tasks, achieving state-of-the-art performance on benchmarks and real-w...

  9. $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

    cs.LG 2024-10 unverdicted novelty 6.0

    π₀ is a vision-language-action flow model trained on diverse multi-platform robot data that supports zero-shot task performance, language instruction following, and efficient fine-tuning for dexterous tasks.

  10. HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory

    cs.RO 2026-06 unverdicted novelty 5.0

    HoloAgent-0 is a unified embodied agent framework with Embodied AgentOS, 3D spatial memory, and embodied skills, deployed and evaluated on real robot hardware for navigation and manipulation tasks.

  11. GeoHAT: Geometry-Adaptive Hybrid Action Transformer for Mobile Manipulation

    cs.RO 2026-06 unverdicted novelty 5.0

    GeoHAT reports a 79.3% mean success rate on the ManiSkill-HAB mobile manipulation benchmark (23.7% above the strongest baseline) by using gated geometric token injection and a hybrid whole-body action decoder.

  12. Make Your VLA More Robust Without More Data By Interleaving Motion Planning

    cs.RO 2026-05 unverdicted novelty 5.0

    MPVI interleaves model-based motion planning with VLAs via VLM completion checking to achieve 113% higher task progress on BEHAVIOR-1K without extra data.

  13. Visibility-Aware Mobile Grasping in Dynamic Environments

    cs.RO 2026-05 unverdicted novelty 5.0

    A visibility-aware mobile grasping system with iterative whole-body planning and behavior-tree subgoal generation achieves 68.8% success in unknown static and 58% in dynamic environments, outperforming a baseline by 2...

  14. Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models

    cs.RO 2024-09 unverdicted novelty 5.0

    Digital twin representations from vision foundation models enable LLM-based planning for robust peg transfer and gauze retrieval on the dVRK surgical platform with claimed generalizability.

  15. A Scalable Embodied Intelligence Platform for Seamless Real-to-Sim-to-Real Transfer of Household Mobile Manipulation Tasks

    cs.RO 2026-06 unverdicted novelty 4.0

    BestMan is a robotics platform with ASG for scene reconstruction, simulation-guided skill learning, and HUM middleware to enable seamless real-to-sim-to-real transfer in household mobile manipulation.

  16. General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling

    cs.CV 2026-05 unverdicted novelty 4.0

    GAM framework uses arc-length parameterization for temporal invariance and schema-affine factorization for geometric invariance to build a covariant action manifold integrated into VLA models for improved generalizati...

  17. Visibility-Aware Mobile Grasping in Dynamic Environments

    cs.RO 2026-05 unverdicted novelty 4.0

    A unified visibility-aware mobile grasping system using whole-body planning, active perception, and behavior trees improves success rates in unknown static and dynamic environments.

  18. Open-Architecture End-to-End System for Real-World Autonomous Robot Navigation

    cs.RO 2024-10 unverdicted novelty 4.0

    Presents an open ROS2-based end-to-end navigation system for quadruped robots achieving over 88% success in zero-shot real-world indoor navigation tasks using semantic scene graphs and LLM planning.