Pith. sign in

Innovative Integration of Visual Foundation Model with a Robotic Arm on a Mobile Platform

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

In the rapidly advancing field of robotics, the fusion of state-of-the-art visual technologies with mobile robotic arms has emerged as a critical integration. This paper introduces a novel system that combines the Segment Anything model (SAM) -- a transformer-based visual foundation model -- with a robotic arm on a mobile platform. The design of integrating a depth camera on the robotic arm's end-effector ensures continuous object tracking, significantly mitigating environmental uncertainties. By deploying on a mobile platform, our grasping system has an enhanced mobility, playing a key role in dynamic environments where adaptability are critical. This synthesis enables dynamic object segmentation, tracking, and grasping. It also elevates user interaction, allowing the robot to intuitively respond to various modalities such as clicks, drawings, or voice commands, beyond traditional robotic systems. Empirical assessments in both simulated and real-world demonstrate the system's capabilities. This configuration opens avenues for wide-ranging applications, from industrial settings, agriculture, and household tasks, to specialized assignments and beyond.

citation-role summary

extension 1

citation-polarity summary

fields

cs.RO 1

years

2026 1

verdicts

CONDITIONAL 1

roles

extension 1

polarities

extend 1

representative citing papers

Kitchen Robotic Manipulation utilizing Foundation Models

cs.RO · 2026-08-04 · conditional · novelty 4.0

A modular perception pipeline using off-the-shelf foundation models achieves 89.12% ADI on a custom kitchen dishware dataset and performs real robot sink-to-dishwasher and cup-stacking tasks without retraining.

citing papers explorer

Showing 1 of 1 citing paper.

  • Kitchen Robotic Manipulation utilizing Foundation Models cs.RO · 2026-08-04 · conditional · none · ref 26 · internal anchor

    A modular perception pipeline using off-the-shelf foundation models achieves 89.12% ADI on a custom kitchen dishware dataset and performs real robot sink-to-dishwasher and cup-stacking tasks without retraining.