Pith. sign in

REVIEW 3 cited by

DINOBot: Robot Manipulation via Retrieval and Alignment with Vision Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13181 v1 pith:SYOV6MK3 submitted 2024-02-20 cs.RO cs.LG

classification cs.ROcs.LG
keywords dinobotobjectnovelvisionfeaturesfoundationimage-levellearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose DINOBot, a novel imitation learning framework for robot manipulation, which leverages the image-level and pixel-level capabilities of features extracted from Vision Transformers trained with DINO. When interacting with a novel object, DINOBot first uses these features to retrieve the most visually similar object experienced during human demonstrations, and then uses this object to align its end-effector with the novel object to enable effective interaction. Through a series of real-world experiments on everyday tasks, we show that exploiting both the image-level and pixel-level properties of vision foundation models enables unprecedented learning efficiency and generalisation. Videos and code are available at https://www.robot-learning.uk/dinobot.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Robots Acquire Manipulation Skills in Seconds from a Single Human Video

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single human video is enough for a robot to acquire a new manipulation skill at inference time, with no parameter updates, reaching 62% success on 50 novel tasks.

  2. MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

    cs.RO 2025-08 conditional novelty 6.0 of 10

    MimicFunc transfers tool-using skills from one human video to novel tools by aligning function-centric keypoint frames, achieving 79.5% success across five manipulation tasks.

  3. CodeDiffuser: Attention-Enhanced Diffusion Policy via VLM-Generated Code for Instruction Ambiguity

    cs.RO 2025-06

Pith tools