Pith. sign in

REVIEW 2 cited by

Deep Object-Centric Representations for Generalizable Robot Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1708.04225 v3 pith:GVWVJILB submitted 2017-08-14 cs.RO cs.AIcs.CV

classification cs.ROcs.AIcs.CV
keywords objectsmanipulationattentiondemonstrationsgeneralizablelearnedlearningobject-centric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robotic manipulation in complex open-world scenarios requires both reliable physical manipulation skills and effective and generalizable perception. In this paper, we propose a method where general purpose pretrained visual models serve as an object-centric prior for the perception system of a learned policy. We devise an object-level attentional mechanism that can be used to determine relevant objects from a few trajectories or demonstrations, and then immediately incorporate those objects into a learned policy. A task-independent meta-attention locates possible objects in the scene, and a task-specific attention identifies which objects are predictive of the trajectories. The scope of the task-specific attention is easily adjusted by showing demonstrations with distractor objects or with diverse relevant objects. Our results indicate that this approach exhibits good generalization across object instances using very few samples, and can be used to learn a variety of manipulation tasks using reinforcement learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. See Less, Specify More: Visual Evidence Budgets for Generalizable VLAs

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    S2 improves generalization in vision-language-action models by using goal-preserving refined language guidance and explicit visual evidence budgets, raising mean subtask success from 54.2% to 79.0% on eight real-robot...

  2. P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies

    cs.RO 2024-12 conditional novelty 5.0 of 10

    P3-PO feeds robot policies human-prescribed semantic keypoints, propagated by correspondence and tracking, and reports strong generalization gains on real manipulation tasks.

Pith tools