Pith. sign in

Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Generic re-usable pre-trained image representation encoders have become a standard component of methods for many computer vision tasks. As visual representations for robots however, their utility has been limited, leading to a recent wave of efforts to pre-train robotics-specific image encoders that are better suited to robotic tasks than their generic counterparts. We propose Scene Objects From Transformers, abbreviated as SOFT, a wrapper around pre-trained vision transformer (PVT) models that bridges this gap without any further training. Rather than construct representations out of only the final layer activations, SOFT individuates and locates object-like entities from PVT attentions, and describes them with PVT activations, producing an object-centric embedding. Across standard choices of generic pre-trained vision transformers PVT, we demonstrate in each case that policies trained on SOFT(PVT) far outstrip standard PVT representations for manipulation tasks in simulated and real settings, approaching the state-of-the-art robotics-aware representations. Code, appendix and videos: https://sites.google.com/view/robot-soft/

fields

cs.AI 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Is an object-centric representation beneficial for robotic manipulation ?

cs.AI · 2025-06-24 · reject · novelty 4.0

Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.

citing papers explorer

Showing 1 of 1 citing paper.

  • Is an object-centric representation beneficial for robotic manipulation ? cs.AI · 2025-06-24 · reject · none · ref 28 · internal anchor

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.