Pith. sign in

Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Unsupervised video-based object-centric learning is a promising avenue to learn structured representations from large, unlabeled video collections, but previous approaches have only managed to scale to real-world datasets in restricted domains. Recently, it was shown that the reconstruction of pre-trained self-supervised features leads to object-centric representations on unconstrained real-world image datasets. Building on this approach, we propose a novel way to use such pre-trained features in the form of a temporal feature similarity loss. This loss encodes semantic and temporal correlations between image patches and is a natural way to introduce a motion bias for object discovery. We demonstrate that this loss leads to state-of-the-art performance on the challenging synthetic MOVi datasets. When used in combination with the feature reconstruction loss, our model is the first object-centric video model that scales to unconstrained video datasets such as YouTube-VIS.

fields

cs.AI 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Is an object-centric representation beneficial for robotic manipulation ?

cs.AI · 2025-06-24 · reject · novelty 4.0

Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.

citing papers explorer

Showing 1 of 1 citing paper.

  • Is an object-centric representation beneficial for robotic manipulation ? cs.AI · 2025-06-24 · reject · none · ref 41 · internal anchor

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.