Pith. sign in

Is an Object-Centric Video Representation Beneficial for Transfer?

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

The objective of this work is to learn an object-centric video representation, with the aim of improving transferability to novel tasks, i.e., tasks different from the pre-training task of action classification. To this end, we introduce a new object-centric video recognition model based on a transformer architecture. The model learns a set of object-centric summary vectors for the video, and uses these vectors to fuse the visual and spatio-temporal trajectory 'modalities' of the video clip. We also introduce a novel trajectory contrast loss to further enhance objectness in these summary vectors. With experiments on four datasets -- SomethingSomething-V2, SomethingElse, Action Genome and EpicKitchens -- we show that the object-centric model outperforms prior video representations (both object-agnostic and object-aware), when: (1) classifying actions on unseen objects and unseen environments; (2) low-shot learning of novel classes; (3) linear probe to other downstream tasks; as well as (4) for standard action classification.

fields

cs.AI 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Is an object-centric representation beneficial for robotic manipulation ?

cs.AI · 2025-06-24 · reject · novelty 4.0

Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.

citing papers explorer

Showing 1 of 1 citing paper.

  • Is an object-centric representation beneficial for robotic manipulation ? cs.AI · 2025-06-24 · reject · none · ref 42 · internal anchor

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to unseen distractors, although the comparison uses unequal pre-training protocols.