Pith. sign in

REVIEW 14 cited by

Conditional Object-Centric Learning from Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.12594 v2 pith:K6TWWFUE submitted 2021-11-24 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords objectsdatamodelmodelsobject-centricrealisticvideobiases
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models with object-centric inductive biases can learn to segment and represent meaningful objects from the statistical structure of the data alone without the need for any supervision. However, such fully-unsupervised methods still fail to scale to diverse realistic data, despite the use of increasingly complex inductive biases such as priors for the size of objects or the 3D geometry of the scene. In this paper, we instead take a weakly-supervised approach and focus on how 1) using the temporal dynamics of video data in the form of optical flow and 2) conditioning the model on simple object location cues can be used to enable segmenting and tracking objects in significantly more realistic synthetic data. We introduce a sequential extension to Slot Attention which we train to predict optical flow for realistic looking synthetic scenes and show that conditioning the initial state of this model on a small set of hints, such as center of mass of objects in the first frame, is sufficient to significantly improve instance segmentation. These benefits generalize beyond the training distribution to novel objects, novel backgrounds, and to longer video sequences. We also find that such initial-state-conditioning can be used during inference as a flexible interface to query the model for specific objects or parts of objects, which could pave the way for a range of weakly-supervised approaches and allow more effective interaction with trained models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.

  2. Factored Latent Action World Models

    cs.LG 2026-02 conditional novelty 6.0 of 10

    FLAM splits a scene into separate factors, each with its own latent action, and reports better video prediction and downstream policy learning than monolithic latent-action models.

  3. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  4. Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SlotSAR fuses wavelet scattering features with a SAR foundation model's semantic features to make slot attention separate targets from clutter in SAR images, improving segmentation metrics on ATRNet-STAR.

  5. Discovering and using Spelke segments

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SpelkeNet, a self-supervised video world model, discovers Spelke segments in static images by aggregating motion correlations across imagined pokes.

  6. Dyn-O: Building Structured World Models with Object-Centric Representations

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Dyn-O learns object-centric world models directly from pixels in complex Procgen games, using SAM2-guided slot attention and Mamba state-space dynamics, and reports better rollout prediction than DreamerV3.

  7. Identifiable Object Representations under Spatial Ambiguities

    cs.LG 2025-06 reject novelty 6.0 of 10

    VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.

  8. Object-Centric Representations Improve Policy Generalization in Robot Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Slot-based object-centric representations, especially a video model pretrained on robot data, improve policy generalization under visual distribution shifts in simulated and real-world manipulation tasks.

  9. Dreamweaver: Learning Compositional World Models from Pixels

    cs.CV 2025-01 conditional novelty 6.0 of 10

    An unsupervised recurrent block-slot model that discovers static and dynamic concept blocks from raw video and recombines them to imagine novel future videos.

  10. Leveraging Color Channel Independence for Improved Unsupervised Object Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Adding the HSV saturation channel to the RGB reconstruction target improves Slot Attention object discovery and disentanglement across several multi-object datasets.

  11. ES-Merging: Biological MLLM Merging via Embedding Space Signals

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    ES-Merging estimates layer-wise and element-wise merge coefficients from coarse- and fine-grained embedding signals and claims better cross-modal reasoning and single-modal knowledge preservation than parameter-space merging.

  12. Efficient Object-centric Representation Learning with Pre-trained Geometric Prior

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Using CroCo's geometric features as both the encoder and reconstruction target improves object discovery in slot-attention video models, with a one-pass attentional decoder.

  13. Is an object-centric representation beneficial for robotic manipulation ?

    cs.AI 2025-06 reject novelty 4.0 of 10

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...

  14. On the Benefits of Instance Decomposition in Video Prediction Models

    cs.CV 2025-01 reject novelty 4.0 of 10

    Explicit instance decomposition with per-class shared weights improves latent-transformer video prediction in the paper's experiments, but the claimed advantage is weakened by mismatched parameter counts and test-set ...

Pith tools