Pith. sign in

REVIEW 13 cited by

Conditional Object-Centric Learning from Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.12594 v2 pith:K6TWWFUE submitted 2021-11-24 cs.CV cs.LGstat.ML

Conditional Object-Centric Learning from Video

classification cs.CV cs.LGstat.ML
keywords objectsdatamodelmodelsobject-centricrealisticvideobiases
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Object-centric representations are a promising path toward more systematic generalization by providing flexible abstractions upon which compositional world models can be built. Recent work on simple 2D and 3D datasets has shown that models with object-centric inductive biases can learn to segment and represent meaningful objects from the statistical structure of the data alone without the need for any supervision. However, such fully-unsupervised methods still fail to scale to diverse realistic data, despite the use of increasingly complex inductive biases such as priors for the size of objects or the 3D geometry of the scene. In this paper, we instead take a weakly-supervised approach and focus on how 1) using the temporal dynamics of video data in the form of optical flow and 2) conditioning the model on simple object location cues can be used to enable segmenting and tracking objects in significantly more realistic synthetic data. We introduce a sequential extension to Slot Attention which we train to predict optical flow for realistic looking synthetic scenes and show that conditioning the initial state of this model on a small set of hints, such as center of mass of objects in the first frame, is sufficient to significantly improve instance segmentation. These benefits generalize beyond the training distribution to novel objects, novel backgrounds, and to longer video sequences. We also find that such initial-state-conditioning can be used during inference as a flexible interface to query the model for specific objects or parts of objects, which could pave the way for a range of weakly-supervised approaches and allow more effective interaction with trained models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

    cs.RO 2026-05 unverdicted novelty 7.0

    OA-WAM uses persistent address vectors and dynamic content vectors in object slots to enable addressable world-action prediction, improving robustness on manipulation benchmarks under scene changes.

  2. Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning

    cs.CV 2026-06 unverdicted novelty 6.0

    DSSA decouples per-frame appearance from temporal identity in slot attention mechanisms to reduce slot swapping and improve temporal consistency in video object segmentation.

  3. InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

    cs.CV 2026-05 unverdicted novelty 6.0

    InfoGeo reformulates cross-view geo-localization as an information bottleneck that aligns object-centric structural relations across views while suppressing view-specific noise.

  4. Unsupervised Learning of Inter-Object Relationships via Group Homomorphism

    cs.LG 2026-04 unverdicted novelty 6.0

    An unsupervised model integrates group homomorphism to segment objects and map relative motions like approaching or receding into a one-dimensional additive latent space from unlabeled dynamic images.

  5. OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation

    cs.RO 2026-04 unverdicted novelty 6.0

    OFlow unifies temporal foresight and object-aware reasoning inside a shared latent space via flow matching to improve VLA robustness in robotic manipulation under distribution shifts.

  6. Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0

    A scene-agnostic object codebook learned via unsupervised object-centric learning provides consistent identity-anchored representations for 3D Gaussians across multiple scenes.

  7. Factored Latent Action World Models

    cs.LG 2026-02 conditional novelty 6.0

    FLAM splits a scene into separate factors, each with its own latent action, and reports better video prediction and downstream policy learning than monolithic latent-action models.

  8. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  9. Learning Object-Centric Representations in SAR Images with Multi-Level Feature Fusion

    cs.CV 2025-09 conditional novelty 6.0

    SlotSAR fuses wavelet scattering features with a SAR foundation model's semantic features to make slot attention separate targets from clutter in SAR images, improving segmentation metrics on ATRNet-STAR.

  10. InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

    cs.CV 2026-05 unverdicted novelty 5.0

    InfoGeo reformulates cross-view geo-localization as an information bottleneck that aligns object-centric structural relations across views while minimizing view-specific noise.

  11. InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

    cs.CV 2026-05 unverdicted novelty 5.0

    InfoGeo applies an information bottleneck to object-centric learning for improved cross-view generalization in UAV geo-localization.

  12. InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization

    cs.CV 2026-05 unverdicted novelty 5.0

    InfoGeo reformulates cross-view geo-localization as an information bottleneck that aligns object-centric structural relations while suppressing view-specific noise, outperforming prior methods on benchmarks.

  13. ES-Merging: Biological MLLM Merging via Embedding Space Signals

    cs.LG 2026-03 unverdicted novelty 5.0

    ES-Merging estimates layer-wise and element-wise merge coefficients from coarse- and fine-grained embedding signals and claims better cross-modal reasoning and single-modal knowledge preservation than parameter-space merging.