Pith. sign in

REVIEW 8 cited by

Object-Centric Learning with Slot Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.15055 v2 pith:JFNEWJQU submitted 2020-06-26 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords representationsattentionlearningobject-centricslotabstractobjectperceptual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed representations that do not capture the compositional properties of natural scenes. In this paper, we present the Slot Attention module, an architectural component that interfaces with perceptual representations such as the output of a convolutional neural network and produces a set of task-dependent abstract representations which we call slots. These slots are exchangeable and can bind to any object in the input by specializing through a competitive procedure over multiple rounds of attention. We empirically demonstrate that Slot Attention can extract object-centric representations that enable generalization to unseen compositions when trained on unsupervised object discovery and supervised property prediction tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 218 citations worldwide. Full citation record

  1. Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Goal-conditioned world models transcribe instructions instead of perceiving spatial relations when the instruction names the scored quantity, and removing the goal from the dynamics fixes it.

  2. Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

    cs.CV 2026-08 reject novelty 6.0 of 10

    ViWi uses material slots and a simulated RF descriptor to predict voxel-level Young's modulus, Poisson's ratio, and density, reporting gains over prior work on a synthetic benchmark.

  3. Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models

    cs.SD 2026-07 unverdicted novelty 6.0 of 10

    ORCA splits Q-Former queries into orthogonally constrained groups, reversing directional collapse and speaker-indistinguishability in audio-LLM connectors and gaining 26.4 points on SAKURA multi-hop reasoning.

  4. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0 of 10

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  5. Successes and Limitations of Object-centric Models at Compositional Generalisation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Object-centric models handle novel combinations of object properties when all local features are present, and can even extrapolate to unseen shapes on the Pentomino dataset.

  6. Mixing Configurations for Downstream Prediction

    cs.LG 2025-10 conditional novelty 5.0 of 10

    Mixing multiple resolution clusterings of an embedding, aligned between train and test, and fused by attention, improves downstream regression and classification over single-resolution baselines.

  7. Agent Identity Evals: Measuring Agentic Identity

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Introduces Agent Identity Evals (AIE), five similarity-based metrics for LMA identity stability, with pilot experiments showing identifiability always at zero and no statistical support.

  8. Is an object-centric representation beneficial for robotic manipulation ?

    cs.AI 2025-06 reject novelty 4.0 of 10

    Evaluating the object-centric SAVi encoder against the global DINO and R3M representations on three simulated manipulation tasks, the authors find SAVi is the only model to solve the pick task and is more robust to un...

Pith tools