Pith. sign in

REVIEW 7 cited by

GENESIS: Generative Scene Inference and Sampling with Object-Centric Latent Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.13052 v4 pith:MOEAJTSZ submitted 2019-07-30 cs.LG cs.CVcs.NEcs.ROstat.ML

classification cs.LGcs.CVcs.NEcs.ROstat.ML
keywords generativescenesgenesisobject-centricscenefashionlatentlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generative latent-variable models are emerging as promising tools in robotics and reinforcement learning. Yet, even though tasks in these domains typically involve distinct objects, most state-of-the-art generative models do not explicitly capture the compositional nature of visual scenes. Two recent exceptions, MONet and IODINE, decompose scenes into objects in an unsupervised fashion. Their underlying generative processes, however, do not account for component interactions. Hence, neither of them allows for principled sampling of novel scenes. Here we present GENESIS, the first object-centric generative model of 3D visual scenes capable of both decomposing and generating scenes by capturing relationships between scene components. GENESIS parameterises a spatial GMM over images which is decoded from a set of object-centric latent variables that are either inferred sequentially in an amortised fashion or sampled from an autoregressive prior. We train GENESIS on several publicly available datasets and evaluate its performance on scene generation, decomposition, and semi-supervised learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Steering Optimisation Trajectories in Diffusion Representation Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SteeringDRL identifies two optimization regimes in diffusion autoencoders and uses gated residual U-Nets with a log SNR curriculum to steer training toward disentangled representations, improving performance across mu...

  2. Dyn-O: Building Structured World Models with Object-Centric Representations

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Dyn-O learns object-centric world models directly from pixels in complex Procgen games, using SAM2-guided slot attention and Mamba state-space dynamics, and reports better rollout prediction than DreamerV3.

  3. IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.

  4. Identifiable Object Representations under Spatial Ambiguities

    cs.LG 2025-06 reject novelty 6.0 of 10

    VISA learns view-invariant object representations by aggregating probabilistic slots across multiple unlabeled viewpoints, with an identifiability analysis up to affine and permutation equivalence.

  5. Hierarchical Compact Clustering Attention (COCA) for Unsupervised Object-Centric Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    COCA-Net introduces compactness-guided hierarchical clustering within an attention architecture, achieving state-of-the-art unsupervised object segmentation on synthetic multi-object images.

  6. Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A survey organizes multimodal reasoning research into a staged roadmap and proposes native large multimodal reasoning models that unify perception, generation, and agentic planning.

  7. On the Benefits of Instance Decomposition in Video Prediction Models

    cs.CV 2025-01 reject novelty 4.0 of 10

    Explicit instance decomposition with per-class shared weights improves latent-transformer video prediction in the paper's experiments, but the claimed advantage is weakened by mismatched parameter counts and test-set ...

Pith tools