Pith. sign in

Joint Generative Modeling of Grounded Scene Graphs and Images via Diffusion Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We introduce a framework for joint grounded scene graph - image generation, a challenging task involving high-dimensional, multi-modal structured data. To effectively model this complex joint distribution, we adopt a factorized approach: first generating a grounded scene graph, followed by image generation conditioned on the generated grounded scene graph. While conditional image generation has been widely explored in the literature, our primary focus is on the generation of grounded scene graphs from noise, which provides efficient and interpretable control over the image generation process. This task requires generating plausible grounded scene graphs with heterogeneous attributes for both nodes (objects) and edges (relations among objects), encompassing continuous attributes (e.g., object bounding boxes) and discrete attributes (e.g., object and relation categories). To address this challenge, we introduce DiffuseSG, a novel diffusion model that jointly models the heterogeneous node and edge attributes. We explore different encoding strategies to effectively handle the categorical data. Leveraging a graph transformer as the denoiser, DiffuseSG progressively refines grounded scene graph representations in a continuous space before discretizing them to generate structured outputs. Additionally, we introduce an IoU-based regularization term to enhance empirical performance. Our model outperforms existing methods in grounded scene graph generation on the VG and COCO-Stuff datasets, excelling in both standard and newly introduced metrics that more accurately capture the task's complexity. Furthermore, we demonstrate the broader applicability of DiffuseSG in two important downstream tasks: 1) achieving superior results in a range of grounded scene graph completion tasks, and 2) enhancing grounded scene graph detection models by leveraging additional training samples generated by DiffuseSG.

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

TrajFlow: Multi-modal Motion Prediction via Flow Matching

cs.CV · 2025-06-10 · conditional · novelty 6.0

TrajFlow uses flow matching with a multi-query transformer to predict multiple trajectories in one pass and a Plackett-Luce ranking loss to improve confidence scores, reporting small SOTA gains on WOMD.

citing papers explorer

Showing 1 of 1 citing paper.

  • TrajFlow: Multi-modal Motion Prediction via Flow Matching cs.CV · 2025-06-10 · conditional · none · ref 49 · internal anchor

    TrajFlow uses flow matching with a multi-query transformer to predict multiple trajectories in one pass and a Plackett-Luce ranking loss to improve confidence scores, reporting small SOTA gains on WOMD.