Pith. sign in

REVIEW 2 cited by

DiffBlender: Composable and Versatile Multimodal Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15194 v3 pith:WKL3INGZ submitted 2023-05-24 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffblenderdiffusiongenerationmultimodalconditionaldiversemodalitiesmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this study, we aim to enhance the capabilities of diffusion-based text-to-image (T2I) generation models by integrating diverse modalities beyond textual descriptions within a unified framework. To this end, we categorize widely used conditional inputs into three modality types: structure, layout, and attribute. We propose a multimodal T2I diffusion model, which is capable of processing all three modalities within a single architecture without modifying the parameters of the pre-trained diffusion model, as only a small subset of components is updated. Our approach sets new benchmarks in multimodal generation through extensive quantitative and qualitative comparisons with existing conditional generation methods. We demonstrate that DiffBlender effectively integrates multiple sources of information and supports diverse applications in detailed image synthesis. The code and demo are available at https://github.com/sungnyun/diffblender.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnyI2V: Animating Any Conditional Image with Motion Control

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.

  2. MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    MixDiffusion derives a joint noise prediction as the sum of per-condition noise estimates minus the base model, enabling multi-condition control without training.

Pith tools