Pith. sign in

REVIEW 2 cited by

DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14216 v2 pith:FTW3KB6H submitted 2023-12-21 cs.CV

classification cs.CV
keywords imagesdiffusiondistributiongenerationpromptsdiversedreamdistributionlearned
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The popularization of Text-to-Image (T2I) diffusion models enables the generation of high-quality images from text descriptions. However, generating diverse customized images with reference visual attributes remains challenging. This work focuses on personalizing T2I diffusion models at a more abstract concept or category level, adapting commonalities from a set of reference images while creating new instances with sufficient variations. We introduce a solution that allows a pretrained T2I diffusion model to learn a set of soft prompts, enabling the generation of novel images by sampling prompts from the learned distribution. These prompts offer text-guided editing capabilities and additional flexibility in controlling variation and mixing between multiple distributions. We also show the adaptability of the learned prompt distribution to other tasks, such as text-to-3D. Finally we demonstrate effectiveness of our approach through quantitative analysis including automatic evaluation and human assessment. Project website: https://briannlongzhao.github.io/DreamDistribution

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Orchestrating the Symphony of Prompt Distribution Learning for Human-Object Interaction Detection

    cs.CV 2024-12 conditional novelty 5.0 of 10

    InterProDa models each HOI category as a Gaussian distribution over soft prompt embeddings, sampled with a learnable noise factor, and reports SOTA results on HICO-DET and V-COCO.

  2. MTFusion: Reconstructing Any 3D Object from Single Image Using Multi-word Textual Inversion

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A two-stage pipeline that extracts a multi-word text embedding from a single image and uses it, with a FlexiCubes-based SDF network, to reconstruct a textured 3D mesh.

Pith tools