Pith. sign in

REVIEW 45 cited by

DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.14896 v4 pith:WSGSFIJH submitted 2022-10-26 cs.CV cs.AIcs.HCcs.LG

classification cs.CVcs.AIcs.HCcs.LG
keywords promptsdiffusiondbmodelsdatasetimagesmodelpromptusers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

With recent advancements in diffusion models, users can generate high-quality images by writing text prompts in natural language. However, generating images with desired details requires proper prompts, and it is often unclear how a model reacts to different prompts or what the best prompts are. To help researchers tackle these critical challenges, we introduce DiffusionDB, the first large-scale text-to-image prompt dataset totaling 6.5TB, containing 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users. We analyze the syntactic and semantic characteristics of prompts. We pinpoint specific hyperparameter values and prompt styles that can lead to model errors and present evidence of potentially harmful model usage, such as the generation of misinformation. The unprecedented scale and diversity of this human-actuated dataset provide exciting research opportunities in understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools to help users more easily use these models. DiffusionDB is publicly available at: https://poloclub.github.io/diffusiondb.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 45 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A learnable frequency-domain watermarking scheme for latent diffusion models uses a Slerp-based attack simulation during training and claims lower BER with up to 256-bit payloads.

  2. Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation

    cs.HC 2026-07 accept novelty 7.0 of 10

    A novel 6M-image Pixiv dataset shows open-source image generation has long-tail model usage, slow life cycles with version inertia, and surging multi-LoRA customization linked to higher engagement.

  3. Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Text-to-3D models lose prompt sensitivity for out-of-distribution shapes due to sink traps but retain geometric diversity via unconditional priors, enabling a decoupled inversion method for robust editing.

  4. Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Enforcing Hermitian symmetry when embedding watermarks in the Fourier domain of latent diffusion models improves detection robustness and image fidelity.

  5. Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space

    cs.CV 2025-04 conditional novelty 7.0 of 10

    Prefilled autoregression in a shared embedding space lets Nexus-Gen unify image understanding, generation, and editing, achieving competitive benchmark scores with a 7B model.

  6. Multimodal LLMs Can Reason about Aesthetics in Zero-Shot

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A zero-shot two-stage prompting baseline (ArtCoT) makes multimodal LLMs' aesthetic judgments align substantially better with human expert rankings in pairwise artwork comparisons.

  7. KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Using learned 32×32 Kronecker block transforms as an online activation smoother improves W4A4 image quality of PixArt-Sigma, SANA, and FLUX.1-schnell over SVDQuant and LoRaQ, with a kernel up to 14% faster than SmoothQuant.

  8. Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Diffusion image models can be aligned without human labels by supervising every denoising step with score targets from original versus degraded prompts.

  9. Elastic ViTs from Pretrained Models without Retraining

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A single-shot, label-free, retraining-free structured pruning method generates elastic ViTs at any sparsity by reweighting gradient-based importance scores with block correlations learned by an evolutionary strategy.

  10. TetriServe: Efficiently Serving Mixed DiT Workloads

    cs.LG 2025-10 conditional novelty 6.0 of 10

    TetriServe's step-level, deadline-aware sequence parallelism improves SLO attainment for mixed-resolution diffusion transformer serving by up to 32% over fixed-SP systems.

  11. Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A multi-agent prompt-refinement system using pairwise AI judging and targeted edit signals outperforms prior automated methods on complex text-to-image tasks.

  12. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.

  13. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  14. Guiding Noisy Label Conditional Diffusion Models with Score-based Discriminator Correction

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A discriminator trained to distinguish clean from corrupt image-label pairs can be used during sampling to correct the score of a noisy-label conditional diffusion model, improving class-wise fidelity without retraining.

  15. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  16. Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    1D binary image latents reduce a 1024x1024 image to 128 discrete tokens and support text-to-image generation with diffusion and autoregressive models.

  17. Ambient Diffusion Omni: Training Good Models with Bad Data

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.

  18. SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models

    cs.CV 2025-05 reject novelty 6.0 of 10

    SEED is a 91,526-image benchmark of diffusion-generated sequential facial edits with sequence, mask, and prompt annotations, and FAITH adds DWT high-frequency cues to a transformer for edit-sequence detection.

  19. Towards Self-Improvement of Diffusion Models via Group Preference Optimization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Group Preference Optimization (GPO) uses standardized rewards over self-generated image groups to improve diffusion models' counting, text rendering, and prompt alignment without human preference annotations.

  20. GIFDL: Generated Image Fluctuation Distortion Learning for Enhancing Steganographic Security

    cs.CR 2025-04 conditional novelty 6.0 of 10

    GIFDL trains a GAN-based steganographic distortion model with fluctuation images from a text-to-image generator, raising steganalysis detection error by an average of 3.30% over the GMAN baseline.

  21. SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

    cs.HC 2025-02 conditional novelty 6.0 of 10

    SketchFlex combines sketch-aware prompt recommendation with decompose-and-recompose shape refinement to help novices generate multi-object images from rough region sketches.

  22. T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    T2ISafety is a large annotated benchmark plus a fine-tuned MLLM evaluator (ImageGuard) for measuring toxicity, privacy, and fairness in text-to-image models.

  23. Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A data-free method that tunes only the diagonal scales of Winograd transform matrices enables accurate 8-bit fully quantized Winograd convolution for diffusion models and ResNets.

  24. Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Prompt-A-Video refines text prompts for video diffusion models via evolutionary search and DPO alignment, improving generated video quality on Open-Sora and CogVideoX.

  25. Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Summarizing LLM-obfuscated text-to-image prompts before classification improves content-moderation F1 scores on the new ATTIP dataset.

  26. The Art of Deception: Color Visual Illusions and Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DDIM inversion in diffusion models produces brightness and color shifts that track human visual illusions, and a diffusion-based optimizer can generate new illusions in realistic images that fool human observers.

  27. FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The paper introduces FiVA, a ~1M-image synthetic dataset with fine-grained visual attribute labels, and FiVA-Adapter, a diffusion adapter that transfers and combines attributes like lighting, color, and motion from re...

  28. The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Transfer-based no-box watermark evasion largely fails without aligned surrogate models, and a simple one-surrogate perturbation (OFT) matches or exceeds the expensive optimization-based attack in 11 of 12 tested confi...

  29. Omegance: A Single Parameter for Various Granularities in Diffusion-Based Synthesis

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Scaling the denoising noise by one knob, omega, controls granularity in diffusion outputs globally, per region, or per timestep.

  30. DiffServe: Efficiently Serving Text-to-Image Diffusion Models with Query-Aware Model Scaling

    cs.DC 2024-11 conditional novelty 6.0 of 10

    DiffServe combines a discriminator-based diffusion model cascade with MILP-based resource allocation to serve text-to-image queries with higher quality and fewer SLO violations.

  31. Reward Fine-Tuning Two-Step Diffusion Models via Learning Differentiable Latent-Space Surrogate Reward

    cs.LG 2024-11 conditional novelty 6.0 of 10

    A learned latent-space surrogate reward enables stable fine-tuning of one/two-step diffusion models with arbitrary, non-differentiable reward signals, outperforming policy-gradient baselines.

  32. ComfyGI: Automatic Improvement of Image Generation Workflows

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A genetic-improvement search over ComfyUI workflow JSONs, guided by ImageReward, raises median reward scores by about 50% and wins about 90% of human preference comparisons.

  33. Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    The paper introduces Image Regeneration, an evaluation benchmark where text-to-image models must reproduce a reference image from MLLM-generated prompts, along with the ImageRepainter framework and two new datasets.

  34. Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact unified model that reuses a frozen VLM encoder and hybrid continuous/discrete tokens reaches competitive image understanding and generation with 15.6M training images and about $2,000 in compute.

  35. Cost-Aware Routing for Efficient Text-To-Image Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A cost-aware router selects per prompt the best among nine pre-trained text-to-image models, beating every single model on the quality-versus-cost frontier.

  36. Exploring Language Patterns of Prompts in Text-to-Image Generation and Their Impact on Visual Diversity

    cs.HC 2025-04 conditional novelty 5.0 of 10

    User prompts on CivitAI became more lexically repetitive over seven months, and higher prompt similarity was associated with lower visual diversity in generated images.

  37. EliGen: Entity-Level Controlled Image Generation with Regional Attention

    cs.CV 2025-01 conditional novelty 5.0 of 10

    EliGen adds entity-level control to FLUX by composing attention masks that bind each entity prompt to its spatial region, then fine-tunes with LoRA on 500k synthetic annotated images.

  38. VersaGen: Unleashing Versatile Visual Control for Text-to-Image Synthesis

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VersaGen enables users to control text-to-image diffusion models with partial sketch inputs at object and scene levels, with automatic localization and adaptive control strength.

  39. Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration

    cs.CV 2026-06 conditional novelty 4.5 of 10

    A DAAM-based visual analytics workflow links step-resolved token attention trajectories, phase summaries, and spatial competition maps for Stable Diffusion-class models on a 60-prompt benchmark.

  40. NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

    cs.CV 2025-05 conditional novelty 4.0 of 10

    The NTIRE 2025 challenge report compares 20 methods for fine-grained text-to-image quality assessment, introduces the EvalMuse-Structure dataset, and finds every participating team outperformed the baselines.

  41. DejAIvu: Identifying and Explaining AI Art on the Web in Real-Time with Saliency Maps

    cs.CV 2025-02 conditional novelty 4.0 of 10

    DejAIvu is a browser extension that classifies images as AI-generated or human-made in real time and highlights the evidence with saliency heatmaps.

  42. Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback

    cs.AI 2024-12 conditional novelty 4.0 of 10

    The paper proposes learning from language feedback to synthesize multimodal preference pairs, but the evidence is weakened by an undefined improvement metric and small, unvalidated effect sizes.

  43. Detecting Facial Image Manipulations with Multi-Layer CNN Models

    cs.CV 2024-12 reject novelty 4.0 of 10

    Extending MesoNet to six convolutional layers and swapping the deepfake training images for Stable Diffusion images raises three-class accuracy from 43% to 76% on a small in-distribution dataset, without external validation.

  44. Artificial Intelligence Policy Framework for Institutions

    cs.CY 2024-12 reject novelty 2.0 of 10

    A proposed AI governance flowchart for institutions, assembled from standard principles, without evidence that it works.

  45. Visual question answering based evaluation metrics for text-to-image generation

    cs.CV 2024-11 reject novelty 2.0 of 10

    A text-to-image evaluation metric that combines ChatGPT-generated yes/no questions answered by a VQA model with a no-reference image quality score, with adjustable weights.

Pith tools