Pith. sign in

REVIEW 44 cited by

Text-to-image Diffusion Models in Generative AI: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.07909 v3 pith:PP4G2WDI submitted 2023-03-14 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords diffusionimagemodelsgenerationsurveytext-to-imagebeyondprogress
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This survey reviews the progress of diffusion models in generating images from text, ~\textit{i.e.} text-to-image diffusion models. As a self-contained work, this survey starts with a brief introduction of how diffusion models work for image synthesis, followed by the background for text-conditioned image synthesis. Based on that, we present an organized review of pioneering methods and their improvements on text-to-image generation. We further summarize applications beyond image generation, such as text-guided generation for various modalities like videos, and text-guided image editing. Beyond the progress made so far, we discuss existing challenges and promising future directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 44 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 76 citations worldwide. Full citation record

  1. Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation

    cs.HC 2026-07 accept novelty 7.0 of 10

    A novel 6M-image Pixiv dataset shows open-source image generation has long-tail model usage, slow life cycles with version inertia, and surging multi-LoRA customization linked to higher engagement.

  2. BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF

    cs.LG 2025-06 conditional novelty 7.0 of 10

    BadReward uses clean-label feature-collision images to poison CLIP-based reward models so that a text-to-image model produces target attributes (e.g., glasses, skin tone, blood) when the trigger phrase is present.

  3. EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios

    cs.MM 2026-08 conditional novelty 6.0 of 10

    In emergency scenes, text-to-image models skew who appears and who helps, favoring men, middle-aged, and lighter-skinned people, and a soft-token tweak reduces the gap.

  4. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  5. DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DiffRIS combines frozen Stable Diffusion and CLIP encoders with a new text adapter and decoder to set a new state-of-the-art mean IoU on three referring remote sensing image segmentation benchmarks.

  6. Rethinking Machine Unlearning in Image Generation Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new taxonomy and multi-aspect evaluation framework for image generation unlearning, with a curated dataset, shows that ten existing unlearning methods perform poorly on preservation and robustness.

  7. Latent Guidance in Diffusion Models for Perceptual Evaluations

    cs.CV 2025-05 conditional novelty 6.0 of 10

    LGDM uses perceptual guidance during Stable Diffusion sampling and aggregates multi-scale, multi-timestep U-Net features to predict human-rated image quality, achieving state-of-the-art correlations on ten NR-IQA datasets.

  8. Text2CT: Towards 3D CT Volume Generation from Free-text Descriptions Using Diffusion Model

    eess.IV 2025-05 conditional novelty 6.0 of 10

    Text2CT is a unified 3D diffusion model that generates 512x512x192 chest CT volumes from free-text clinical descriptions, outperforming prior text-to-CT baselines on FID, CLIP score, and data augmentation.

  9. On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A training-free pipeline makes diffusion text-to-video generation run on an iPhone 15 Pro with quality close to GPU output, at the cost of slower generation.

  10. Concept Guided Co-salient Object Detection

    cs.CV 2024-12 conditional novelty 6.0 of 10

    ConceptCoSOD extracts a shared text embedding from an image group and guides co-salient object segmentation with it, outperforming five baselines on three clean and five corrupted datasets.

  11. AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AsymRnR selectively reduces query and key/value tokens in video DiT attention to cut FLOPs and latency by 10 to 30 percent with minor or no VBench score change.

  12. Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.

  13. Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries

    cs.IR 2024-12 conditional novelty 6.0 of 10

    Needle generates AI-made query images from text, embeds them with an ensemble of visual models, and uses nearest-neighbor search to retrieve matching real images, beating zero-shot text-image baselines on complex queries.

  14. Any-Resolution AI-Generated Image Detection by Spectral Learning

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SPAI uses spectral reconstruction similarity from a frozen masked-frequency ViT plus attention pooling to reach 91.0 average AUC for AI-generated image detection across 13 unseen generators.

  15. JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.

  16. Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models

    cs.LG 2025-12 conditional novelty 5.0 of 10

    A concept trained on visually uniform images tends to dominate and suppress a second concept in multi-concept text-to-image generation, a failure the authors name DvD.

  17. GeoLoom: High-quality Geometric Diagram Generation from Textual Input

    cs.CV 2025-12 conditional novelty 5.0 of 10

    Natural-language geometry descriptions can be autoformalized into a custom geometry language and converted to coordinates by Monte Carlo optimization, yielding usable diagrams in seconds for about 81-85% of test problems.

  18. ShaLa: Multimodal Shared Latent Space Modelling

    cs.LG 2025-08 reject novelty 5.0 of 10

    VIPER-R1 fine-tunes a vision-language model to read kinematic plots and propose symbolic equations, then refines them with symbolic regression, but its final metric is computed on the same data used for the refinement fit.

  19. VideoGuard: Protecting Video Content from Unauthorized Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    VideoGuard adds joint, motion-aware perturbations to videos to block unauthorized diffusion-model editing.

  20. Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Ethics experts' definitions of five ethical theories, rendered as DALL-E 3 images and re-evaluated by the same experts, yield eight themes showing how morality, society, and learned associations shape and bias the mod...

  21. Learning Graph Representation of Agent Diffusers

    cs.LG 2025-05 reject novelty 5.0 of 10

    LGR-AD claims to improve text-to-image generation by coordinating multiple diffusion models via a graph convolutional network, but the loss function's KL term contradicts the stated goal of maximizing diversity.

  22. Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Masking the image-feature dimensions most correlated with the style reference's content text reduces content leakage and improves text fidelity in text-to-image style transfer diffusion models.

  23. DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Applying pairwise direct preference optimization to score distillation makes text-to-3D outputs better aligned with human preferences and more controllable.

  24. Exploring the latent space of diffusion models directly through singular value decomposition

    cs.CV 2025-02 reject novelty 5.0 of 10

    The authors report that singular value decomposition of diffusion latent codes reveals stable, order-mobile attribute directions and propose Attribute Vector Integration, a per-pair MLP-based editor that transfers tex...

  25. Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation

    cs.MM 2024-12 conditional novelty 5.0 of 10

    An automatic music-visualizer pipeline that uses genre-guided image generation and audio-energy-controlled frame interpolation beats linear interpolation on a new synchrony metric.

  26. Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Blending a safe-caption denoising branch with the original-prompt branch during early diffusion steps generates safer images without editing or erasing concepts.

  27. Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration

    cs.CV 2026-06 conditional novelty 4.5 of 10

    A DAAM-based visual analytics workflow links step-resolved token attention trajectories, phase summaries, and spatial competition maps for Stable Diffusion-class models on a 60-prompt benchmark.

  28. Effectively obtaining acoustic, visual and textual data from videos

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.

  29. T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation

    cs.CV 2025-07 reject novelty 4.0 of 10

    A 1,200-prompt benchmark across six world-knowledge domains reports that ten state-of-the-art text-to-video models average below 0.70 on a 0 to 1 scale for producing videos consistent with real-world knowledge.

  30. SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A spatial-frequency feature fusion network with attention achieves improved accuracy on remote sensing image forgery detection and introduces a stable-diffusion-based benchmark.

  31. Toward Rich Video Human-Motion2D Generation

    cs.CV 2025-06 reject novelty 4.0 of 10

    A new 150K-video 2D skeleton dataset with text captions and a diffusion model for single- and double-character motion generation, though the claimed FID-rewarded RL training is misrepresented.

  32. Breaking the Barriers of Text-Hungry and Audio-Deficient AI

    cs.SD 2025-06 reject novelty 4.0 of 10

    A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.

  33. Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

    cs.CV 2025-05 reject novelty 4.0 of 10

    Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.

  34. MRI Image Generation Based on Text Prompts

    eess.IV 2025-05 conditional novelty 4.0 of 10

    Fine-tuning Stable Diffusion with MRI-text pairs yields plausible brain MRI images by field strength and modality, and synthetic images appear to improve a small MRI contrast classification task.

  35. Generative Distribution Prediction: A Unified Approach to Multimodal Learning

    stat.ML 2025-02 conditional novelty 4.0 of 10

    GDP predicts by sampling candidate outputs from a conditional generative model and selecting the one with the lowest average loss; its excess risk is bounded by generation error plus a vanishing sampling term.

  36. FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation

    cs.CV 2025-01 reject novelty 4.0 of 10

    FlexMotion reports a latent diffusion model with a physics-aware multimodal autoencoder and a ControlNet-style module for controlling joint locations, contact forces, joint actuations, and muscle activations in genera...

  37. Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model

    cs.CV 2025-01 reject novelty 4.0 of 10

    A new 10.5M-pair remote sensing dataset and a 1.3B diffusion model generate resolution-controlled satellite imagery from text, with large reported gains on the RSICD benchmark.

  38. Enhancing Diffusion Models for Inverse Problems with Covariance-Aware Posterior Sampling

    cs.CV 2024-12 reject novelty 4.0 of 10

    CA-DPS approximates the covariance of the reverse diffusion process via a finite-difference Hessian estimate and uses it to improve posterior sampling for linear inverse problems.

  39. Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A review of non-centralized continual learning that taxonomizes data-, model-, and device-level methods and benchmarks twelve federated continual learning methods on six datasets.

  40. DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis

    cs.GR 2025-07 reject novelty 3.0 of 10

    A softmax and spatial attention fusion of SDXL base and refiner latents yields an ImageNet FID drop of about 1 point, but the result is not statistically supported.

  41. Text to Image Generation and Editing: A Survey

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A broad survey of text-to-image generation and editing research from 2021 to 2024, organized by architecture and comparison tables.

  42. Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review

    eess.IV 2025-05 conditional novelty 2.0 of 10

    A literature review that organizes diffusion-model work for hyperspectral imaging into eight task categories and compiles comparative performance tables from prior papers.

  43. Foundations of GenIR

    cs.IR 2025-01 unverdicted novelty 1.0 of 10

    A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.

  44. Text-to-Image Synthesis: A Decade Survey

    cs.CV 2024-11 conditional novelty 1.0 of 10

    A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.

Pith tools