REVIEW 44 cited by
Text-to-image Diffusion Models in Generative AI: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This survey reviews the progress of diffusion models in generating images from text, ~\textit{i.e.} text-to-image diffusion models. As a self-contained work, this survey starts with a brief introduction of how diffusion models work for image synthesis, followed by the background for text-conditioned image synthesis. Based on that, we present an organized review of pioneering methods and their improvements on text-to-image generation. We further summarize applications beyond image generation, such as text-guided generation for various modalities like videos, and text-guided image editing. Beyond the progress made so far, we discuss existing challenges and promising future directions.
Forward citations
Cited by 44 Pith papers
-
Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation
A novel 6M-image Pixiv dataset shows open-source image generation has long-tail model usage, slow life cycles with version inertia, and surging multi-LoRA customization linked to higher engagement.
-
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
BadReward uses clean-label feature-collision images to poison CLIP-based reward models so that a text-to-image model produces target attributes (e.g., glasses, skin tone, blood) when the trigger phrase is present.
-
EmergencyBias: Bias in Text-to-Image Models under Emergency Scenarios
In emergency scenes, text-to-image models skew who appears and who helps, favoring men, middle-aged, and lighter-skinned people, and a soft-token tweak reduces the gap.
-
Testing chatbots on the creation of encoders for audio conditioned image generation
All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.
-
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
DiffRIS combines frozen Stable Diffusion and CLIP encoders with a new text adapter and decoder to set a new state-of-the-art mean IoU on three referring remote sensing image segmentation benchmarks.
-
Rethinking Machine Unlearning in Image Generation Models
A new taxonomy and multi-aspect evaluation framework for image generation unlearning, with a curated dataset, shows that ten existing unlearning methods perform poorly on preservation and robustness.
-
Latent Guidance in Diffusion Models for Perceptual Evaluations
LGDM uses perceptual guidance during Stable Diffusion sampling and aggregates multi-scale, multi-timestep U-Net features to predict human-rated image quality, achieving state-of-the-art correlations on ten NR-IQA datasets.
-
Text2CT: Towards 3D CT Volume Generation from Free-text Descriptions Using Diffusion Model
Text2CT is a unified 3D diffusion model that generates 512x512x192 chest CT volumes from free-text clinical descriptions, outperforming prior text-to-CT baselines on FID, CLIP score, and data augmentation.
-
On-device Sora: Enabling Training-Free Diffusion-based Text-to-Video Generation for Mobile Devices
A training-free pipeline makes diffusion text-to-video generation run on an iPhone 15 Pro with quality close to GPU output, at the cost of slower generation.
-
Concept Guided Co-salient Object Detection
ConceptCoSOD extracts a shared text embedding from an image group and guides co-salient object segmentation with it, outperforming five baselines on three clean and five corrupted datasets.
-
AsymRnR: Video Diffusion Transformers Acceleration with Asymmetric Reduction and Restoration
AsymRnR selectively reduces query and key/value tokens in video DiT attention to cut FLOPs and latency by 10 to 30 percent with minor or no VBench score change.
-
Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics
A new benchmark and fine-tuned models show that multi-task training improves average perceptual-similarity accuracy on known tasks but does not generalize to out-of-distribution perceptual tasks.
-
Needle: A Generative AI-Powered Multi-modal Database for Answering Complex Natural Language Queries
Needle generates AI-made query images from text, embeds them with an ensemble of visual models, and uses nearest-neighbor search to retrieve matching real images, beating zero-shot text-image baselines on complex queries.
-
Any-Resolution AI-Generated Image Detection by Spectral Learning
SPAI uses spectral reconstruction similarity from a frozen masked-frequency ViT plus attention pooling to reach 91.0 average AUC for AI-generated image detection across 13 unseen generators.
-
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
JuZhou 1.0 is a 0.387B-parameter T2I diffusion model with 4-step inference achieving 0.69 GenEval, trained on 9M Chinese pairs using Sugon K100 accelerators and deployable on Android/iOS devices.
-
Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
A concept trained on visually uniform images tends to dominate and suppress a second concept in multi-concept text-to-image generation, a failure the authors name DvD.
-
GeoLoom: High-quality Geometric Diagram Generation from Textual Input
Natural-language geometry descriptions can be autoformalized into a custom geometry language and converted to coordinates by Monte Carlo optimization, yielding usable diagrams in seconds for about 81-85% of test problems.
-
ShaLa: Multimodal Shared Latent Space Modelling
VIPER-R1 fine-tunes a vision-language model to read kinematic plots and propose symbolic equations, then refines them with symbolic regression, but its final metric is computed on the same data used for the refinement fit.
-
VideoGuard: Protecting Video Content from Unauthorized Editing
VideoGuard adds joint, motion-aware perturbations to videos to block unauthorized diffusion-model editing.
-
Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art
Ethics experts' definitions of five ethical theories, rendered as DALL-E 3 images and re-evaluated by the same experts, yield eight themes showing how morality, society, and learned associations shape and bias the mod...
-
Learning Graph Representation of Agent Diffusers
LGR-AD claims to improve text-to-image generation by coordinating multiple diffusion models via a graph convolutional network, but the loss function's KL term contradicts the stated goal of maximizing diversity.
-
Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models
Masking the image-feature dimensions most correlated with the style reference's content text reduces content leakage and improves text fidelity in text-to-image style transfer diffusion models.
-
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
Applying pairwise direct preference optimization to score distillation makes text-to-3D outputs better aligned with human preferences and more controllable.
-
Exploring the latent space of diffusion models directly through singular value decomposition
The authors report that singular value decomposition of diffusion latent codes reveals stable, order-mobile attribute directions and propose Attribute Vector Integration, a per-pair MLP-based editor that transfers tex...
-
Combining Genre Classification and Harmonic-Percussive Features with Diffusion Models for Music-Video Generation
An automatic music-visualizer pipeline that uses genre-guided image generation and audio-energy-controlled frame interpolation beats linear interpolation on a new synchrony metric.
-
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
Blending a safe-caption denoising branch with the original-prompt branch during early diffusion steps generates safer images without editing or erasing concepts.
-
Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
A DAAM-based visual analytics workflow links step-resolved token attention trajectories, phase summaries, and spatial competition maps for Stable Diffusion-class models on a 60-prompt benchmark.
-
Effectively obtaining acoustic, visual and textual data from videos
A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.
-
T2VWorldBench: A Benchmark for Evaluating World Knowledge in Text-to-Video Generation
A 1,200-prompt benchmark across six world-knowledge domains reports that ten state-of-the-art text-to-video models average below 0.70 on a 0 to 1 scale for producing videos consistent with real-world knowledge.
-
SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection
A spatial-frequency feature fusion network with attention achieves improved accuracy on remote sensing image forgery detection and introduces a stable-diffusion-based benchmark.
-
Toward Rich Video Human-Motion2D Generation
A new 150K-video 2D skeleton dataset with text captions and a diffusion model for single- and double-character motion generation, though the claimed FID-rewarded RL training is misrepresented.
-
Breaking the Barriers of Text-Hungry and Audio-Deficient AI
A proposed audio-native translation framework called MAST with fractional diffusion is described, but no evidence is given that it produces working translations.
-
Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization
Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.
-
MRI Image Generation Based on Text Prompts
Fine-tuning Stable Diffusion with MRI-text pairs yields plausible brain MRI images by field strength and modality, and synthetic images appear to improve a small MRI contrast classification task.
-
Generative Distribution Prediction: A Unified Approach to Multimodal Learning
GDP predicts by sampling candidate outputs from a conditional generative model and selecting the one with the lowest average loss; its excess risk is bounded by generation error plus a vanishing sampling term.
-
FlexMotion: Lightweight, Physics-Aware, and Controllable Human Motion Generation
FlexMotion reports a latent diffusion model with a physics-aware multimodal autoencoder and a ControlNet-style module for controlling joint locations, contact forces, joint actuations, and muscle activations in genera...
-
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
A new 10.5M-pair remote sensing dataset and a 1.3B diffusion model generate resolution-controlled satellite imagery from text, with large reported gains on the RSICD benchmark.
-
Enhancing Diffusion Models for Inverse Problems with Covariance-Aware Posterior Sampling
CA-DPS approximates the covariance of the reverse diffusion process via a finite-difference Hessian estimate and uses it to improve posterior sampling for linear inverse problems.
-
Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey
A review of non-centralized continual learning that taxonomizes data-, model-, and device-level methods and benchmarks twelve federated continual learning methods on six datasets.
-
DLSF: Dual-Layer Synergistic Fusion for High-Fidelity Image Syn-thesis
A softmax and spatial attention fusion of SDXL base and refiner latents yields an ImageNet FID drop of about 1 point, but the result is not statistically supported.
-
Text to Image Generation and Editing: A Survey
A broad survey of text-to-image generation and editing research from 2021 to 2024, organized by architecture and comparison tables.
-
Diffusion Models for Hyperspectral Image Analysis: A Comprehensive Review
A literature review that organizes diffusion-model work for hyperspectral imaging into eight task categories and compiles comparative performance tables from prior papers.
-
Foundations of GenIR
A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.