REVIEW 4 major objections 8 minor 122 references
From Noise to Nuance: Advances in Deep Generative Image Models
T0 review · 4 major / 8 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This survey argues that deep generative image models shifted decisively from GANs to diffusion and transformer architectures after 2021, driven by latent-space efficiency and control, with resource-conscious and interpretable systems…
desk verdict A broad but unreliable survey of deep generative image models; the equation errors and copy-paste problems mean it should not be trusted as a reference until extensively revised. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's organizing object is the latent-space diffusion process with cross-attention conditioning: an autoencoder compresses images into a latent, a denoising network operates there, and cross-attention lets text or other signals guide each step. For efficiency, the other central object is the consistency function, which collapses the probability-flow trajectory so that every point on it maps to the same clean image, enabling one-step or few-step generation when learned or distilled. Low-rank adapters and quantization are the supporting machinery for deploying these models cheaply.
What would settle it
A reproducible audit that compares the SDXL-Lightning discriminator equations and the consistency-model loss with the original released implementations would settle the fidelity question: any mismatch in the loss or discriminator definitions would show that the survey's summaries of those models are unreliable. For the historical claim itself, an independent benchmark in which few-step distilled models consistently beat multi-step diffusion on every standard quality metric would weaken the paper's claim that quality-speed trade-offs remain a critical challenge.
Extended reading notes
Core claim
The central claim is that since 2021 image generation has shifted from GANs to diffusion models, and then to compute-efficient variants: latent diffusion in a compressed space, transformer-based diffusion, autoregressive and masked token transformers, and consistency models that map any noisy point on a trajectory to its clean origin. The paper further claims that conditioning mechanisms, including cross-attention on text, structural control inputs, and style-transfer adapters, turned these models into controllable tools. Taken together, the survey's account says the field's progress is best understood as a race between generative quality and computational cost, and that the remaining frontier is resource-conscious, interpretable architectures.
Load-bearing premise
The survey's narrative is trustworthy only if its reproductions of primary-source equations and model descriptions are faithful, and that fidelity is not independently verified.
Editorial extensions
If this is right
- If the survey's historical account is right, future image-generation research will keep operating in learned latent spaces, making the quality of the autoencoder a first-order performance factor.
- If consistency and distillation methods truly preserve quality at few steps, interactive and real-time image editing on consumer hardware becomes feasible.
- If the claimed efficiency techniques work as described, quantization and low-rank adaptation will become standard deployment practice for large text-to-image models.
- If multi-component prompt comprehension remains hard, evaluation will need to move toward compositional benchmarks rather than single-prompt quality scores.
Reading between the lines
- An unstated consequence of the survey's framing is that further advances in generation quality are likely to come from better learned autoencoders and noise schedules rather than from scaling the denoising network alone.
- Because the survey presents equations without code or independent verification, a reader should treat its reproduced formulas as pointers to the original papers rather than as authoritative derivations.
- A testable extension the survey leaves implicit: a systematic comparison of consistency losses with and without the stop-gradient rule across data scales could reveal whether the consistency property itself, rather than the time-sampling schedule, drives the reported efficiency.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a survey of deep generative image models, tracing the transition from GANs to diffusion- and transformer-based architectures. It covers DDPMs, latent diffusion, Stable Diffusion variants (SD1.x–SD3, SDXL-Turbo, SDXL-Lightning), consistency models, Imagen, DALL-E series, and practical topics such as quantization, PEFT, inpainting/outpainting, ControlNet, style transfer, and evaluation metrics. The authors argue that the field has undergone a paradigm shift since 2021 and that, despite advances in quality and efficiency, resource constraints and interpretability remain open problems. The paper concludes by outlining future research directions including neural architecture optimization and explainable generation frameworks.
Significance. As a survey, the paper's value depends on the fidelity of its technical summaries and the comprehensiveness of its coverage. The standard DDPM equations (Eqs. 2–8) and the general LDM framework (Eqs. 10–18) are reproduced correctly, and the paper covers an impressively broad set of topics, including efficiency methods, control, and evaluation. However, the survey offers no original experiments, analysis, or falsifiable predictions, and its reliability is undermined by a copied first-person fragment, a duplicated Consistency Models section, and at least two load-bearing equation inaccuracies. If corrected, the survey could serve as a useful entry point, but in its current form its technical descriptions cannot be fully trusted.
major comments (4)
- [Section II.B.2.d, Eq. (19)] The LDM training objective is written as E[||ε − ε_θ(z_t, t)||²], omitting the conditioning input c that appears in the same paper's Eqs. (13), (17), and (18) and in the original LDM objective (Rombach et al., Eq. 3). Because conditioning is the mechanism that makes text-to-image synthesis possible, this misstates the core model and is internally inconsistent. The equation should read E[||ε − ε_θ(z_t, t, c)||²] (or, in original notation, with the text-encoder embedding τ_θ(y) as the conditioning input).
- [Section II.C, Eq. (24)] The consistency loss as written evaluates f_θ on the same input x at two times t1 and t2, which does not reflect the consistency-model training procedure. In Song et al. [49], the loss compares the model outputs at two points x_t and x_t′ that are noisy versions of the same data sample along the probability-flow ODE, and it uses a target network (e.g., an EMA of f_θ). The current formula also omits the target network and the sampling of noisy points, so it does not enforce the self-consistency property described in Eq. (23). Please replace it with the actual consistency distillation or consistency training objective and state the appropriate sampling procedure for x, t1, and t2.
- [Section II.B.3.b] The sentence 'Our models are available as both LoRA and full UNet weights' is a first-person statement that belongs to the SDXL-Lightning authors, not to this survey. It appears without quotation marks or attribution, which is a verbatim copying problem. This casts doubt on the reliability of the surrounding technical content (Eqs. 21–22); the authors should rewrite all such material in third-person form and verify every equation against the source.
- [Section II.B.4 and Section II.C] Consistency Models are presented twice: first as a short subsection under 'Diffusion Model Breakthroughs' (II.B.4) and then as a full section 'Consistency Models for Efficient Image Generation' (II.C). The two treatments are inconsistent: II.B.4 defines the consistency function via the probability-flow ODE, while II.C.1 uses an arbitrary-image formulation, and Eq. (24) in II.C.2 conflicts with the II.B.4 definition. This duplication suggests an unfinished editorial pass and should be resolved by merging the two accounts into one accurate presentation.
minor comments (8)
- [Abstract / Section I] The abstract states that the field has undergone a 'paradigm shift since 2021,' but the paper does not provide a dated timeline or evidence that 2021 is the correct boundary; please support this claim with a brief historical analysis or soften the phrasing.
- [References] Several references are self-citations to the authors' prior work (e.g., [20], [30], [47], [98], [102], [117]) and are used for efficiency, evaluation, and safety claims; please identify these as the authors' own work or replace them with independent sources.
- [Section II.B.3.b] The SDXL-Lightning paragraph does not cite [44] in the text; the reference appears only in the bibliography. Please add the citation and ensure that all equations and claims in that subsection are attributed to the source.
- [Section I.A] There are grammatical slips, e.g., 'Image generation models has been through significant changes' and 'These foundation models marks a significant milestone'; they should be corrected.
- [Section V.A] The sentence '[98] also proposed proposed CLIP-based metrics' contains a duplicated word 'proposed'.
- [Section I] The survey does not describe its literature selection methodology, so a reader cannot assess completeness or bias; I recommend adding a short scope and selection-criteria paragraph.
- [Section II.C.2] The claim that the time sampling strategy T 'focus[es] on regions where the model's outputs are most sensitive to noise' is vague and unreferenced; please define or cite the weighting scheme.
- [Tables I and II] Tables I and II summarize model series but do not include citations to the primary sources in the table captions or rows; please add them.
Circularity Check
No significant circularity; the survey's narrative rests on external literature, with only minor non-load-bearing self-citations.
full rationale
This is a survey paper, not a derivation paper. Its central claim—that image generation has undergone a paradigm shift since 2021 and that efficiency and interpretability challenges remain—is a synthesis of externally cited results (e.g., DDPM [3], LDM [16], SDXL [39], Consistency Models [49]), not a prediction derived from fitted parameters or from the authors' own prior work. I checked the seven self-citations ([20], [23], [30], [47], [98], [102], [117]): they support auxiliary points such as efficient architectures, UI/UX applications, LLM safety, multimodal embeddings, and CLIP-based metrics. None is load-bearing for the paradigm-shift narrative; removing them would not collapse the survey's main argument, which is independently supported by primary external sources. There is no uniqueness theorem imported from the authors' prior work, no ansatz smuggled in via self-citation, and no fitted quantity renamed as a prediction. The in-scope textual anomalies are accuracy and fidelity issues rather than circularity: Eq. (19) omits the conditioning c that appears in Eq. (18), Eq. (24) simplifies the Consistency Models objective away from the probability-flow ODE sampling used in [49], and Section II.B.3.b contains the verbatim first-person sentence 'Our models are available as both LoRA and full UNet weights,' indicating copied source text. These problems undermine the survey's reliability as a reproduction of primary equations, but they do not make any claim reduce to its own inputs. Accordingly, the circularity score is 1: minor self-citation exists, but the central claim has independent content and no step is circular by construction.
Assumptions & free parameters
assumptions (2)
- domain assumption The reproduced DDPM and LDM equations (Eqs. 2-20) accurately represent the original formulations of Ho et al. and Rombach et al.
- domain assumption The descriptions of model capabilities (SDXL-Lightning, Consistency Models, DALL-E series) faithfully reflect the claims of the cited source papers.
Cite this review
Pith. "Pith review of From Noise to Nuance: Advances in Deep Generative Image Models." pith.science (2026). https://pith.science/paper/ZO3RBTBD
@misc{pith2026241209656,
author = {Pith},
title = {Pith review of: From Noise to Nuance: Advances in Deep Generative Image Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZO3RBTBD}},
note = {Machine review of arXiv:2412.09656}
}
read the original abstract
Deep learning-based image generation has undergone a paradigm shift since 2021, marked by fundamental architectural breakthroughs and computational innovations. Through reviewing architectural innovations and empirical results, this paper analyzes the transition from traditional generative methods to advanced architectures, with focus on compute-efficient diffusion models and vision transformer architectures. We examine how recent developments in Stable Diffusion, DALL-E, and consistency models have redefined the capabilities and performance boundaries of image synthesis, while addressing persistent challenges in efficiency and quality. Our analysis focuses on the evolution of latent space representations, cross-attention mechanisms, and parameter-efficient training methodologies that enable accelerated inference under resource constraints. While more efficient training methods enable faster inference, advanced control mechanisms like ControlNet and regional attention systems have simultaneously improved generation precision and content customization. We investigate how enhanced multi-modal understanding and zero-shot generation capabilities are reshaping practical applications across industries. Our analysis demonstrates that despite remarkable advances in generation quality and computational efficiency, critical challenges remain in developing resource-conscious architectures and interpretable generation systems for industrial applications. The paper concludes by mapping promising research directions, including neural architecture optimization and explainable generation frameworks.
Figures
Reference graph
Works this paper leans on
-
[49]
Y . Song, P . Dhariwal, M. Chen, and I. Sutskever, “Consis tency models,” arXiv preprint arXiv:2303.01469 , 2023
arXiv 2023
-
[1]
Generative adversar ial net- works,
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Ward e-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversar ial net- works,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
-
[2]
Unsupervised repr esentation learning with deep convolutional generative adversarial n etworks,
A. Radford, L. Metz, and S. Chintala, “Unsupervised repr esentation learning with deep convolutional generative adversarial n etworks,” in International Conference on Learning Representations , 2016
2016
-
[3]
Denoising diffusion proba bilistic mod- els,
J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion proba bilistic mod- els,” in Advances in Neural Information Processing Systems , pp. 6840– 6851, 2020
2020
-
[4]
Score-based generative modeling through sto chastic differential equations,
Y . Song, J. Sohl-Dickstein, D. P . Kingma, A. Kumar, S. Erm on, and B. Poole, “Score-based generative modeling through sto chastic differential equations,” arXiv preprint arXiv:2011.13456 , 2020
arXiv 2011
-
[5]
Diffusion models beat gans on image synthesis,
P . Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Advances in neural information processing systems , vol. 34, pp. 8780–8794, 2021
2021
-
[6]
Imagenet: A large-scale hierarchical image database,
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei , “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition , pp. 248–255, IEEE, 2009
2009
-
[7]
Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs,
C. Schuhmann, R. V encu, R. Beaumont, R. Kaczmarczyk, C. M ullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki, “Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs, ” arXiv preprint arXiv:2111.02114, 2021
arXiv 2021
Show all 122 references
-
[8]
In-datacenter performance analysis of a tensor processing unit,
N. P . Jouppi, C. Y oung, N. Patil, D. Patterson, G. Agrawal , R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th annual international symposium on computer archit ecture, ...
2017
-
[9]
Scaling l aws for neural language models,
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Che ss, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling l aws for neural language models,” arXiv preprint arXiv:2001.08361 , 2020
2001 arXiv
-
[10]
Megatron-lm: Training multi-billion parameter lan guage models using model parallelism,
M. Shoeybi, M. Patwary, R. Puri, P . LeGresley, J. Casper , and B. Catan- zaro, “Megatron-lm: Training multi-billion parameter lan guage models using model parallelism,” arXiv preprint arXiv:1909.08053 , 2019
1909 arXiv
-
[11]
Scaling language models: Methods, analysis & insights from training gopher,
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Y oung, et al., “Scaling language models: Methods, analysis & insights from training gopher, ” arXiv preprint arXiv:2112.11446, 2021
2021 arXiv
-
[12]
Attention is all you need,
A. V aswani, “Attention is all you need,” Advances in Neural Informa- tion Processing Systems , 2017
2017
-
[13]
H ierarchi- cal text-conditional image generation with clip latents,
A. Ramesh, P . Dhariwal, A. Nichol, C. Chu, and M. Chen, “H ierarchi- cal text-conditional image generation with clip latents,” arXiv preprint arXiv:2204.06125, vol. 1, no. 2, p. 3, 2022
2022 arXiv
-
[14]
Photorealistic text-to-image diffusion models with dee p lan- guage understanding,
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Dent on, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salima ns, et al. , “Photorealistic text-to-image diffusion models with dee p lan- guage understanding,” Advances in neural information processing systems, vol...
2022
-
[15]
Zero-shot text-to-image generation,
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. V oss, A. Radford , M. Chen, and I. Sutskever, “Zero-shot text-to-image generation,” i n International conference on machine learning , pp. 8821–8831, Pmlr, 2021
2021
-
[16]
High-resolution image synthesis with latent diffusion mo dels,
R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Omm er, “High-resolution image synthesis with latent diffusion mo dels,” in Proceedings of the IEEE/CVF conference on computer vision a nd pattern recognition, pp. 10684–10695, 2022
2022
-
[17]
Improved denoising diffu sion prob- abilistic models,
A. Q. Nichol and P . Dhariwal, “Improved denoising diffu sion prob- abilistic models,” in International conference on machine learning , pp. 8162–8171, PMLR, 2021
2021
-
[18]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Ag arwal, G. Sastry, A. Askell, P . Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning , pp. 8748–8763, PMLR, 2021
2021
-
[19]
Alias-free generative adversarial networks,
T. Karras, M. Aittala, S. Laine, E. H¨ ark¨ onen, J. Hells ten, J. Lehtinen, and T. Aila, “Alias-free generative adversarial networks, ” Advances in neural information processing systems , vol. 34, pp. 852–863, 2021
2021
-
[20]
Surveying the mllm landscape: A meta- review of current surveys,
M. Li, K. Chen, Z. Bi, M. Liu, B. Peng, Q. Niu, J. Liu, J. Wan g, S. Zhang, X. Pan, et al. , “Surveying the mllm landscape: A meta- review of current surveys,” arXiv preprint arXiv:2409.18991 , 2024
2024
-
[21]
A survey of accelerator architectures for deep neural networks,
Y . Chen, Y . Xie, L. Song, F. Chen, and T. Tang, “A survey of accelerator architectures for deep neural networks,” Engineering, vol. 6, no. 3, pp. 264–274, 2020
2020
-
[22]
The age of generative ai and ai-generated everything,
H. Du, D. Niyato, J. Kang, Z. Xiong, P . Zhang, S. Cui, X. Sh en, S. Mao, Z. Han, A. Jamalipour, et al. , “The age of generative ai and ai-generated everything,” IEEE Network , 2024
2024
-
[23]
Llms and diffusion models in ui/ux: Advancing human-computer interaction and design,
L. Sun, M. Qin, and B. Peng, “Llms and diffusion models in ui/ux: Advancing human-computer interaction and design,” OSF Preprints , Oct 2024
2024
-
[24]
Lightweight g enerative adversarial networks for text-guided image manipulation,
B. Li, X. Qi, P . Torr, and T. Lukasiewicz, “Lightweight g enerative adversarial networks for text-guided image manipulation, ” Advances in Neural Information Processing Systems , vol. 33, pp. 22020–22031, 2020
2020
-
[25]
Dpm-sol ver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,
C. Lu, Y . Zhou, F. Bao, J. Chen, C. Li, and J. Zhu, “Dpm-sol ver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps,” Advances in Neural Information Processing Systems , vol. 35, pp. 5775–5787, 2022
2022
-
[26]
The emergence of deepfake technology: A review,
M. Westerlund, “The emergence of deepfake technology: A review,” Technology innovation management review , vol. 9, no. 11, 2019
2019
-
[27]
The cat and mouse game: The ongoing arms race between diffusion models and det ection methods,
L. Laurier, A. Giulietta, A. Octavia, and M. Cleti, “The cat and mouse game: The ongoing arms race between diffusion models and det ection methods,” arXiv preprint arXiv:2410.18866 , 2024
2024 arXiv
-
[28]
Easily a ccessible text-to-image generation amplifies demographic stereotyp es at large scale,
F. Bianchi, P . Kalluri, E. Durmus, F. Ladhak, M. Cheng, D . Nozza, T. Hashimoto, D. Jurafsky, J. Zou, and A. Caliskan, “Easily a ccessible text-to-image generation amplifies demographic stereotyp es at large scale,” in Proceedings of the 2023 ACM Conference on Fairness, Accoun...
2023
-
[29]
The creativity of text-to-image gen eration,
J. Oppenlaender, “The creativity of text-to-image gen eration,” in Pro- ceedings of the 25th international academic mindtrek confe rence, pp. 192–202, 2022
2022
-
[30]
Jailbreaking and mitigation of vuln erabilities in large language models,
B. Peng, Z. Bi, Q. Niu, M. Liu, P . Feng, T. Wang, L. K. Y an, Y . Wen, Y . Zhang, and C. H. Yin, “Jailbreaking and mitigation of vuln erabilities in large language models,” arXiv preprint arXiv:2410.15236 , 2024
2024 arXiv
-
[31]
Scalable diffusion models with t ransformers,
W. Peebles and S. Xie, “Scalable diffusion models with t ransformers,” in Proceedings of the IEEE/CVF International Conference on Co m- puter Vision, pp. 4195–4205, 2023
2023
-
[32]
Sd-dit: Unleashing the power of self-supervised discrimi nation in diffusion transformer,
R. Zhu, Y . Pan, Y . Li, T. Y ao, Z. Sun, T. Mei, and C. W. Chen, “Sd-dit: Unleashing the power of self-supervised discrimi nation in diffusion transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 8435–8445, 2024
2024
-
[33]
Scaling autoregres- sive models for content-rich text-to-image generation,
J. Y u, Y . Xu, J. Y . Koh, T. Luong, G. Baid, Z. Wang, V . V a- sudevan, A. Ku, Y . Y ang, B. K. Ayan, et al. , “Scaling autoregres- sive models for content-rich text-to-image generation,” arXiv preprint arXiv:2206.10789, vol. 2, no. 3, p. 5, 2022
2022 arXiv
-
[34]
Muse: Text-to-image generation via masked generative transform ers,
H. Chang, H. Zhang, J. Barber, A. Maschinot, J. Lezama, L . Jiang, M.-H. Y ang, K. Murphy, W. T. Freeman, M. Rubinstein, et al., “Muse: Text-to-image generation via masked generative transform ers,” arXiv preprint arXiv:2301.00704, 2023
2023 arXiv
-
[35]
Cogview: Mastering text-to-image generation via transformers,
M. Ding, Z. Y ang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin , X. Zou, Z. Shao, H. Y ang, et al., “Cogview: Mastering text-to-image generation via transformers,” Advances in neural information processing systems , vol. 34, pp. 19822–19835, 2021
2021
-
[36]
Cogview2: Faste r and better text-to-image generation via hierarchical transfo rmers,
M. Ding, W. Zheng, W. Hong, and J. Tang, “Cogview2: Faste r and better text-to-image generation via hierarchical transfo rmers,” Advances in Neural Information Processing Systems , vol. 35, pp. 16890–16902, 2022
2022
-
[37]
Cogview3: Finer and faster text-to-im age generation via relay diffusion,
W. Zheng, J. Teng, Z. Y ang, W. Wang, J. Chen, X. Gu, Y . Dong , M. Ding, and J. Tang, “Cogview3: Finer and faster text-to-im age generation via relay diffusion,” arXiv preprint arXiv:2403.05121, 2024
2024 arXiv
-
[38]
Auto-encoding variational bayes,
D. P . Kingma, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[39]
Sdxl: Improving latent diffusion models for high-resolution image synthesis,
D. Podell, Z. English, K. Lacey, A. Blattmann, T. Dockho rn, J. M¨ uller, J. Penna, and R. Rombach, “Sdxl: Improving latent diffusion models for high-resolution image synthesis,” arXiv preprint arXiv:2307.01952, 2023
2023 arXiv
-
[40]
Adve rsarial diffusion distillation,
A. Sauer, D. Lorenz, A. Blattmann, and R. Rombach, “Adve rsarial diffusion distillation,” in European Conference on Computer Vision , pp. 87–103, Springer, 2025
2025
-
[41]
Scaling rectified flow transformers for high-resolution image synthesis,
P . Esser, S. Kulal, A. Blattmann, R. Entezari, J. M¨ ulle r, H. Saini, Y . Levi, D. Lorenz, A. Sauer, F. Boesel, et al. , “Scaling rectified flow transformers for high-resolution image synthesis,” in F orty-first International Conference on Machine Learning , 2024
2024
-
[42]
Fast high-resolution image synthesis with latent ad versarial diffusion distillation,
A. Sauer, F. Boesel, T. Dockhorn, A. Blattmann, P . Esser , and R. Rom- bach, “Fast high-resolution image synthesis with latent ad versarial diffusion distillation,” arXiv preprint arXiv:2403.12015 , 2024
2024 arXiv
-
[43]
Progressive distillation for fa st sampling of diffusion models,
T. Salimans and J. Ho, “Progressive distillation for fa st sampling of diffusion models,” arXiv preprint arXiv:2202.00512 , 2022
2022 arXiv
-
[44]
Sdxl-lightning: Progress ive adversarial diffusion distillation,
S. Lin, A. Wang, and X. Y ang, “Sdxl-lightning: Progress ive adversarial diffusion distillation,” arXiv preprint arXiv:2402.13929 , 2024
2024 arXiv
-
[45]
Google research imagen 2 update
Google Research, “Google research imagen 2 update.” https://blog.google/technology/ai/google-imagen-2/. Accessed: 2024-11-10
2024
-
[46]
Google deepmind imagen 3 update
Google DeepMind, “Google deepmind imagen 3 update.” https://deepmind.com/technologies/imagen-3/ . Accesse d: 2024-11-10
2024
-
[47]
From word vectors to multimodal embeddings: Techniques, applications, and future directions for large language models,
C. Zhang, B. Peng, X. Sun, Q. Niu, J. Liu, K. Chen, M. Li, P . Feng, Z. Bi, M. Liu, et al. , “From word vectors to multimodal embeddings: Techniques, applications, and future directions for large language models,” arXiv preprint arXiv:2411.05036 , 2024
2024
-
[48]
Dall-e 3: Ai that can create images from text wi th improved prompt following
OpenAI, “Dall-e 3: Ai that can create images from text wi th improved prompt following.” https://openai.com/dall-e-3. Access ed: 2024-11-10
2024
-
[50]
Consi stency models made easy,
Z. Geng, A. Pokle, W. Luo, J. Lin, and J. Z. Kolter, “Consi stency models made easy,” arXiv preprint arXiv:2406.14548 , 2024
2024 arXiv
-
[51]
Q-diffusion: Quantizing diffusion models,
X. Li, Y . Liu, L. Lian, H. Y ang, Z. Dong, D. Kang, S. Zhang, and K. Keutzer, “Q-diffusion: Quantizing diffusion models,” i n Proceed- ings of the IEEE/CVF International Conference on Computer V ision, pp. 17535–17545, 2023
2023
-
[52]
Post-traini ng quantiza- tion on diffusion models,
Y . Shang, Z. Y uan, B. Xie, B. Wu, and Y . Y an, “Post-traini ng quantiza- tion on diffusion models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 1972–1981, 2023
1972
-
[53]
Efficientdm: Efficient quantization-aware fine-tuning of low-bit diffusion model s,
Y . He, J. Liu, W. Wu, H. Zhou, and B. Zhuang, “Efficientdm: Efficient quantization-aware fine-tuning of low-bit diffusion model s,” arXiv preprint arXiv:2310.03270, 2023
2023 arXiv
-
[54]
Efficientqat: Efficient quantization-aware traini ng for large language models,
M. Chen, W. Shao, P . Xu, J. Wang, P . Gao, K. Zhang, Y . Qiao, and P . Luo, “Efficientqat: Efficient quantization-aware traini ng for large language models,” arXiv preprint arXiv:2407.11062 , 2024
2024 arXiv
-
[55]
Mixed precision training,
P . Micikevicius, S. Narang, J. Alben, G. Diamos, E. Else n, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. V enkatesh, et al. , “Mixed precision training,” arXiv preprint arXiv:1710.03740 , 2017
2017 arXiv
-
[56]
Adaptive quantization for deep neural network,
Y . Zhou, S.-M. Moosavi-Dezfooli, N.-M. Cheung, and P . F rossard, “Adaptive quantization for deep neural network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, 2018
2018
-
[57]
Parameter-efficie nt transfer learning for nlp,
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. D e Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly, “Parameter-efficie nt transfer learning for nlp,” in International conference on machine learning , pp. 2790–2799, PMLR, 2019
2019
-
[58]
Lora: Low-rank adaptation of large language mo dels,
E. J. Hu, Y . Shen, P . Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language mo dels,” arXiv preprint arXiv:2106.09685 , 2021
2021 arXiv
-
[59]
Qlora: Efficient finetuning of quantized llms,
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoye r, “Qlora: Efficient finetuning of quantized llms,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[60]
Diffstyler: Diffusion-based localized image s tyle transfer,
S. Li, “Diffstyler: Diffusion-based localized image s tyle transfer,” arXiv preprint arXiv:2403.18461, 2024
2024
-
[61]
Lcm-lora: A universal stable-diffusion acceleration module,
S. Luo, Y . Tan, S. Patil, D. Gu, P . von Platen, A. Passos, L . Huang, J. Li, and H. Zhao, “Lcm-lora: A universal stable-diffusion acceleration module,” arXiv preprint arXiv:2311.05556 , 2023
2023 arXiv
-
[62]
T2i- adapter: Learning adapters to dig out more controllable abi lity for text- to-image diffusion models,
C. Mou, X. Wang, L. Xie, Y . Wu, J. Zhang, Z. Qi, and Y . Shan, “T2i- adapter: Learning adapters to dig out more controllable abi lity for text- to-image diffusion models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, pp. 4296–4304, 2024
2024
-
[63]
Ip-adapter: Text compatible image prompt adapter for text-to-image diffusi on models,
H. Y e, J. Zhang, S. Liu, X. Han, and W. Y ang, “Ip-adapter: Text compatible image prompt adapter for text-to-image diffusi on models,” arXiv preprint arXiv:2308.06721 , 2023
2023 arXiv
-
[64]
Optimizing prompts f or text-to-image generation,
Y . Hao, Z. Chi, L. Dong, and F. Wei, “Optimizing prompts f or text-to-image generation,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[65]
Sparsity-and hybrid ity-inspired visual parameter-efficient fine-tuning for medical diagnos is,
M. Liu, L. Xu, S. Liu, and J. Zhang, “Sparsity-and hybrid ity-inspired visual parameter-efficient fine-tuning for medical diagnos is,” in In- ternational Conference on Medical Image Computing and Comp uter- Assisted Intervention, pp. 627–637, Springer, 2024
2024
-
[66]
Zero: M emory optimizations toward training trillion parameter models,
S. Rajbhandari, J. Rasley, O. Ruwase, and Y . He, “Zero: M emory optimizations toward training trillion parameter models, ” in SC20: In- ternational Conference for High Performance Computing, Ne tworking, Storage and Analysis , pp. 1–16, IEEE, 2020
2020
-
[67]
Efficient large-scale language model training on gpu clus ters using megatron-lm,
D. Narayanan, M. Shoeybi, J. Casper, P . LeGresley, M. Pa twary, V . Korthikanti, D. V ainbrand, P . Kashinkunti, J. Bernauer,B. Catanzaro, et al. , “Efficient large-scale language model training on gpu clus ters using megatron-lm,” in Proceedings of the International Conferenc...
2021
-
[68]
xdit: an infer ence engine for diffusion transformers (dits) with massive para llelism,
J. Fang, J. Pan, X. Sun, A. Li, and J. Wang, “xdit: an infer ence engine for diffusion transformers (dits) with massive para llelism,” arXiv preprint arXiv:2411.01738, 2024
2024 arXiv
-
[69]
Deep generative model for imag e inpainting with local binary pattern learning and spatial a ttention,
H. Wu, J. Zhou, and Y . Li, “Deep generative model for imag e inpainting with local binary pattern learning and spatial a ttention,” IEEE Transactions on Multimedia , vol. 24, pp. 4016–4027, 2021
2021
-
[70]
Gen erative image inpainting with contextual attention,
J. Y u, Z. Lin, J. Y ang, X. Shen, X. Lu, and T. S. Huang, “Gen erative image inpainting with contextual attention,” in Proceedings of the IEEE conference on computer vision and pattern recognition , pp. 5505–5514, 2018
2018
-
[71]
Grig: Few-shot generative residual image inpaint ing,
W. Lu, X. Jiang, X. Jin, Y .-L. Y ang, M. Gong, T. Wang, K. Sh i, and H. Zhao, “Grig: Few-shot generative residual image inpaint ing,” arXiv preprint arXiv:2304.12035, 2023
2023 arXiv
-
[72]
Latentpaint : Image inpainting in latent space with diffusion models,
C. Corneanu, R. Gadde, and A. M. Martinez, “Latentpaint : Image inpainting in latent space with diffusion models,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 4334–4343, 2024
2024
-
[73]
Repaint: Inpainting using denoising diffusio n probabilis- tic models,
A. Lugmayr, M. Danelljan, A. Romero, F. Y u, R. Timofte, a nd L. V an Gool, “Repaint: Inpainting using denoising diffusio n probabilis- tic models,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pp. 11461–11471, 2022
2022
-
[74]
Gradpaint: Gradi ent-guided inpainting with diffusion models,
A. Grechka, G. Couairon, and M. Cord, “Gradpaint: Gradi ent-guided inpainting with diffusion models,” Computer Vision and Image Under- standing, vol. 240, p. 103928, 2024
2024
-
[75]
Inout: Diverse image outpainting via gan inversio n,
Y .-C. Cheng, C. H. Lin, H.-Y . Lee, J. Ren, S. Tulyakov, an d M.- H. Y ang, “Inout: Diverse image outpainting via gan inversio n,” in Proceedings of the IEEE/CVF Conference on Computer Vision a nd Pattern Recognition, pp. 11431–11440, 2022
2022
-
[76]
Painting outside as inside: Edge guided image outpainting via bidirec- tional rearrangement with progressive step learning,
K. Kim, Y . Y un, K.-W. Kang, K. Kong, S. Lee, and S.-J. Kang , “Painting outside as inside: Edge guided image outpainting via bidirec- tional rearrangement with progressive step learning,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision, pp...
2021
-
[77]
Towards reliable image outpainting: Learning structure-aware multimodal fusion with depth guidance,
L. Zhang, C. Lin, K. Liao, and Y . Zhao, “Towards reliable image outpainting: Learning structure-aware multimodal fusion with depth guidance,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pp. 1–5, IEEE, 2023
2023
-
[78]
Em ergent correspondence from image diffusion,
L. Tang, M. Jia, Q. Wang, C. P . Phoo, and B. Hariharan, “Em ergent correspondence from image diffusion,” Advances in Neural Information Processing Systems, vol. 36, pp. 1363–1389, 2023
2023
-
[79]
Zero-1-to-3: Zero-shot one image to 3d object ,
R. Liu, R. Wu, B. V an Hoorick, P . Tokmakov, S. Zakharov, a nd C. V ondrick, “Zero-1-to-3: Zero-shot one image to 3d object ,” in Proceedings of the IEEE/CVF international conference on co mputer vision, pp. 9298–9309, 2023
2023
-
[80]
Wonder3d: Single image to 3d using cross-domain diffusion,
X. Long, Y .-C. Guo, C. Lin, Y . Liu, Z. Dou, L. Liu, Y . Ma, S. -H. Zhang, M. Habermann, C. Theobalt, et al. , “Wonder3d: Single image to 3d using cross-domain diffusion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 9970– 9980, 2024
2024
-
[81]
Make-it-3d: High-fidelity 3d creation from a single image w ith diffu- sion prior,
J. Tang, T. Wang, B. Zhang, T. Zhang, R. Yi, L. Ma, and D. Ch en, “Make-it-3d: High-fidelity 3d creation from a single image w ith diffu- sion prior,” in Proceedings of the IEEE/CVF international conference on computer vision , pp. 22819–22829, 2023
2023
-
[82]
Adding conditional c ontrol to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional c ontrol to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , pp. 3836–3847, 2023
2023
-
[83]
Controlnet- xs: Designing an efficient and effective architecture for controlling tex t-to-image diffusion models,
D. Zavadski, J.-F. Feiden, and C. Rother, “Controlnet- xs: Designing an efficient and effective architecture for controlling tex t-to-image diffusion models,” arXiv preprint arXiv:2312.06573 , 2023
2023 arXiv
-
[84]
Uni-controlnet: All-in-one control to text-to-ima ge diffusion models,
S. Zhao, D. Chen, Y .-C. Chen, J. Bao, S. Hao, L. Y uan, and K .-Y . K. Wong, “Uni-controlnet: All-in-one control to text-to-ima ge diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[85]
Styledrop: Text-to-image generation in any style,
K. Sohn, N. Ruiz, K. Lee, D. C. Chin, I. Blok, H. Chang, J. B arber, L. Jiang, G. Entis, Y . Li, et al. , “Styledrop: Text-to-image generation in any style,” arXiv preprint arXiv:2306.00983 , 2023
2023 arXiv
-
[86]
Any-to-any style transfer: M aking picasso and da vinci collaborate,
S. Liu, J. Y e, and X. Wang, “Any-to-any style transfer: M aking picasso and da vinci collaborate,” arXiv preprint arXiv:2304.09728 , 2023
2023 arXiv
-
[87]
Clipstyler: Image style transfer w ith a single text condition,
G. Kwon and J. C. Y e, “Clipstyler: Image style transfer w ith a single text condition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 18062–18071, 2022
2022
-
[88]
Styleshot: A snapshot on any style,
J. Gao, Y . Liu, Y . Sun, Y . Tang, Y . Zeng, K. Chen, and C. Zha o, “Styleshot: A snapshot on any style,” arXiv preprint arXiv:2407.01414, 2024
2024 arXiv
-
[89]
Metrop olis theorem and its applications in single image detail enhance ment,
H. Jiang, M. Asad, J. Liu, H. Zhang, and D. Cheng, “Metrop olis theorem and its applications in single image detail enhance ment,” arXiv preprint arXiv:2302.09762, 2023
2023 arXiv
-
[90]
Multi-scale image decomposition using a l ocal statistical edge model,
K.-M. Wong, “Multi-scale image decomposition using a l ocal statistical edge model,” in 2021 IEEE 7th International Conference on Virtual Reality (ICVR) , pp. 10–18, IEEE, 2021
2021
-
[91]
Crnet: A detail-preserving network for unified i mage restoration and enhancement task,
K. Y ang, T. Hu, K. Dai, G. Chen, Y . Cao, W. Dong, P . Wu, Y . Zh ang, and Q. Y an, “Crnet: A detail-preserving network for unified i mage restoration and enhancement task,” arXiv preprint arXiv:2404.14132 , 2024
2024 arXiv
-
[92]
Ecaformer: Low-light image enhancement using cross attention,
Y . Ruan, H. Ma, W. Li, and X. Wang, “Ecaformer: Low-light image enhancement using cross attention,” arXiv preprint arXiv:2406.13281 , 2024
2024 arXiv
-
[93]
Gans trained by a two time-scale update rule converge to a lo cal nash equilibrium,
M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter, “Gans trained by a two time-scale update rule converge to a lo cal nash equilibrium,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[94]
On aliased resizing and surprising subtleties in gan evaluation,
G. Parmar, R. Zhang, and J.-Y . Zhu, “On aliased resizing and surprising subtleties in gan evaluation,” in CVPR, 2022
2022
-
[95]
Demysti- fying mmd gans,
M. Bi´ nkowski, D. J. Sutherland, M. Arbel, and A. Gretto n, “Demysti- fying mmd gans,” arXiv preprint arXiv:1801.01401 , 2018
2018 arXiv
-
[96]
The unreasonable effectiveness of deep features as a perceptua l metric,
R. Zhang, P . Isola, A. A. Efros, E. Shechtman, and O. Wang , “The unreasonable effectiveness of deep features as a perceptua l metric,” in Proceedings of the IEEE conference on computer vision and pa ttern recognition, pp. 586–595, 2018
2018
-
[97]
Imagereward: Learning and evaluating human preferences f or text- to-image generation,
J. Xu, X. Liu, Y . Wu, Y . Tong, Q. Li, M. Ding, J. Tang, and Y . Dong, “Imagereward: Learning and evaluating human preferences f or text- to-image generation,” Advances in Neural Information Processing Systems, vol. 36, 2024
2024
-
[98]
Clip-agiqa: Boos ting the performance of ai-generated image quality assessment with clip,
Z. Tang, Z. Wang, B. Peng, and J. Dong, “Clip-agiqa: Boos ting the performance of ai-generated image quality assessment with clip,” arXiv preprint arXiv:2408.15098, 2024
2024 arXiv
-
[99]
Human evaluation of text-to-image models on a multi-task benchma rk,
V . Petsiuk, A. E. Siemenn, S. Surbehera, Z. Chin, K. Tyse r, G. Hunter, A. Raghavan, Y . Hicke, B. A. Plummer, O. Kerret, et al. , “Human evaluation of text-to-image models on a multi-task benchma rk,” arXiv preprint arXiv:2211.12112, 2022
2022 arXiv
-
[100]
Hum an-ai co- creation: evaluating the impact of large-scale text-to-im age generative models on the creative process,
T. Turchi, S. Carta, L. Ambrosini, and A. Malizia, “Hum an-ai co- creation: evaluating the impact of large-scale text-to-im age generative models on the creative process,” in International Symposium on End User Development, pp. 35–51, Springer, 2023
2023
-
[101]
Evaluating text-to-visual generation with im age-to-text generation,
Z. Lin, D. Pathak, B. Li, J. Li, X. Xia, G. Neubig, P . Zhan g, and D. Ramanan, “Evaluating text-to-visual generation with im age-to-text generation,” in European Conference on Computer Vision, pp. 366–384, Springer, 2025
2025
-
[102]
Llms and diffusion models i n ui/ux: Advancing human-computer interaction and design,
L. Sun, M. Qin, and B. Peng, “Llms and diffusion models i n ui/ux: Advancing human-computer interaction and design,” OSF Preprints , Oct 2024
2024
-
[103]
Clip- score: A reference-free evaluation metric for image captio ning,
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y . Cho i, “Clip- score: A reference-free evaluation metric for image captio ning,” arXiv preprint arXiv:2104.08718, 2021
2021 arXiv
-
[104]
Enhanc- ing semantic fidelity in text-to-image synthesis: Attentio n regulation in diffusion models,
Y . Zhang, T. T. Tzun, L. W. Hern, T. Sim, and K. Kawaguchi , “Enhanc- ing semantic fidelity in text-to-image synthesis: Attentio n regulation in diffusion models,” arXiv preprint arXiv:2403.06381 , 2024
2024 arXiv
-
[105]
Hypernymy understan ding evalu- ation of text-to-image models via wordnet hierarchy,
A. Baryshnikov and M. Ryabinin, “Hypernymy understan ding evalu- ation of text-to-image models via wordnet hierarchy,” arXiv preprint arXiv:2310.09247, 2023
2023 arXiv
-
[106]
Memory-efficient fine-tuni ng for quantized diffusion model,
H. Ryu, S. Lim, and H. Shim, “Memory-efficient fine-tuni ng for quantized diffusion model,” in European Conference on Computer Vision, pp. 356–372, Springer, 2025
2025
-
[107]
Mlperf inference benchmark,
V . J. Reddi, C. Cheng, D. Kanter, P . Mattson, G. Schmuel ling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou, et al., “Mlperf inference benchmark,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) , pp. 446–459, IEEE, 2020
2020
-
[108]
Carbon emission s and large neural network training,
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Mungu ia, D. Rothchild, D. So, M. Texier, and J. Dean, “Carbon emission s and large neural network training,” arXiv preprint arXiv:2104.10350, 2021
2021 arXiv
-
[109]
The c hallenges of image generation models in generating multi-component i mages,
T. Y . Foong, S. Kotyan, P . Y . Mao, and D. V . V argas, “The c hallenges of image generation models in generating multi-component i mages,” arXiv preprint arXiv:2311.13620 , 2023
2023 arXiv
-
[110]
Prefpa int: Aligning image inpainting diffusion model with human prefe rence,
K. Liu, Z. Zhu, C. Li, H. Liu, H. Zeng, and J. Hou, “Prefpa int: Aligning image inpainting diffusion model with human prefe rence,” arXiv preprint arXiv:2410.21966 , 2024
2024 arXiv
-
[111]
Playground v2. 5: Three insights towards enhancing aesthe tic quality in text-to-image generation,
D. Li, A. Kamko, E. Akhgari, A. Sabet, L. Xu, and S. Doshi , “Playground v2. 5: Three insights towards enhancing aesthe tic quality in text-to-image generation,” arXiv preprint arXiv:2402.17245 , 2024
2024 arXiv
-
[112]
Ai gciqa2023: A large-scale image quality assessment database for ai gene rated images: from the perspectives of quality, authenticity and correspon- dence,
J. Wang, H. Duan, J. Liu, S. Chen, X. Min, and G. Zhai, “Ai gciqa2023: A large-scale image quality assessment database for ai gene rated images: from the perspectives of quality, authenticity and correspon- dence,” in CAAI International Conference on Artificial Intelligence , p...
2023
-
[113]
Emu: Enhancing image generation models using photogenic needles in a haystack,
X. Dai, J. Hou, C.-Y . Ma, S. Tsai, J. Wang, R. Wang, P . Zha ng, S. V andenhende, X. Wang, A. Dubey, et al. , “Emu: Enhancing image generation models using photogenic needles in a haystack,” arXiv preprint arXiv:2309.15807, 2023
2023 arXiv
-
[114]
Promptmagician: Interactive prompt engineering for text- to-image creation,
Y . Feng, X. Wang, K. K. Wong, S. Wang, Y . Lu, M. Zhu, B. Wan g, and W. Chen, “Promptmagician: Interactive prompt engineering for text- to-image creation,” IEEE Transactions on Visualization and Computer Graphics, 2023
2023
-
[115]
Beau ti- fulprompt: Towards automatic prompt engineering for text- to-image synthesis,
T. Cao, C. Wang, B. Liu, Z. Wu, J. Zhu, and J. Huang, “Beau ti- fulprompt: Towards automatic prompt engineering for text- to-image synthesis,” arXiv preprint arXiv:2311.06752 , 2023
2023 arXiv
-
[116]
Neuroprompts: An a daptive framework to optimize prompts for text-to-image generatio n,
S. Rosenman, V . Lal, and P . Howard, “Neuroprompts: An a daptive framework to optimize prompts for text-to-image generatio n,” arXiv preprint arXiv:2311.12229, 2023
2023 arXiv
-
[117]
Securing large language models: Addressing bias, misinformation, a nd prompt attacks,
B. Peng, K. Chen, M. Li, P . Feng, Z. Bi, J. Liu, and Q. Niu, “Securing large language models: Addressing bias, misinformation, a nd prompt attacks,” arXiv preprint arXiv:2409.08087 , 2024
2024
-
[118]
Sneakypro mpt: Jail- breaking text-to-image generative models,
Y . Y ang, B. Hui, H. Y uan, N. Gong, and Y . Cao, “Sneakypro mpt: Jail- breaking text-to-image generative models,” in 2024 IEEE symposium on security and privacy (SP) , pp. 897–912, IEEE, 2024
2024
-
[119]
Guardt2i: Defend- ing text-to-image models from adversarial prompts,
Y . Y ang, R. Gao, X. Y ang, J. Zhong, and Q. Xu, “Guardt2i: Defend- ing text-to-image models from adversarial prompts,” arXiv preprint arXiv:2403.01446, 2024
2024 arXiv
-
[120]
Unsafebench: Benchmarking image safety classifiers on rea l-world and ai-generated images,
Y . Qu, X. Shen, Y . Wu, M. Backes, S. Zannettou, and Y . Zha ng, “Unsafebench: Benchmarking image safety classifiers on rea l-world and ai-generated images,” arXiv preprint arXiv:2405.03486 , 2024
2024 arXiv
-
[121]
Turbovit: Generating fast vi- sion transformers via generative architecture search,
A. Wong, S. Abbasi, and S. Nair, “Turbovit: Generating fast vi- sion transformers via generative architecture search,” arXiv preprint arXiv:2308.11421, 2023
2023 arXiv
-
[122]
Ultra-high-resolution image synthesis wit h pyramid diffusion model,
J. Y ang, “Ultra-high-resolution image synthesis wit h pyramid diffusion model,” arXiv preprint arXiv:2403.12915 , 2024
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.