REVIEW 43 references
Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generative AI presents transformative potential across various domains, from creative arts to scientific visualization. However, the utility of AI-generated imagery is often compromised by visual flaws, including anatomical inaccuracies, improper object placements, and misplaced textual elements. These imperfections pose significant challenges for practical applications. To overcome these limitations, we introduce \textit{Yuan}, a novel framework that autonomously corrects visual imperfections in text-to-image synthesis. \textit{Yuan} uniquely conditions on both the textual prompt and the segmented image, generating precise masks that identify areas in need of refinement without requiring manual intervention -- a common constraint in previous methodologies. Following the automated masking process, an advanced inpainting module seamlessly integrates contextually coherent content into the identified regions, preserving the integrity and fidelity of the original image and associated text prompts. Through extensive experimentation on publicly available datasets such as ImageNet100 and Stanford Dogs, along with a custom-generated dataset, \textit{Yuan} demonstrated superior performance in eliminating visual imperfections. Our approach consistently achieved higher scores in quantitative metrics, including NIQE, BRISQUE, and PI, alongside favorable qualitative evaluations. These results underscore \textit{Yuan}'s potential to significantly enhance the quality and applicability of AI-generated images across diverse fields.
Reference graph
Works this paper leans on
-
[1]
Borch, C.; and Hee Min, B. 2022. Toward a sociology of machine learning explainability: Human--machine interaction in deep neural network-based automated trading. Big Data & Society, 9(2): 20539517221111361
2022
-
[2]
Brock, A.; Donahue, J.; and Simonyan, K. 2018. Large scale GAN training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096
arXiv 2018
-
[3]
Cetinic, E.; and She, J. 2022. Understanding and creating art with AI: Review and outlook. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 18(2): 1--22
2022
-
[4]
Chen, X.; Wang, W.; Bender, C.; Ding, Y.; Jia, R.; Li, B.; and Song, D. 2021. Refit: a unified watermark removal framework for deep learning systems with limited data. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, 321--335
2021
-
[5]
Gafni, O.; Polyak, A.; Ashual, O.; Sheynin, S.; Parikh, D.; and Taigman, Y. 2022. Make-a-scene: Scene-based text-to-image generation with human priors. In European Conference on Computer Vision, 89--106. Springer
2022
-
[6]
H.; Chechik, G.; and Cohen-Or, D
Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618
arXiv 2022
-
[7]
Gandikota, R.; Orgad, H.; Belinkov, Y.; Materzy \'n ska, J.; and Bau, D. 2024. Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 5111--5120
2024
-
[8]
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014. Generative adversarial nets. Advances in neural information processing systems, 27
2014
Show all 43 references
-
[9]
Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33: 6840--6851
2020
-
[10]
Hong, S.; Lee, J.; and Woo, S. S. 2024. All but One: Surgical Concept Erasing with Model Preservation in Text-to-Image Diffusion Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 21143--21151
2024
-
[11]
Karras, T.; Laine, S.; and Aila, T. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4401--4410
2019
-
[12]
Khosla, A.; Jayadevaprakash, N.; Yao, B.; and Li, F.-F. 2011. Novel dataset for fine-grained image categorization: Stanford dogs. In CVPR-W, volume 2. Citeseer
2011
-
[13]
P.; and Welling, M
Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114
2013 arXiv
-
[14]
Le, H.; and Samaras, D. 2019. Shadow removal via shadow image decomposition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8578--8587
2019
-
[15]
S.; Hou, Q.; Wang, Y.; and Yang, J
Li, S.; van de Weijer, J.; Hu, T.; Khan, F. S.; Hou, Q.; Wang, Y.; and Yang, J. 2024. Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models. arXiv preprint arXiv:2402.05375
2024 arXiv
-
[16]
Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023. GLIGEN: Open-set grounded text-to-image generation. In CVPR, 22511--22521
2023
-
[17]
Liu, Z.; Yin, H.; Wu, X.; Wu, Z.; Mi, Y.; and Wang, S. 2021. From shadow generation to shadow removal. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4927--4936
2021
-
[18]
Luo, X.; Li, Y.; Chang, H.; Liu, C.; Milanfar, P.; and Yang, F. 2023. DVMark: a deep multiscale framework for video watermarking. IEEE Transactions on Image Processing
2023
-
[19]
completely blind
Mittal, A.; Soundararajan, R.; and Bovik, A. C. 2012. Making a “completely blind” image quality analyzer. IEEE Signal processing letters, 20(3): 209--212
2012
-
[20]
O.; Cohen, N.; Mittal, G.; and Hegde, C
Pham, M.; Marshall, K. O.; Cohen, N.; Mittal, G.; and Hegde, C. 2023. Circumventing concept erasure methods for text-to-image generative models. In The Twelfth International Conference on Learning Representations
2023
-
[21]
O.; Hegde, C.; and Cohen, N
Pham, M.; Marshall, K. O.; Hegde, C.; and Cohen, N. 2024. Robust Concept Erasure Using Task Vectors. arXiv preprint arXiv:2404.03631
2024 arXiv
-
[22]
Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2): 3
2022 arXiv
-
[23]
Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021. Zero-shot text-to-image generation. In International conference on machine learning, 8821--8831. Pmlr
2021
-
[24]
Ray, A.; and Roy, S. 2020. Recent trends in image watermarking techniques for copyright protection: a survey. International Journal of Multimedia Information Retrieval, 9(4): 249--270
2020
-
[25]
C.; and Fei-Fei, L
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015. ImageNet Large Scale Visual Recognition Challenge . International Journal of Computer Vision (IJCV), 115(3): 211--252
2015
-
[26]
L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al
Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing syste...
2022
-
[27]
Singh, N.; Jain, M.; and Sharma, S. 2013. A survey of digital watermarking techniques. International Journal of Modern Communication Technologies and Research, 1(6): 265852
2013
-
[28]
P.; Kumar, A.; Ermon, S.; and Poole, B
Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456
2020 arXiv
-
[29]
Sun, J.; Wang, X.; Shi, Y.; Wang, L.; Wang, J.; and Liu, Y. 2022. Ide-3d: Interactive disentangled editing for high-resolution 3d-aware portrait synthesis. ACM Transactions on Graphics (ToG), 41(6): 1--10
2022
-
[30]
Tsai, Y.-L.; Hsu, C.-Y.; Xie, C.; Lin, C.-H.; Chen, J.-Y.; Li, B.; Chen, P.-Y.; Yu, C.-M.; and Huang, C.-Y. 2023. Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models? arXiv preprint arXiv:2310.10012
2023 arXiv
-
[31]
Wang, T.; Zhang, Y.; Qi, S.; Zhao, R.; Xia, Z.; and Weng, J. 2023. Security and privacy on generative data in aigc: A survey. arXiv preprint arXiv:2309.09435
2023 arXiv
-
[32]
Wang, Y.; Huang, W.; Sun, F.; Xu, T.; Rong, Y.; and Huang, J. 2020. Deep multimodal fusion by channel exchanging. Advances in neural information processing systems, 33: 4835--4845
2020
-
[33]
C.; Sheikh, H
Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4): 600--612
2004
-
[34]
Wu, J.; Gan, W.; Chen, Z.; Wan, S.; and Lin, H. 2023. Ai-generated content (aigc): A survey. arXiv preprint arXiv:2304.06632
2023 arXiv
-
[35]
Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; and Yuan, L. 2024. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4818--4829
2024
-
[36]
Xiong, T.; Wu, Y.; Xie, E.; Li, Z.; and Liu, X. 2024. Editing Massive Concepts in Text-to-Image Diffusion Models. arXiv preprint arXiv:2403.13807
2024
-
[37]
Xu, T.; Zhang, P.; Huang, Q.; Zhang, H.; Gan, Z.; Huang, X.; and He, X. 2018. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1316--1324
2018
-
[38]
Yang, Y.; Mu, K.; and Deng, R. H. 2022. Lightweight privacy-preserving GAN framework for model training and image synthesis. IEEE Transactions on Information Forensics and Security, 17: 1083--1098
2022
-
[39]
Z.; Yang, X.; Lin, L.; Lee, M
Yao, B. Z.; Yang, X.; Lin, L.; Lee, M. W.; and Zhu, S.-C. 2010. I2t: Image parsing to text description. Proceedings of the IEEE, 98(8): 1485--1508
2010
-
[40]
Zhang, Z.; and Schomaker, L. 2021. Dtgan: Dual attention generative adversarial networks for text-to-image generation. In 2021 international joint conference on neural networks (IJCNN), 1--8. IEEE
2021
-
[41]
Zhao, M.; Zhang, L.; Zheng, T.; Kong, Y.; and Yin, B. 2024. Separable Multi-Concept Erasure from Diffusion Models. arXiv preprint arXiv:2402.05947
2024 arXiv
-
[42]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...
-
[43]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Discussion (0). Continue with ORCID to comment.