Pith. sign in

REVIEW 43 references

Yuan: Yielding Unblemished Aesthetics Through A Unified Network for Visual Imperfections Removal in Generated Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.08505 v1 pith:VVOPTKME submitted 2025-01-15 cs.CV eess.IV

classification cs.CVeess.IV
keywords yuanimperfectionstextitvisualacrossai-generatedimageimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative AI presents transformative potential across various domains, from creative arts to scientific visualization. However, the utility of AI-generated imagery is often compromised by visual flaws, including anatomical inaccuracies, improper object placements, and misplaced textual elements. These imperfections pose significant challenges for practical applications. To overcome these limitations, we introduce \textit{Yuan}, a novel framework that autonomously corrects visual imperfections in text-to-image synthesis. \textit{Yuan} uniquely conditions on both the textual prompt and the segmented image, generating precise masks that identify areas in need of refinement without requiring manual intervention -- a common constraint in previous methodologies. Following the automated masking process, an advanced inpainting module seamlessly integrates contextually coherent content into the identified regions, preserving the integrity and fidelity of the original image and associated text prompts. Through extensive experimentation on publicly available datasets such as ImageNet100 and Stanford Dogs, along with a custom-generated dataset, \textit{Yuan} demonstrated superior performance in eliminating visual imperfections. Our approach consistently achieved higher scores in quantitative metrics, including NIQE, BRISQUE, and PI, alongside favorable qualitative evaluations. These results underscore \textit{Yuan}'s potential to significantly enhance the quality and applicability of AI-generated images across diverse fields.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

43 extracted references · 1 canonical work pages

  1. [1]

    Borch, C.; and Hee Min, B. 2022. Toward a sociology of machine learning explainability: Human--machine interaction in deep neural network-based automated trading. Big Data & Society, 9(2): 20539517221111361

  2. [2]

    Brock, A.; Donahue, J.; and Simonyan, K. 2018. Large scale GAN training for high fidelity natural image synthesis. arXiv preprint arXiv:1809.11096

  3. [3]

    Cetinic, E.; and She, J. 2022. Understanding and creating art with AI: Review and outlook. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 18(2): 1--22

  4. [4]

    Chen, X.; Wang, W.; Bender, C.; Ding, Y.; Jia, R.; Li, B.; and Song, D. 2021. Refit: a unified watermark removal framework for deep learning systems with limited data. In Proceedings of the 2021 ACM Asia Conference on Computer and Communications Security, 321--335

  5. [5]

    Gafni, O.; Polyak, A.; Ashual, O.; Sheynin, S.; Parikh, D.; and Taigman, Y. 2022. Make-a-scene: Scene-based text-to-image generation with human priors. In European Conference on Computer Vision, 89--106. Springer

  6. [6]

    H.; Chechik, G.; and Cohen-Or, D

    Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A. H.; Chechik, G.; and Cohen-Or, D. 2022. An image is worth one word: Personalizing text-to-image generation using textual inversion. arXiv preprint arXiv:2208.01618

  7. [7]

    Gandikota, R.; Orgad, H.; Belinkov, Y.; Materzy \'n ska, J.; and Bau, D. 2024. Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 5111--5120

  8. [8]

    Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014. Generative adversarial nets. Advances in neural information processing systems, 27

Show all 43 references
  1. [9]

    Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems, 33: 6840--6851

  2. [10]

    Hong, S.; Lee, J.; and Woo, S. S. 2024. All but One: Surgical Concept Erasing with Model Preservation in Text-to-Image Diffusion Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, 21143--21151

  3. [11]

    Karras, T.; Laine, S.; and Aila, T. 2019. A style-based generator architecture for generative adversarial networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4401--4410

  4. [12]

    Khosla, A.; Jayadevaprakash, N.; Yao, B.; and Li, F.-F. 2011. Novel dataset for fine-grained image categorization: Stanford dogs. In CVPR-W, volume 2. Citeseer

  5. [13]

    P.; and Welling, M

    Kingma, D. P.; and Welling, M. 2013. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114

  6. [14]

    Le, H.; and Samaras, D. 2019. Shadow removal via shadow image decomposition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 8578--8587

  7. [15]

    S.; Hou, Q.; Wang, Y.; and Yang, J

    Li, S.; van de Weijer, J.; Hu, T.; Khan, F. S.; Hou, Q.; Wang, Y.; and Yang, J. 2024. Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models. arXiv preprint arXiv:2402.05375

  8. [16]

    Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; and Lee, Y. J. 2023. GLIGEN: Open-set grounded text-to-image generation. In CVPR, 22511--22521

  9. [17]

    Liu, Z.; Yin, H.; Wu, X.; Wu, Z.; Mi, Y.; and Wang, S. 2021. From shadow generation to shadow removal. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 4927--4936

  10. [18]

    Luo, X.; Li, Y.; Chang, H.; Liu, C.; Milanfar, P.; and Yang, F. 2023. DVMark: a deep multiscale framework for video watermarking. IEEE Transactions on Image Processing

  11. [19]

    completely blind

    Mittal, A.; Soundararajan, R.; and Bovik, A. C. 2012. Making a “completely blind” image quality analyzer. IEEE Signal processing letters, 20(3): 209--212

  12. [20]

    O.; Cohen, N.; Mittal, G.; and Hegde, C

    Pham, M.; Marshall, K. O.; Cohen, N.; Mittal, G.; and Hegde, C. 2023. Circumventing concept erasure methods for text-to-image generative models. In The Twelfth International Conference on Learning Representations

  13. [21]

    O.; Hegde, C.; and Cohen, N

    Pham, M.; Marshall, K. O.; Hegde, C.; and Cohen, N. 2024. Robust Concept Erasure Using Task Vectors. arXiv preprint arXiv:2404.03631

  14. [22]

    Ramesh, A.; Dhariwal, P.; Nichol, A.; Chu, C.; and Chen, M. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2): 3

  15. [23]

    Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; and Sutskever, I. 2021. Zero-shot text-to-image generation. In International conference on machine learning, 8821--8831. Pmlr

  16. [24]

    Ray, A.; and Roy, S. 2020. Recent trends in image watermarking techniques for copyright protection: a survey. International Journal of Multimedia Information Retrieval, 9(4): 249--270

  17. [25]

    C.; and Fei-Fei, L

    Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015. ImageNet Large Scale Visual Recognition Challenge . International Journal of Computer Vision (IJCV), 115(3): 211--252

  18. [26]

    L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al

    Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E. L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing syste...

  19. [27]

    Singh, N.; Jain, M.; and Sharma, S. 2013. A survey of digital watermarking techniques. International Journal of Modern Communication Technologies and Research, 1(6): 265852

  20. [28]

    P.; Kumar, A.; Ermon, S.; and Poole, B

    Song, Y.; Sohl-Dickstein, J.; Kingma, D. P.; Kumar, A.; Ermon, S.; and Poole, B. 2020. Score-based generative modeling through stochastic differential equations. arXiv preprint arXiv:2011.13456

  21. [29]

    Sun, J.; Wang, X.; Shi, Y.; Wang, L.; Wang, J.; and Liu, Y. 2022. Ide-3d: Interactive disentangled editing for high-resolution 3d-aware portrait synthesis. ACM Transactions on Graphics (ToG), 41(6): 1--10

  22. [30]

    Tsai, Y.-L.; Hsu, C.-Y.; Xie, C.; Lin, C.-H.; Chen, J.-Y.; Li, B.; Chen, P.-Y.; Yu, C.-M.; and Huang, C.-Y. 2023. Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models? arXiv preprint arXiv:2310.10012

  23. [31]

    Wang, T.; Zhang, Y.; Qi, S.; Zhao, R.; Xia, Z.; and Weng, J. 2023. Security and privacy on generative data in aigc: A survey. arXiv preprint arXiv:2309.09435

  24. [32]

    Wang, Y.; Huang, W.; Sun, F.; Xu, T.; Rong, Y.; and Huang, J. 2020. Deep multimodal fusion by channel exchanging. Advances in neural information processing systems, 33: 4835--4845

  25. [33]

    C.; Sheikh, H

    Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4): 600--612

  26. [34]

    Wu, J.; Gan, W.; Chen, Z.; Wan, S.; and Lin, H. 2023. Ai-generated content (aigc): A survey. arXiv preprint arXiv:2304.06632

  27. [35]

    Xiao, B.; Wu, H.; Xu, W.; Dai, X.; Hu, H.; Lu, Y.; Zeng, M.; Liu, C.; and Yuan, L. 2024. Florence-2: Advancing a unified representation for a variety of vision tasks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4818--4829

  28. [36]

    Xiong, T.; Wu, Y.; Xie, E.; Li, Z.; and Liu, X. 2024. Editing Massive Concepts in Text-to-Image Diffusion Models. arXiv preprint arXiv:2403.13807

  29. [37]

    Xu, T.; Zhang, P.; Huang, Q.; Zhang, H.; Gan, Z.; Huang, X.; and He, X. 2018. Attngan: Fine-grained text to image generation with attentional generative adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1316--1324

  30. [38]

    Yang, Y.; Mu, K.; and Deng, R. H. 2022. Lightweight privacy-preserving GAN framework for model training and image synthesis. IEEE Transactions on Information Forensics and Security, 17: 1083--1098

  31. [39]

    Z.; Yang, X.; Lin, L.; Lee, M

    Yao, B. Z.; Yang, X.; Lin, L.; Lee, M. W.; and Zhu, S.-C. 2010. I2t: Image parsing to text description. Proceedings of the IEEE, 98(8): 1485--1508

  32. [40]

    Zhang, Z.; and Schomaker, L. 2021. Dtgan: Dual attention generative adversarial networks for text-to-image generation. In 2021 international joint conference on neural networks (IJCNN), 1--8. IEEE

  33. [41]

    Zhao, M.; Zhang, L.; Zheng, T.; Kong, Y.; and Yin, B. 2024. Separable Multi-Concept Erasure from Diffusion Models. arXiv preprint arXiv:2402.05947

  34. [42]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  35. [43]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools