Pith. sign in

REVIEW 9 cited by

Generated Faces in the Wild: Quantitative Comparison of Stable Diffusion, Midjourney and DALL-E 2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.00586 v2 pith:I5IVBP5K submitted 2022-10-02 cs.CV

classification cs.CV
keywords facesdiffusionmodelsstablewildcodecomparisondall-e
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The field of image synthesis has made great strides in the last couple of years. Recent models are capable of generating images with astonishing quality. Fine-grained evaluation of these models on some interesting categories such as faces is still missing. Here, we conduct a quantitative comparison of three popular systems including Stable Diffusion, Midjourney, and DALL-E 2 in their ability to generate photorealistic faces in the wild. We find that Stable Diffusion generates better faces than the other systems, according to the FID score. We also introduce a dataset of generated faces in the wild dubbed GFW, including a total of 15,076 faces. Furthermore, we hope that our study spurs follow-up research in assessing the generative models and improving them. Data and code are available at data and code, respectively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations

    cs.CV 2025-02 conditional novelty 7.0 of 10

    REALEDIT provides a large-scale, real-world image editing dataset from Reddit and demonstrates that training on it improves performance on authentic user requests.

  2. Sketchar: Supporting Character Design and Illustration Prototyping Using Generative AI

    cs.HC 2025-08 conditional novelty 6.0 of 10

    Sketchar, a ChatGPT and DALL-E based prototyping tool, helped game designers without artistic backgrounds generate reference images and character documents that they rated as more supportive of creativity than sketchi...

  3. AnyAni: An Interactive System with Generative AI for Animation Effect Creation and Code Understanding in Web Development

    cs.HC 2025-06 conditional novelty 6.0 of 10

    AnyAni combines LLM generation, a version tree, and video-based checking to help front-end developers create and understand web animations; a nine-person study reports usability gains over a chatbot baseline.

  4. The Role of Urban Designers in the Era of AIGC: An Experimental Study Based on Public Participation

    cs.HC 2024-11 conditional novelty 5.0 of 10

    In a 160-participant experiment, structured designer prompts produced higher-rated AI-generated urban pocket garden images than freeform prompts, but group differences were only marginally significant.

  5. AMCR: A Framework for Assessing and Mitigating Copyright Risks in Generative Models

    cs.LG 2025-08 reject novelty 4.0 of 10

    AMCR combines prompt sanitization, attention-based partial infringement detection, and a similarity-minimizing fine-tuning loss to reduce copyright infringement in text-to-image generation.

  6. EdgeDoc: Hybrid CNN-Transformer Model for Accurate Forgery Detection and Localization in ID Documents

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    A hybrid CNN-transformer with auxiliary noiseprint features achieves competitive forgery detection and localization on ID documents.

  7. A Multi-Agent Framework for Automated Qinqiang Opera Script Generation Using Large Language Models

    cs.AI 2025-04 reject novelty 4.0 of 10

    A prompt-based multi-agent pipeline generates Qinqiang opera scripts, visuals, and audio, but the reported 0.3-point expert-rated improvement over a vague baseline is not statistically supported.

  8. Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review

    eess.IV 2025-05 reject novelty 3.0 of 10

    A survey of DDPM, LDM, and WDM diffusion models for medical imaging, organized around training and inference efficiency.

  9. Novel AI Camera Camouflage: Face Cloaking Without Full Disguise

    cs.CV 2024-12 reject novelty 2.0 of 10

    Subtle cosmetic lines near facial key points and an alpha-layer PNG trick are claimed to hide faces from commercial detectors, but the evidence is anecdotal and not reproducible.

Pith tools