{"id":"d00c8151-43c9-4b8e-a980-7a05ca69331e","arxiv_id":"2411.15516","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Ten popular Stable Diffusion models generated harmful images for most test prompts, showed almost no refusal behavior, and displayed a bias associating Black individuals with violence.","lead":"This study tested ten popular Stable Diffusion image generators with explicit, violent, and celebrity-related prompts, generating about 24,000 images. It found that the models almost never refused, produced harmful images in most cases, and showed biases such as associating Black individuals with violence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated Hive thresholds and the absence of a benign-prompt control make every Section 3 percentage, model ranking, and race-bias ratio unanchored; a detector artifact could masquerade as model harmfulness.","rationale":"The reader's weakest assumption is the right one. The paper's central assertion has two parts. The capability part — that these models can generate harmful images and exhibit racial bias — is qualitatively corroborated by the manual visual analysis in Section 3.1 and by prior literature, so it is not the main risk. The quantitative part — that more than half of all generated images are harmful, that model rankings differ as reported, and that specific race-bias ratios hold — is produced entirely by Hive AI's moderation API, the thresholds in Appendix C, and DeepFace race labels. Section 2 gives no indication that either tool was validated against human ground truth on synthetic images, and the thresholds are asserted without sensitivity analysis. The absence of benign control prompts is especially important: a moderation API with a high base rate on synthetic images would inflate all unsafe ratios and could even invert model rankings, while a well-calibrated detector would leave the conclusions intact. The manual analysis cannot fix this because it is exploratory and not used to recalibrate the automated labels. Therefore the load-bearing concern is not that the models are harmless; it is that the exact quantitative headline, and the strongest derived inference of 'no refusal behavior,' is unanchored. A calibration study with human labels and control prompts would settle the question. If it passes, the quantitative claims are supported and the conditional verdict could move toward acceptance; if it fails, the quantitative claims require substantial revision. Since this is the same concern the reader identified, no verdict change is needed.","tokens_in":6906,"tokens_out":6912,"duration_ms":68651,"concrete_test":"Conduct a calibration study on a stratified sample of 300 generated images spanning all ten models and prompt categories, including images across the Hive score range. Have at least two independent human annotators label each image for the Appendix C categories and, where relevant, race; report agreement (Cohen's kappa and confusion matrix) between human labels and the Hive/DeepFace pipeline. In parallel, generate 20 images per model from 10 clearly benign control prompts (e.g., 'a professional portrait of a person', 'people enjoying a park') using the same pipeline and compute the control unsafe rate. If human/Hive agreement is poor, or if the benign-control unsafe rate approaches the reported 35.5% safe baseline, then all Section 3.2 percentages and rankings must be recomputed with calibrated thresholds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All of the paper's quantitative claims in Section 3.2 rest on a single chain: generated images are scored by Hive AI's moderation API, converted to unsafe labels by the thresholds in Appendix C, and combined with DeepFace race classifications for the bias analyses. Neither tool is validated on AI-generated images. Synthetic images contain unusual textures and anatomical artifacts — the paper itself documents 'vulva-penis hybrids' and distorted bodies — and Hive's probability-like scores may behave very differently on such inputs than on the real-world content they were presumably tuned for. The thresholds in Appendix C are presented as fixed rules with no sensitivity analysis, and the design includes no negative control: no benign prompts were run through the same pipeline to estimate the detector's false-positive base rate on synthetic images. Without such a baseline, the reported 35.5% average safe rate, the model ranking (SD3 at 55.13% safe vs. Epic Realism at 18.38%), and the race-bias ratios (e.g., 24.5% Black in gang violence) cannot be separated from detector behavior. The manual visual analysis in Section 3.1 supports the qualitative capability claim, but it is exploratory and is never used to audit or recalibrate the automated labels, so it does not rescue the quantitative conclusions. Thus the central assertion that 'more than half of images' are harmful, and the derived inference of 'no refusal behavior,' is load-bearing on an unvalidated classifier.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a safety analysis of ten popular Stable Diffusion models (five from Hugging Face, five from Civit AI). Using 50 prompts across nine categories (NSFW, violence, personally sensitive content) and generating 24,000 images, the authors apply Hive AI's vision moderation API, a celebrity classifier, and DeepFace's race classifier to quantify the prevalence of harmful content and investigate racial biases. The central claims are that, on average, only 35.5% of generated images are safe, that models show no refusal behavior, and that generated violence is disproportionately associated with Black individuals while nudity is predominantly White. The authors conclude that open text-to-image models lack adequate safety measures and call for better data curation, prompt moderation, and post-generation content filtering.","tokens_in":7171,"tokens_out":3219,"duration_ms":32551,"significance":"If the quantitative results hold, the paper would provide a useful large-scale empirical snapshot of the safety properties of widely used open-weight text-to-image models. The study has clear strengths: it evaluates ten popular models, uses a substantial generation corpus (24,000 images), includes a manual visual analysis that documents qualitative phenomena, and its model-selection rationale (download counts) is reasonable. However, the central quantitative claims—the average safe rate, the model ranking, and the race-bias ratios—rest on a single unvalidated classification pipeline and lack a negative control. Because the paper's main contribution is precisely these numbers, the missing validation is load-bearing rather than cosmetic.","major_comments":[{"comment":"The quantitative claims in Section 3.2 (average 35.5% safe rate, model ranking in Figure 1, per-category unsafe ratios in Figure 2) rely entirely on Hive AI's moderation API with the fixed thresholds listed in Appendix C. No validation is reported of this API on synthetic, AI-generated images, which differ substantially from real-world content—the paper itself documents 'vulva-penis hybrids,' distorted bodies, and unnatural anatomy in Section 3.1. Without a sensitivity analysis over the thresholds or a human-annotated audit sample of the generated images, the reported percentages are unanchored and could reflect detector miscalibration rather than model behavior.","section":"§3.2, Appendix C"},{"comment":"The experimental design includes no benign-prompt negative control. Without generating images from innocuous prompts and passing them through the same Hive moderation pipeline, the false-positive base rate of the detector on synthetic images is unknown. If the detector spuriously flags synthetic textures as 'bloody' or 'suggestive,' then the average safe rate, the model ordering, and the violence-related bias ratios (e.g., 24.5% Black in gang violence in Section 3.2.3) could be artifacts of the detector rather than properties of the models. The authors should either add such a control or explicitly justify why one is unnecessary.","section":"§3.2, §2 (Methods)"},{"comment":"The claim of a 'complete lack of any refusal behavior or safety measures' is not directly measured by the study. No metric for refusal is defined; no empty outputs, error messages, or text-based refusals are reported; and no benign-prompts baseline is used to contextualize the unsafe-image rate. The inference that models 'comply with the prompts instead of rejecting them' is made solely from the high proportion of images labeled unsafe by Hive. The authors should either define and measure refusal behavior explicitly (e.g., proportion of generation failures, empty outputs, or safety-filter activations) or temper the abstract and discussion to say that the models generated harmful content in a large fraction of trials.","section":"Abstract; §3.2; §4"},{"comment":"The race-bias analysis uses DeepFace to label race in synthetic images, but DeepFace is not validated on AI-generated faces. Given that the paper documents severe anatomical distortions and unrealistic body structures, face detection and race classification may fail at differential rates across images and across racial categories, which would directly bias the reported ratios such as 24.5% Black in gang violence versus 12.36% overall. The authors should validate DeepFace on a sample of their generated images (e.g., by manual inspection or comparison with a second classifier) and should also report the rate at which DeepFace fails to detect a face, since undetected faces are implicitly excluded from the denominator.","section":"§3.2.3, §2 (Methods)"},{"comment":"Several quantitative results are reported without any measure of uncertainty. For instance, the violence-related prompt categories contain only 2 prompts each (Appendix B), so the per-category unsafe ratios in Figure 2 are based on at most 40 images per model (2 prompts × 20 seeds) and have wide confidence intervals. The overall safe rates in Figure 1 are also point estimates with no error bars. The authors should provide confidence intervals, bootstrapped estimates, or per-prompt standard deviations so that the reader can judge whether differences such as SD3 at 55.13% safe versus Epic Realism at 18.38% are statistically reliable or within sampling noise.","section":"Appendix B; Figures 1 and 2"}],"minor_comments":[{"comment":"The prompt list is not included in the paper or appendix; the text states that prompts 'can be made available to other researchers upon reasonable request.' For reproducibility and transparency, the full prompt list should be published as a supplement or in a public repository.","section":"Appendix B"},{"comment":"The phrase 'It is to see that nudity is most often white' is awkward and should be rewritten (e.g., 'Nudity is most often depicted with light skin tones'). Similar phrasing issues appear elsewhere in Section 3.1 and 3.2.","section":"§3.1"},{"comment":"The sentence 'Alternatively, this pattern could indicate that male individuals were overrepresented by transgender individuals in the training data' is a speculative explanation unsupported by the data. It should be removed or clearly labeled as an unverified hypothesis, since the observed anatomical artifacts can be explained by training-data distributional properties without invoking transgender overrepresentation.","section":"§3.2.2"},{"comment":"The model names are not fully consistent: 'EpiCRealism Natural Sin RC1 VAE' is listed in Appendix A while the text and Figure 1 refer to 'Epic Realism.' Please standardize names and use the exact version identifiers.","section":"Appendix A"},{"comment":"Some references contain formatting errors, such as 'Struppek, Lukas; Hintersdorf, Dom; Friedrich, Felix; Br, Manuel' (incomplete author name) and 'In jair 78, pp. 1017–1068.' These should be corrected before publication.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the qualitative finding that open Stable Diffusion models readily generate harmful content is likely robust. However, the quantitative claims are the paper's main contribution, and they are currently supported only by an unvalidated commercial moderation API with no negative control. I believe these issues are fixable within the manuscript's scope by adding validation experiments, sensitivity analyses, and uncertainty estimates. I would also encourage the editor to consider whether the paper's current lack of a data/code release is acceptable for the journal's reproducibility standards."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before citing it. It is a concrete cross-model safety audit of ten popular Stable Diffusion models, including five CivitAI fine-tunes, which is new relative to prior work like Unsafe Diffusion and Stable Bias. The authors generated 24,000 images, used Hive AI's moderation API and DeepFace for race classification, and report large unsafe-image ratios and racial biases. The qualitative finding—these models readily produce explicit, violent, and biased content—is credible and consistent with prior literature. The manual visual analysis, while exploratory, supports the capability claim and documents odd artifacts like vulva-penis hybrids.\n\nThe soft spots are real and load-bearing for the quantitative claims. The stress-test note is on target: every percentage and ranking in Section 3.2 hinges on Hive's thresholds, which are presented as fixed rules with no sensitivity analysis and no validation on AI-generated images. Synthetic images contain unusual textures and anatomical distortions, so Hive's scores could behave differently on them than on real-world content. There is no negative control: no benign prompts were run through the same pipeline to estimate the detector's false-positive base rate. Without that baseline, the reported 35.5% safe rate, the model ordering, and the race-bias ratios (e.g., 24.5% Black in gang violence) cannot be cleanly separated from detector behavior. The claim of a 'complete lack of refusal behavior' is also not directly measured; the authors infer it from high unsafe rates, but they never tested whether models would refuse or hedge. That is an overstatement.\n\nThere are additional weaknesses: percentages lack error bars, the celebrity identification threshold is arbitrary, and no prompts, code, or data are released—only a vague 'available upon request.' These are all addressable.\n\nStill, the paper deserves a serious referee. It is a useful audit with a plausible core result, and the methodological flaws are fixable. I would send it to review but with a clear request for major revision: add a benign-prompt control, validate the classifiers on a sample of generated images, report threshold sensitivity, release the prompts and ideally the images, and temper the refusal-behavior conclusion. If those are done, the quantitative claims would have a solid foundation.\n\nFor a reading group, it's a fine case study in how evaluation pipelines can confound results. I'd bring it up for that reason, but I wouldn't cite the numbers as-is in my own work until the validation is added.","headline":"A concrete but methodologically fragile audit: the core qualitative finding holds, yet the paper's percentages rest on an unvalidated classifier with no benign-prompt baseline.","tokens_in":7681,"tokens_out":1708,"would_cite":false,"duration_ms":17562,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Study: Most Stable Diffusion images from harmful prompts are unsafe","keywords":["text-to-image safety","Stable Diffusion","content moderation","NSFW content","violence bias","racial bias","synthetic media","refusal behavior"],"falsifier":"Take a random sample of the 24,000 generated images, have human annotators independently label them with the same nine categories, and compare their labels to the automated ones; if agreement is low or the racial composition of violent images changes under human labeling, the reported percentages and bias ratios would not survive.","tokens_in":6699,"feed_emoji":"🖼️","tokens_out":9331,"duration_ms":81470,"temperature":0.7,"pith_summary":"Ten widely downloaded open text-to-image models were prompted with 50 harmful prompts spanning nudity, sexual acts, violence, hate crimes, and personal depictions of public figures, and every model responded by generating the requested content instead of refusing. Across 24,000 generated images, only 35.5% avoided being flagged by an automated content-moderation pipeline, making the majority of outputs unsafe by the study's definitions. The paper also reports that nudity and sexual imagery were skewed heavily toward white individuals, while images of gang violence disproportionately showed Black individuals even though no prompt mentioned race. Celebrity depictions were usually too unrealistic to be useful for abuse, though smoking and gambling scenes were the most common recognizable personal-sensitive outputs. These findings matter because these models are widely distributed, run offline, and appear to ship without the refusal behavior found in closed assistants.","feed_headline":"Most Stable Diffusion images from harmful prompts are unsafe","feed_subtitle":"A 24,000-image test of ten popular open models found no refusals and violence skewed toward Black individuals.","key_machinery":"The argument is carried by a standardized image-generation and classification pipeline: each model is run with its own recommended sampling configuration, the same 50 prompts are applied 20 times, and every output is passed to an automated visual moderation system with category-specific thresholds (for example, general_nsfw at 0.7, very_bloody at 0.5, and a combination of identification score and content scores for celebrity images). An image counts as safe only if it exceeds none of the thresholds for its prompt category. A separate face-recognition step labels the race of depicted individuals, and a celebrity classifier determines whether a public figure is recognizable. This pipeline converts 'does the model refuse?' into a measurable safety percentage, and it is what the reported model rankings and bias ratios rest on.","core_discovery":"The paper's central claim is that popular Stable Diffusion models, including base models and fine-tuned community versions, have no effective safety layer: when given prompts for NSFW, violent, or personally sensitive content, they comply. Measured across ten models with 50 prompts applied 20 times each, the safest model still produced harmful images 44.87% of the time and the least safe produced them 81.62% of the time, with an average of 64.5% of all images flagged as unsafe. The same classifier-based analysis found strong representation bias: white individuals dominated nudity and sexual content, while Black individuals were overrepresented in gang-violence images. The authors interpret these patterns as inherited from training data and conclude that current open image generators cannot be considered safe or neutral by default.","pith_inferences":["The paper leaves the safety percentages to a single commercial moderation API; a reasonable extension would be to re-label a random sample by human raters and measure agreement, since API thresholds were not validated on synthetic images.","Because the race classifier was also not validated on generated faces, the reported bias ratios are provisional; human labeling could shift them.","An implication the authors do not develop is that offline distribution makes post-hoc safety unenforceable; the practical intervention point may be the training data and the initial model release.","A testable extension would be to prompt the same models with neutral scenes and measure whether violent context alone shifts the racial composition, separating stereotype bias from prompt compliance."],"forward_implications":["If the results hold, a user can generate explicit or violent images from a majority of harmful prompts on today's most-downloaded open models, with no built-in refusal.","Model rankings by safety shift with the distribution platform: the newest base model was least harmful, while realism-focused fine-tunes were most harmful.","Racial bias in generated violence is not tied to any race word in the prompt, so the models themselves associate Black individuals with violent contexts.","Current celebrity-image risk is low because generated faces are not believable, but the same pipeline would need re-testing as realism improves.","The observed behavior supports adding prompt filters, post-generation content classifiers, and better training-data curation to open image models."],"supporting_citations":[{"why":"Defines the Stable Diffusion architecture that all ten evaluated models are built from, so it supplies the object under test.","marker":"(Rombach et al. 2022)"},{"why":"Earlier demonstration that text-to-image models produce unsafe images and hateful memes, the line of work this study extends.","marker":"(Qu et al. 2023)"},{"why":"Documents misogyny, pornography, and malignant stereotypes in large multimodal training datasets, supporting the data-bias explanation.","marker":"(Birhane et al. 2021)"},{"why":"Documents abusive material in the dataset used to train Stable Diffusion, motivating training-data curation.","marker":"(Thiel 2023)"},{"why":"Earlier analysis of societal representations in diffusion models, used as prior evidence for the bias findings.","marker":"(Luccioni et al. 2023)"},{"why":"Shows how instruction following and safety tuning work in closed models, the contrast that makes the absence of refusal in open models salient.","marker":"(Ouyang et al. 2022)"},{"why":"Shows attackers can bypass text-to-image safety filters, contextualizing why filter-based defenses are not sufficient.","marker":"(Deng and Chen 2024)"}],"fun_headline_variants":["Stable Diffusion models never refuse harmful prompts","No safety layer: Stable Diffusion complies with every harmful prompt","Bias check: Stable Diffusion links Black identity with violence","Test of 10 Stable Diffusion models finds zero refusal behavior","Image generators lack safety: harmful prompts succeed 64% of the time"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated moderation API and the race classifier correctly label AI-generated images as harmful or safe; the paper reports no validation of either tool on synthetic images.","fun_headline_variants_meta":{"raw":{"variants":["Stable Diffusion models never refuse harmful prompts","No safety layer: Stable Diffusion complies with every harmful prompt","Bias check: Stable Diffusion links Black identity with violence","Test of 10 Stable Diffusion models finds zero refusal behavior","Image generators lack safety: harmful prompts succeed 64% of the time"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1499,"prompt_tokens":792,"completion_tokens":707,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":626}},"tokens_in":408,"tokens_out":707,"duration_ms":6141,"temperature":1.0,"reasoning_tokens":626,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:11:51.707927+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the 24,000 generated images, have human annotators independently label them with the same nine categories, and compare their labels to the automated ones; if agreement is low or the racial composition of violent images changes under human labeling, the reported percentages and bias ratios would not survive.","supporting_citations":[{"cited_title":"safe” when they are not detected by any of our classifier’s category, and “unsafe","cited_arxiv_id":null,"evidence_quote":"Documents misogyny, pornography, and malignant stereotypes in large multimodal training datasets, supporting the data-bias explanation."}],"review_version":1}