REVIEW 3 major objections 5 minor 2 cited by
Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read On the open-source image-model hub CivitAI, explicit content rose from 41% to 80% of posted images in two years, while 15.27% of all model adapters are built to mimic real people.
desk verdict The most complete CivitAI audit to date, with a credible but label-dependent NSFW trend; the numbers need a stability check before they are cited as exact. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery that carries the argument has three parts. First, the model adapter—chiefly LoRA (low-rank adaptation, a small trainable module that steers a large image model) and textual inversion—is what lets a user replicate a specific face, body, or style without retraining a foundation model, turning a shared model into a multiplier for unlimited derivative images. Second, the "person of interest" flag in CivitAI metadata identifies which adapters are tailored to real people, giving the authors a population-level count of deepfake-ready assets. Third, the auto-tagging pipeline: captioning systems such as CLIP Interrogator and Danbooru-based taggers supply the text used to train adapters, and the paper shows that more than 70% of adapters with extractable captions used Danbooru-derived tags, whose sexualized vocabulary normalizes explicit content, with 5.6% containing "loli"/"shota" and 2.1% containing "rape". These three components together convert individual user choices into systematic, platform-wide bias.
What would settle it
Take a random sample of images from each month between January 2023 and December 2024, have annotators apply a single fixed explicitness rubric blind to date, and recompute the monthly share of explicit images. If the fixed-rubric share stays roughly flat while the platform-labeled share climbs from 41% to 80%, the headline trend would be shown to be an artifact of classifier or moderation drift.
Extended reading notes
Core claim
On the paper's own terms, the central finding is that CivitAI has become an explicit-content pipeline: the share of posted images labeled NSFW rose from 41% in January 2023 to 80% by December 2024, and image subjects inferred as female outnumber male subjects by 6.24 to 1, are younger on average (23.67 versus 31.16 years), and appear in NSFW contexts more often (85.44% versus 73.66%). It further reports that 15.27% of model adapters (33,804) carry a "person of interest" flag indicating they are tailored to replicate a real individual, with actors, adult performers, and online personalities among the most common targets. The paper attributes these patterns to a layered mechanism: popularity-based feedback loops favor sensational content, and auto-tagging systems built on anime-imageboard vocabularies—used by more than 70% of adapters with identifiable training captions—embed sexualized, occasionally abusive label sets into model training. The authors conclude that personalization hubs are not neutral distribution points but active amplifiers that normalize gendered harm.
Load-bearing premise
The paper's central trend depends on the assumption that the platform's explicit-content labels were equally strict in 2023 and 2024; if the labeler's thresholds shifted, part of the 41-to-80 surge could be classification drift rather than a change in what users actually post.
Editorial extensions
If this is right
- If the current trajectory holds, the NSFW share on model-sharing hubs will continue to climb because engagement incentives reward sensational content, making the rise a structural feature rather than a removable nuisance.
- Deepfake adapters remain available after download and can be redistributed outside the platform, so policy changes that only affect on-platform display or tags will not stop non-consensual imagery from spreading.
- Because most adapters are trained with Danbooru-based auto-tagging, changing the default captioning tool used in training interfaces is a concrete intervention point for reducing explicit and abusive training vocabularies.
- Existing image-perturbation defenses lose effectiveness against newer architectures and against LoRA fine-tuning, so protective tools must be continuously updated to keep pace with model development.
- The biases the paper documents are embedded in models rather than only in individual images, so any downstream consumer product that integrates these open-source models or adapters would inherit the skewed representations.
Reading between the lines
- An implicit testable implication: the 41-to-80 trend should be re-measured on a fixed, human-labeled sample to separate genuine content change from drift in the platform's own classification thresholds.
- The training-tag finding suggests a natural experiment: comparing two otherwise identical LoRA training runs, one captioned with a Danbooru-based tagger and one with a neutral captioner, could isolate how much of the gendered NSFW skew comes from captioning vocabulary rather than from the underlying images.
- If the celebrity-creator leaderboard and tag-based discoverability are as influential as the paper suggests, removing popularity-sorted feeds or sexualized default tags would provide a direct test of whether platform design, not user demand alone, drives the rise.
- The demographic measurements rely on a binary gender classifier; with more inclusive annotation the 6.24:1 ratio might shift, but the qualitative pattern of younger female subjects in explicit contexts is unlikely to vanish.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an exploratory sociotechnical analysis of CivitAI, a major open-source text-to-image model-sharing platform, using metadata for 40,630,560 images and 231,252 models from late November 2022 to December 2024. It reports a rise in the platform-assigned NSFW ratio from 41% in January 2023 to 80% in December 2024, a female-to-male ratio of 6.24:1 among depicted subjects with women depicted younger and more often in NSFW contexts, a tag and training-metadata analysis showing widespread use of Danbooru-based captioning, and an estimate that 15.27% (33,804) of model adapters are flagged as person-of-interest, i.e., tailored to replicate real individuals. The authors interpret these findings through feminist and constructivist frameworks and propose interventions in content moderation, tool design, and platform policy.
Significance. If the headline trends hold, this is a substantial empirical contribution to the literature on AI governance and gender-based harm: it is one of the largest descriptive studies of a text-to-image model-sharing platform, it covers a two-year period at a scale no prior study has matched, and it connects image-level, model-level, and training-pipeline evidence. The paper is also transparent about several limitations, including acknowledged bias in NSFW classifiers, MiVOLO's binary gender classification, and potential misclassification in the LLM-based profession inference. The main caveat is that the central temporal claim rests on platform-assigned NSFW labels whose stability over time is not demonstrated, and the demographic analysis rests on an underspecified 0.1% sample. These issues are fixable and do not undermine the value of the dataset or the cross-sectional measurements, but they do affect how strongly the paper's central narrative can be asserted.
major comments (3)
- [Section 2.2.1, Figure 3, Table S1] The headline NSFW increase from 41% in January 2023 to 80% in December 2024 is computed entirely from CivitAI's browsingLevel field, which the paper states (Section 2.1) is assigned by Amazon Rekognition, unspecified open-source NSFW classifiers, and manual review. The paper does not test whether classifier versions, classification thresholds, or manual-review policies drifted over the two-year period. This is load-bearing because the platform-decay argument in Section 3.1 and the phrase 'disproportionate rise' rest directly on this trend. The paper cites Leu et al. (2024) for cross-sectional classifier bias but does not address temporal drift. I ask the authors to re-annotate a monthly stratified sample of images with one fixed classifier and present both the platform-assigned series and the re-annotated series; at minimum, they should document any known CivitAI moderation-policy changes during 2023-2024 and perform a sensitivity analysis using alternative NSFW thresholds. Without such a check, the temporal trend may partly reflect a change in measurement rather than a change in uploaded content.
- [Section 2.2.2, Table S2] The demographic analysis is based on a 0.1% subsample described only as 'temporally balanced and distributionally representative' (Section 2.2.2). No sampling protocol is given: the text does not specify the strata (e.g., month, NSFW level, or image type), the random seed, the exclusion criteria, or the confidence threshold used to discard images with no detected human subjects. Table S2 suggests that monthly sample sizes scale roughly with total image volume, but the reader cannot determine whether images were drawn uniformly within months or whether any NSFW/SFW balancing was applied. Because the female-to-male ratio, the mean ages, and the NSFW co-occurrence rates are all derived from this sample, the authors should specify the exact sampling procedure, release the sample image identifiers, and report the number of images excluded at each step.
- [Section 2.3.2, Figure 6] The checkpoint inference experiment compares outputs from the ten most downloaded checkpoints using the prompts 'woman' and 'man', but the caption states that 'initial noise seeds were iterated to ensure a representative, uncurated sample' without specifying the iteration rule, the number of seeds tried, or the criteria for stopping. This makes the comparison non-reproducible and raises the possibility of selective presentation. Since this experiment is illustrative rather than central to the paper's main claims, the issue is not blocking, but the text should either state the exact seed-iteration protocol or present results over a fixed set of seeds.
minor comments (5)
- [Section 2.1] The sentence 'The dataset covers 40,630,560 images and 231,252 models as of 19 th January' is missing a year; it should read 'as of January 19, 2025' to match the later text.
- [Figure 6 caption] The caption says 'as of 19 th January 2024', while the main text (Section 2.3.2) says the checkpoints were selected 'as of January 19 th, 2025'. The caption year appears to be a typo and should be corrected.
- [Section 2.1] In the sentence describing safetensors files, 'as described by by321 (2022)' contains a duplicated 'by'; this should be 'as described by321 (2022)' or 'as described by by321 (2022)' should be rephrased.
- [Figure 5] The tag co-occurrence network contains a full-width comma in 'girl,'; this appears to be a typo and should be replaced with a regular comma.
- [Section 2.3.4] The profession and country inference using DeepSeek is acknowledged as potentially error-prone, but the paper does not report any validation of the LLM annotations against a human-annotated sample. A short inter-annotator agreement or manual-validation subsection would strengthen confidence in the sunburst-chart results.
Circularity Check
No significant circularity: the headline NSFW, demographic, and person-of-interest quantities are direct measurements from collected platform data and external classifiers, not outputs derived from fitted parameters or self-citations.
full rationale
The paper's central claims are empirical measurements rather than derivations. The NSFW ratio trend (41% to 80%) is computed directly from CivitAI's browsing-level labels in the authors' own extended dataset, the female-to-male ratio and age estimates come from the external MiVOLO classifier applied to a sampled subset, and the 15.27% person-of-interest figure is a count of platform-provided POI flags. No equation in the paper maps a fitted parameter into a predicted quantity, and no target result is defined in terms of the inputs that supposedly produce it. The citation to the authors' prior Civiverse work (Palmini et al., 2024) is used only as contextual precedent; the current paper recomputes the trend on its own 40,630,560-image dataset, so the earlier study is not load-bearing evidence for the headline finding. The acknowledged limitations concerning CivitAI's browsing-level classifiers, MiVOLO bias, and age-estimation bias are measurement-validity concerns rather than circularity: the paper claims to describe the platform's labeled content distribution, and those labels are an input measurement, not a construct secretly defined by the conclusion. There is no self-citation chain invoked to force an interpretation, no uniqueness theorem imported from the authors, and no ansatz smuggled in via citation. The derivation chain is therefore self-contained with respect to the paper's stated claims, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- Caption-source matching threshold =
5 matches
- MiVOLO sample size =
40,636 images (0.1%)
- Checkpoint inference seed iteration =
unstated
assumptions (4)
- domain assumption CivitAI NSFW browsing levels are a consistent measure of content explicitness across 2022-2024
- domain assumption MiVOLO's binary gender and age predictions are sufficiently accurate to support ratio claims
- domain assumption LLM-based profession and country inference for deepfake targets is reliable
- domain assumption Danbooru-derived tag taxonomy accurately categorizes training captions
Cite this review
Pith. "Pith review of Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm." pith.science (2026). https://pith.science/paper/5ZG7J3Z2
@misc{pith2026250504600,
author = {Pith},
title = {Pith review of: Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm},
year = {2026},
howpublished = {\url{https://pith.science/paper/5ZG7J3Z2}},
note = {Machine review of arXiv:2505.04600}
}
read the original abstract
Open-source text-to-image (TTI) pipelines have become dominant in the landscape of AI-generated visual content, driven by technological advances that enable users to personalize models through adapters tailored to specific tasks. While personalization methods such as LoRA offer unprecedented creative opportunities, they also facilitate harmful practices, including the generation of non-consensual deepfakes and the amplification of misogynistic or hypersexualized content. This study presents an exploratory sociotechnical analysis of CivitAI, the most active platform for sharing and developing open-source TTI models. Drawing on a dataset of more than 40 million user-generated images and over 230,000 models, we find a disproportionate rise in not-safe-for-work (NSFW) content and a significant number of models intended to mimic real individuals. We also observe a strong influence of internet subcultures on the tools and practices shaping model personalizations and resulting visual media. In response to these findings, we contextualize the emergence of exploitative visual media through feminist and constructivist perspectives on technology, emphasizing how design choices and community dynamics shape platform outcomes. Building on this analysis, we propose interventions aimed at mitigating downstream harm, including improved content moderation, rethinking tool design, and establishing clearer platform policies to promote accountability and consent.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns
A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.
-
Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)
The dominant real-world use of generative-image abuse is non-consensual intimate imagery, yet the AI/ML research field focuses almost exclusively on viewer deception.
Reference graph
Works this paper leans on
-
[1]
Abbate J (2012) Introduction: Rediscovering Women’s History in Computing , chapter
work page 2012
-
[2]
The MIT Press, pp. 1–10. DOI:10.7551/mitpress/9014.001.0001. Ajder H, Patrini G, Cavalli F and Cullen L (2019) The state of deepfakes: Landscape, threats, and impact. URL https://regmedia.co.uk/2019/10/08/deepfake_report.pdf . Deeptrace. (Accessed 2 May 2025). Anderson M (2025) Civitai tightens deepfake rules under pressure from mastercard and visa. URL h...
arXiv 2019
-
[3]
London: Palgrave Macmillan UK, pp. 14–26. DOI:10.1007/978-1-349-19798-9
-
[6]
Cui J and Araujo DA (2024) Rethinking use-restricted open-source licenses for regulating abuse of generative models. Big Data & Society 11(1). DOI:10.1177/20539517241229699. DeepSeek-AI et al. (2025) Deepseek-v3 technical report. DOI:10.48550/arXiv.2412.19437. Doane MA (1991) Femmes Fatales: Feminism, Film Theory, Psychoanalysis . Film - Women’s Studies. ...
-
[9]
Oxford University Press, pp. 55–77. DOI: 10.1093/acprof:oso/9780190604981.003.0003. Montani I and Honnibal M (2021) spacy 3: Industrial-strength natural language processing in python. URL https://spacy.io. SpaCy. (Accessed 2 May 2025). Mulvey L (1989) Visual Pleasure and Narrative Cinema , chapter
arXiv 2021
-
[11]
In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (AIES)
Naik R and Nushi B (2023) Social biases through the text-to-image generation lens. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (AIES). ACM, pp. 786–808. DOI: 10.1145/3600211.3604711. Ng K (2025) Pakistan: Us teen shot dead by father over tiktok videos. URL https://www.bbc.co m/news/articles/cy8pvw3xxxeo. BBC News. (Accessed ...
arXiv 2023
-
[15]
could be categorized as child pornography,
Leu W, Nakashima Y and Garcia N (2024) Auditing image-based NSFW classifiers for content filtering. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency . ACM, pp. 1163–1173. DOI:10.1145/3630106.3658963. Li Q, Zhang S, Kasper AT, Ashkinaze J, Eaton AA, Schoenebeck S and Gilbert E (2024) Reporting non-consensual intimate media: An audi...
arXiv 2024
-
[16]
ACM, pp. 620:1–620:25. DOI:10.1145/3613904.3642877. 23 Zhang Y, Jiang L, Turk G and Yang D (2024b) Auditing gender presentation differences in text-to- image models. In: Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO) . ACM, pp. 12:1–12:10. DOI:10.1145/3689904.3694710. Zhao Z, Duan J, Xu K, Wa...
arXiv 2024
Show all 17 references
-
[17]
Y ear Month T otal Images NSFW T rue NSFW F alse Avg
The table reports the total number of generated images, the number classified as NSFW (Not Safe for Work) and non-NSFW, the average number of images generated per day, and the NSFW ratio (the proportion of NSFW images relative to the total image count). Y ear Month T otal Imag...
2022
-
[18]
Wildcards
The NSFW ratio is calculated as the proportion of NSFW-true images out of the total monthly count. 27 Table S2: Addendum to Section 3.2, MiVOLO Age and Gender Estimations for 2023–2024 Month Total Images Total Persons No Per- sons ♀ (%) ♂ (%) ♀ ♂ ♀ Age (Mean ± SD) ♂ Age (Mean ...
2023
-
[30]
Wang W, Bai H, Huang J, Wan Y, Yuan Y, Qiu H, Peng N and Lyu MR (2024) New job, new gender? measuring the social bias in image generation models
DOI:10.1007/s11229-022-04012-2. Wang W, Bai H, Huang J, Wan Y, Yuan Y, Qiu H, Peng N and Lyu MR (2024) New job, new gender? measuring the social bias in image generation models. In: Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Au...
2024
-
[37]
Kim K (2024) KichangKim/DeepDanbooru
DOI:10.48550/arXiv.2407.06863. Kim K (2024) KichangKim/DeepDanbooru. URL https://github.com/KichangKim/DeepDanbooru. GitHub. (Accessed 26 July 2024). King D (2024) The algorithmic gaze: Representations of women in AI art. URL https://www.le random.art/editorial/the-algorithmic...
-
[62]
CivitAI (2023) REST API reference
DOI:10.1093/geront/gnab167. CivitAI (2023) REST API reference. URLh t t p s : / / g i t h u b . c o m / c i v i t a i / c i v i t a i / w i k i / R E ST -A P I -R e f e r e n c e. GitHub. (Accessed 14 January 2025). CivitAI (2025) Content moderation — civitai. URL https://educ...
2023 doi
- [81]
-
[139]
8748–8763
PMLR, pp. 8748–8763. Rombach R, Blattmann A, Lorenz D, Esser P and Ommer B (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10684–10695. DOI:10.1109/CVPR52688....
2022
-
[2023]
DOI:10.48550/arXiv.2208.01618
Kigali, Rwanda. DOI:10.48550/arXiv.2208.01618. Gavrilovi´ c Nilsson M, Tzani-Pepelasis C, Ioannou M and Lester D (2019) Understanding the link between sextortion and suicide. International Journal of Cyber Criminology 13(1): 55–69. DOI: 10.5281/zenodo.3402357. Gervais SJ, Vesc...
-
[2024]
6949–6958
ACM, pp. 6949–6958. DOI:10.1145/3664647.3681052. Widder DG, West S and Whittaker M (2023) Open (for business): Big tech, concentrated power, and the political economy of open ai. SSRN Electronic Journal DOI:10.2139/ssrn.4543807. Wu Y, Nakashima Y and Garcia N (2024) Stable dif...
2023
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.