Pith. sign in

REVIEW 3 major objections 5 minor 2 cited by

Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read On the open-source image-model hub CivitAI, explicit content rose from 41% to 80% of posted images in two years, while 15.27% of all model adapters are built to mimic real people.

desk verdict The most complete CivitAI audit to date, with a credible but label-dependent NSFW trend; the numbers need a stability check before they are cited as exact. read the letter →

arxiv 2505.04600 v2 pith:5ZG7J3Z2 submitted 2025-05-07 cs.CY

classification cs.CY
keywords generativeAItext-to-imagemodelsCivitmodelpersonalizationNSFWcontentdeepfakesgenderbiasplatformdecay
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is an empirical study of CivitAI, the largest hub for sharing and downloading open-source text-to-image models and the small adapters that personalize them. Drawing on metadata for more than 40 million images and over 230,000 models, it argues that the platform's content has shifted sharply toward explicit imagery, that female-presenting subjects are over-represented, younger, and more often sexualized, and that a substantial share of adapters are designed to replicate identifiable real people. The authors read these patterns as a self-reinforcing sociotechnical cycle rather than isolated misuse: engagement incentives reward sensational content, subculture-derived auto-tagging systems inject sexualized vocabulary into model training, and policy gaps let deepfake adapters circulate. If the picture is right, open-source image generation is not merely reflecting existing biases but actively normalizing misogynistic and exploitative imagery at scale.

What carries the argument

The machinery that carries the argument has three parts. First, the model adapter—chiefly LoRA (low-rank adaptation, a small trainable module that steers a large image model) and textual inversion—is what lets a user replicate a specific face, body, or style without retraining a foundation model, turning a shared model into a multiplier for unlimited derivative images. Second, the "person of interest" flag in CivitAI metadata identifies which adapters are tailored to real people, giving the authors a population-level count of deepfake-ready assets. Third, the auto-tagging pipeline: captioning systems such as CLIP Interrogator and Danbooru-based taggers supply the text used to train adapters, and the paper shows that more than 70% of adapters with extractable captions used Danbooru-derived tags, whose sexualized vocabulary normalizes explicit content, with 5.6% containing "loli"/"shota" and 2.1% containing "rape". These three components together convert individual user choices into systematic, platform-wide bias.

What would settle it

Take a random sample of images from each month between January 2023 and December 2024, have annotators apply a single fixed explicitness rubric blind to date, and recompute the monthly share of explicit images. If the fixed-rubric share stays roughly flat while the platform-labeled share climbs from 41% to 80%, the headline trend would be shown to be an artifact of classifier or moderation drift.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central finding is that CivitAI has become an explicit-content pipeline: the share of posted images labeled NSFW rose from 41% in January 2023 to 80% by December 2024, and image subjects inferred as female outnumber male subjects by 6.24 to 1, are younger on average (23.67 versus 31.16 years), and appear in NSFW contexts more often (85.44% versus 73.66%). It further reports that 15.27% of model adapters (33,804) carry a "person of interest" flag indicating they are tailored to replicate a real individual, with actors, adult performers, and online personalities among the most common targets. The paper attributes these patterns to a layered mechanism: popularity-based feedback loops favor sensational content, and auto-tagging systems built on anime-imageboard vocabularies—used by more than 70% of adapters with identifiable training captions—embed sexualized, occasionally abusive label sets into model training. The authors conclude that personalization hubs are not neutral distribution points but active amplifiers that normalize gendered harm.

Load-bearing premise

The paper's central trend depends on the assumption that the platform's explicit-content labels were equally strict in 2023 and 2024; if the labeler's thresholds shifted, part of the 41-to-80 surge could be classification drift rather than a change in what users actually post.

Editorial extensions

If this is right

  • If the current trajectory holds, the NSFW share on model-sharing hubs will continue to climb because engagement incentives reward sensational content, making the rise a structural feature rather than a removable nuisance.
  • Deepfake adapters remain available after download and can be redistributed outside the platform, so policy changes that only affect on-platform display or tags will not stop non-consensual imagery from spreading.
  • Because most adapters are trained with Danbooru-based auto-tagging, changing the default captioning tool used in training interfaces is a concrete intervention point for reducing explicit and abusive training vocabularies.
  • Existing image-perturbation defenses lose effectiveness against newer architectures and against LoRA fine-tuning, so protective tools must be continuously updated to keep pace with model development.
  • The biases the paper documents are embedded in models rather than only in individual images, so any downstream consumer product that integrates these open-source models or adapters would inherit the skewed representations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit testable implication: the 41-to-80 trend should be re-measured on a fixed, human-labeled sample to separate genuine content change from drift in the platform's own classification thresholds.
  • The training-tag finding suggests a natural experiment: comparing two otherwise identical LoRA training runs, one captioned with a Danbooru-based tagger and one with a neutral captioner, could isolate how much of the gendered NSFW skew comes from captioning vocabulary rather than from the underlying images.
  • If the celebrity-creator leaderboard and tag-based discoverability are as influential as the paper suggests, removing popularity-sorted feeds or sexualized default tags would provide a direct test of whether platform design, not user demand alone, drives the rise.
  • The demographic measurements rely on a binary gender classifier; with more inclusive annotation the 6.24:1 ratio might shift, but the qualitative pattern of younger female subjects in explicit contexts is unlikely to vanish.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents an exploratory sociotechnical analysis of CivitAI, a major open-source text-to-image model-sharing platform, using metadata for 40,630,560 images and 231,252 models from late November 2022 to December 2024. It reports a rise in the platform-assigned NSFW ratio from 41% in January 2023 to 80% in December 2024, a female-to-male ratio of 6.24:1 among depicted subjects with women depicted younger and more often in NSFW contexts, a tag and training-metadata analysis showing widespread use of Danbooru-based captioning, and an estimate that 15.27% (33,804) of model adapters are flagged as person-of-interest, i.e., tailored to replicate real individuals. The authors interpret these findings through feminist and constructivist frameworks and propose interventions in content moderation, tool design, and platform policy.

Significance. If the headline trends hold, this is a substantial empirical contribution to the literature on AI governance and gender-based harm: it is one of the largest descriptive studies of a text-to-image model-sharing platform, it covers a two-year period at a scale no prior study has matched, and it connects image-level, model-level, and training-pipeline evidence. The paper is also transparent about several limitations, including acknowledged bias in NSFW classifiers, MiVOLO's binary gender classification, and potential misclassification in the LLM-based profession inference. The main caveat is that the central temporal claim rests on platform-assigned NSFW labels whose stability over time is not demonstrated, and the demographic analysis rests on an underspecified 0.1% sample. These issues are fixable and do not undermine the value of the dataset or the cross-sectional measurements, but they do affect how strongly the paper's central narrative can be asserted.

major comments (3)
  1. [Section 2.2.1, Figure 3, Table S1] The headline NSFW increase from 41% in January 2023 to 80% in December 2024 is computed entirely from CivitAI's browsingLevel field, which the paper states (Section 2.1) is assigned by Amazon Rekognition, unspecified open-source NSFW classifiers, and manual review. The paper does not test whether classifier versions, classification thresholds, or manual-review policies drifted over the two-year period. This is load-bearing because the platform-decay argument in Section 3.1 and the phrase 'disproportionate rise' rest directly on this trend. The paper cites Leu et al. (2024) for cross-sectional classifier bias but does not address temporal drift. I ask the authors to re-annotate a monthly stratified sample of images with one fixed classifier and present both the platform-assigned series and the re-annotated series; at minimum, they should document any known CivitAI moderation-policy changes during 2023-2024 and perform a sensitivity analysis using alternative NSFW thresholds. Without such a check, the temporal trend may partly reflect a change in measurement rather than a change in uploaded content.
  2. [Section 2.2.2, Table S2] The demographic analysis is based on a 0.1% subsample described only as 'temporally balanced and distributionally representative' (Section 2.2.2). No sampling protocol is given: the text does not specify the strata (e.g., month, NSFW level, or image type), the random seed, the exclusion criteria, or the confidence threshold used to discard images with no detected human subjects. Table S2 suggests that monthly sample sizes scale roughly with total image volume, but the reader cannot determine whether images were drawn uniformly within months or whether any NSFW/SFW balancing was applied. Because the female-to-male ratio, the mean ages, and the NSFW co-occurrence rates are all derived from this sample, the authors should specify the exact sampling procedure, release the sample image identifiers, and report the number of images excluded at each step.
  3. [Section 2.3.2, Figure 6] The checkpoint inference experiment compares outputs from the ten most downloaded checkpoints using the prompts 'woman' and 'man', but the caption states that 'initial noise seeds were iterated to ensure a representative, uncurated sample' without specifying the iteration rule, the number of seeds tried, or the criteria for stopping. This makes the comparison non-reproducible and raises the possibility of selective presentation. Since this experiment is illustrative rather than central to the paper's main claims, the issue is not blocking, but the text should either state the exact seed-iteration protocol or present results over a fixed set of seeds.
minor comments (5)
  1. [Section 2.1] The sentence 'The dataset covers 40,630,560 images and 231,252 models as of 19 th January' is missing a year; it should read 'as of January 19, 2025' to match the later text.
  2. [Figure 6 caption] The caption says 'as of 19 th January 2024', while the main text (Section 2.3.2) says the checkpoints were selected 'as of January 19 th, 2025'. The caption year appears to be a typo and should be corrected.
  3. [Section 2.1] In the sentence describing safetensors files, 'as described by by321 (2022)' contains a duplicated 'by'; this should be 'as described by321 (2022)' or 'as described by by321 (2022)' should be rephrased.
  4. [Figure 5] The tag co-occurrence network contains a full-width comma in 'girl,'; this appears to be a typo and should be replaced with a regular comma.
  5. [Section 2.3.4] The profession and country inference using DeepSeek is acknowledged as potentially error-prone, but the paper does not report any validation of the LLM annotations against a human-annotated sample. A short inter-annotator agreement or manual-validation subsection would strengthen confidence in the sunburst-chart results.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline NSFW, demographic, and person-of-interest quantities are direct measurements from collected platform data and external classifiers, not outputs derived from fitted parameters or self-citations.

full rationale

The paper's central claims are empirical measurements rather than derivations. The NSFW ratio trend (41% to 80%) is computed directly from CivitAI's browsing-level labels in the authors' own extended dataset, the female-to-male ratio and age estimates come from the external MiVOLO classifier applied to a sampled subset, and the 15.27% person-of-interest figure is a count of platform-provided POI flags. No equation in the paper maps a fitted parameter into a predicted quantity, and no target result is defined in terms of the inputs that supposedly produce it. The citation to the authors' prior Civiverse work (Palmini et al., 2024) is used only as contextual precedent; the current paper recomputes the trend on its own 40,630,560-image dataset, so the earlier study is not load-bearing evidence for the headline finding. The acknowledged limitations concerning CivitAI's browsing-level classifiers, MiVOLO bias, and age-estimation bias are measurement-validity concerns rather than circularity: the paper claims to describe the platform's labeled content distribution, and those labels are an input measurement, not a construct secretly defined by the conclusion. There is no self-citation chain invoked to force an interpretation, no uniqueness theorem imported from the authors, and no ansatz smuggled in via citation. The derivation chain is therefore self-contained with respect to the paper's stated claims, and the appropriate circularity score is 0.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new free parameters in a modeling sense; the three entries above are hand-set analysis choices that directly shape the reported statistics. The domain assumptions are all made explicit in the text, and the authors acknowledge many of the underlying limitations.

free parameters (3)
  • Caption-source matching threshold = 5 matches
    Captions with fewer than five matches to a system's characteristic vocabulary are labeled 'unknown' (Section 2.4). This hand-set threshold determines the reported 70% Danbooru share.
  • MiVOLO sample size = 40,636 images (0.1%)
    The demographic analysis uses a hand-chosen 0.1% sample described only as temporally balanced; the sample construction is not specified so this choice directly sets the reported ratios.
  • Checkpoint inference seed iteration = unstated
    Section 2.3.2 says seeds were iterated for a representative sample, but no rule is given; this can bias the illustrative gender comparison.
assumptions (4)
  • domain assumption CivitAI NSFW browsing levels are a consistent measure of content explicitness across 2022-2024
    The central NSFW trend (Figure 3) relies on labels assigned by Amazon Rekognition and manual review without testing for drift in classifier thresholds or moderation policy.
  • domain assumption MiVOLO's binary gender and age predictions are sufficiently accurate to support ratio claims
    The 6.24:1 female-to-male ratio and 85.44% NSFW co-occurrence rate depend on MiVOLO; the authors note the classifier ignores non-binary identities and may have uneven error rates.
  • domain assumption LLM-based profession and country inference for deepfake targets is reliable
    Section 2.3.4 uses spaCy NER and DeepSeek to assign professions and countries; uncertain entries are excluded, which can bias the profile of 16,078 individuals.
  • domain assumption Danbooru-derived tag taxonomy accurately categorizes training captions
    Section 2.4 maps training tags to the Danbooru taxonomy to count 'explicit', 'loli/shota', and 'rape' keywords; this assumes the taxonomy matches how tags are used in practice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm." pith.science (2026). https://pith.science/paper/5ZG7J3Z2

@misc{pith2026250504600,
  author       = {Pith},
  title        = {Pith review of: Perpetuating Misogyny with Generative AI: How Model Personalization Normalizes Gendered Harm},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5ZG7J3Z2}},
  note         = {Machine review of arXiv:2505.04600}
}
read the original abstract

Open-source text-to-image (TTI) pipelines have become dominant in the landscape of AI-generated visual content, driven by technological advances that enable users to personalize models through adapters tailored to specific tasks. While personalization methods such as LoRA offer unprecedented creative opportunities, they also facilitate harmful practices, including the generation of non-consensual deepfakes and the amplification of misogynistic or hypersexualized content. This study presents an exploratory sociotechnical analysis of CivitAI, the most active platform for sharing and developing open-source TTI models. Drawing on a dataset of more than 40 million user-generated images and over 230,000 models, we find a disproportionate rise in not-safe-for-work (NSFW) content and a significant number of models intended to mimic real individuals. We also observe a strong influence of internet subcultures on the tools and practices shaping model personalizations and resulting visual media. In response to these findings, we contextualize the emergence of exploitative visual media through feminist and constructivist perspectives on technology, emphasizing how design choices and community dynamics shape platform outcomes. Building on this analysis, we propose interventions aimed at mitigating downstream harm, including improved content moderation, rethinking tool design, and establishing clearer platform policies to promote accountability and consent.

Figures

Figures reproduced from arXiv: 2505.04600 by the authors.

Figure 1
Figure 1. Random sample of images taken from CivitAI in March 2024 with different NSFW levels [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Schematic overview of the CivitAI platform, showing layered relationships between the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 4
Figure 4. Inferred age and gender predictions using MiVOLO on a 0.1% sample of temporally representative images from the 40 million image dataset. without human subjects. The results are shown in [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: Co-occurrence network of the 300 most frequent promotional tags across all assets shared [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Top 10 most popular standalone models (type: checkpoint) as of 19 [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Top 30 most popular model adapters on CivitAI, including the number of downloads, [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Sunburst charts depicting the distribution of 16,078 individuals targeted by deepfake mod [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 9
Figure 9. Figure 9: Subdivision of the 40,000 most downloaded CivitAI assets by type, presence of extractable [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video Deepfake Abuse: How Company Choices Predictably Shape Misuse Patterns

    cs.CY 2025-11 conditional novelty 6.0 of 10

    A few open-weight video models and distribution platforms dominate the creation and spread of NSFW AI video, making developer and platform choices the main intervention points for reducing non-consensual deepfake abuse.

  2. Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)

    cs.AI 2026-05 conditional novelty 5.0 of 10

    The dominant real-world use of generative-image abuse is non-consensual intimate imagery, yet the AI/ML research field focuses almost exclusively on viewer deception.

Reference graph

Works this paper leans on

17 extracted references · 4 canonical work pages · cited by 2 Pith papers

  1. [1]

    Abbate J (2012) Introduction: Rediscovering Women’s History in Computing , chapter

  2. [2]

    The MIT Press, pp. 1–10. DOI:10.7551/mitpress/9014.001.0001. Ajder H, Patrini G, Cavalli F and Cullen L (2019) The state of deepfakes: Landscape, threats, and impact. URL https://regmedia.co.uk/2019/10/08/deepfake_report.pdf . Deeptrace. (Accessed 2 May 2025). Anderson M (2025) Civitai tightens deepfake rules under pressure from mastercard and visa. URL h...

  3. [3]

    London: Palgrave Macmillan UK, pp. 14–26. DOI:10.1007/978-1-349-19798-9

  4. [6]

    Big Data & Society 11(1)

    Cui J and Araujo DA (2024) Rethinking use-restricted open-source licenses for regulating abuse of generative models. Big Data & Society 11(1). DOI:10.1177/20539517241229699. DeepSeek-AI et al. (2025) Deepseek-v3 technical report. DOI:10.48550/arXiv.2412.19437. Doane MA (1991) Femmes Fatales: Feminism, Film Theory, Psychoanalysis . Film - Women’s Studies. ...

  5. [9]

    Oxford University Press, pp. 55–77. DOI: 10.1093/acprof:oso/9780190604981.003.0003. Montani I and Honnibal M (2021) spacy 3: Industrial-strength natural language processing in python. URL https://spacy.io. SpaCy. (Accessed 2 May 2025). Mulvey L (1989) Visual Pleasure and Narrative Cinema , chapter

  6. [11]

    In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (AIES)

    Naik R and Nushi B (2023) Social biases through the text-to-image generation lens. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (AIES). ACM, pp. 786–808. DOI: 10.1145/3600211.3604711. Ng K (2025) Pakistan: Us teen shot dead by father over tiktok videos. URL https://www.bbc.co m/news/articles/cy8pvw3xxxeo. BBC News. (Accessed ...

  7. [15]

    could be categorized as child pornography,

    Leu W, Nakashima Y and Garcia N (2024) Auditing image-based NSFW classifiers for content filtering. In: The 2024 ACM Conference on Fairness, Accountability, and Transparency . ACM, pp. 1163–1173. DOI:10.1145/3630106.3658963. Li Q, Zhang S, Kasper AT, Ashkinaze J, Eaton AA, Schoenebeck S and Gilbert E (2024) Reporting non-consensual intimate media: An audi...

  8. [16]

    id ": 4 0 6 0 3 6 5 1 , 3

    ACM, pp. 620:1–620:25. DOI:10.1145/3613904.3642877. 23 Zhang Y, Jiang L, Turk G and Yang D (2024b) Auditing gender presentation differences in text-to- image models. In: Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO) . ACM, pp. 12:1–12:10. DOI:10.1145/3689904.3694710. Zhao Z, Duan J, Xu K, Wa...

Show all 17 references
  1. [17]

    Y ear Month T otal Images NSFW T rue NSFW F alse Avg

    The table reports the total number of generated images, the number classified as NSFW (Not Safe for Work) and non-NSFW, the average number of images generated per day, and the NSFW ratio (the proportion of NSFW images relative to the total image count). Y ear Month T otal Imag...

  2. [18]

    Wildcards

    The NSFW ratio is calculated as the proportion of NSFW-true images out of the total monthly count. 27 Table S2: Addendum to Section 3.2, MiVOLO Age and Gender Estimations for 2023–2024 Month Total Images Total Persons No Per- sons ♀ (%) ♂ (%) ♀ ♂ ♀ Age (Mean ± SD) ♂ Age (Mean ...

  3. [30]

    Wang W, Bai H, Huang J, Wan Y, Yuan Y, Qiu H, Peng N and Lyu MR (2024) New job, new gender? measuring the social bias in image generation models

    DOI:10.1007/s11229-022-04012-2. Wang W, Bai H, Huang J, Wan Y, Yuan Y, Qiu H, Peng N and Lyu MR (2024) New job, new gender? measuring the social bias in image generation models. In: Proceedings of the 32nd ACM International Conference on Multimedia, MM 2024, Melbourne, VIC, Au...

  4. [37]

    Kim K (2024) KichangKim/DeepDanbooru

    DOI:10.48550/arXiv.2407.06863. Kim K (2024) KichangKim/DeepDanbooru. URL https://github.com/KichangKim/DeepDanbooru. GitHub. (Accessed 26 July 2024). King D (2024) The algorithmic gaze: Representations of women in AI art. URL https://www.le random.art/editorial/the-algorithmic...

  5. [62]

    CivitAI (2023) REST API reference

    DOI:10.1093/geront/gnab167. CivitAI (2023) REST API reference. URLh t t p s : / / g i t h u b . c o m / c i v i t a i / c i v i t a i / w i k i / R E ST -A P I -R e f e r e n c e. GitHub. (Accessed 14 January 2025). CivitAI (2025) Content moderation — civitai. URL https://educ...

  6. [81]

    PMLR, pp. 77–91. Buschek C and Thorp J (2025) Models all the way down. URLhttps://knowingmachines.org/models- all-the-way. Knowing Machines. (Accessed 1 January 2025). by321 (2022) by321/safetensors util. URL https://github.com/by321/safetensors_util . GitHub. (Accessed 26 Jul...

  7. [139]

    8748–8763

    PMLR, pp. 8748–8763. Rombach R, Blattmann A, Lorenz D, Esser P and Ommer B (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). pp. 10684–10695. DOI:10.1109/CVPR52688....

  8. [2023]

    DOI:10.48550/arXiv.2208.01618

    Kigali, Rwanda. DOI:10.48550/arXiv.2208.01618. Gavrilovi´ c Nilsson M, Tzani-Pepelasis C, Ioannou M and Lester D (2019) Understanding the link between sextortion and suicide. International Journal of Cyber Criminology 13(1): 55–69. DOI: 10.5281/zenodo.3402357. Gervais SJ, Vesc...

  9. [2024]

    6949–6958

    ACM, pp. 6949–6958. DOI:10.1145/3664647.3681052. Widder DG, West S and Whittaker M (2023) Open (for business): Big tech, concentrated power, and the political economy of open ai. SSRN Electronic Journal DOI:10.2139/ssrn.4543807. Wu Y, Nakashima Y and Garcia N (2024) Stable dif...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.