{"id":"13d8f503-d142-4ce4-9ebe-e77a846f4634","arxiv_id":"2505.04600","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On CivitAI, NSFW content grew from 41% to 80% of shared images in two years, 15% of model adapters replicate real individuals, and female subjects are portrayed younger and more sexualized than male subjects.","lead":"This study analyzes over 40 million images and 230,000 models from CivitAI, the largest open-source text-to-image model hub. It documents that the share of not-safe-for-work content rose from 41% to 80% over two years, that generated women are far more often depicted as young and sexualized than men, and that over 33,000 model adapters mimic real people.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline NSFW rise (41% to 80%) depends on CivitAI's own browsingLevel labels, and the paper never tests whether CivitAI's classifiers or thresholds drifted over 2023–2024; re-annotating a monthly sample with one fixed classifier would settle it.","rationale":"The reader's weakest assumption identifies the same spot I would: the NSFW trend is both the headline quantitative result and the one that depends on a platform-controlled label whose temporal stability is untested. I checked the alternative candidate weaknesses before settling on this one. The demographic sample (Section 2.2.2) is indeed under-specified, but Table S2 shows monthly sample sizes roughly proportional to monthly volumes, so the sampling concern is less likely to overturn the female-to-male ratio. The checkpoint comparison in Section 2.3.2 uses a vague 'seeds were iterated' rule, but that experiment is qualitative and illustrative, not the paper's central claim. The POI count in Section 2.3.4 depends on CivitAI's `poi` flag and on LLM-based name resolution, but even a factor-of-two error there would not change the core narrative of widespread real-person adapters; the NSFW trend, by contrast, is a precise ratio whose entire increase could be an artifact of moderation changes. The proposed fixed-label re-annotation is a single, feasible check that would either confirm or refute the trend. If the trend survives, the paper's central empirical claim is solid and the CONDITIONAL verdict can be upgraded; if it does not, the headline needs substantial revision. For now, keeping the CONDITIONAL verdict is appropriate, so I recommend no change to the reader's verdict.","tokens_in":25251,"tokens_out":4724,"duration_ms":44836,"concrete_test":"Take a stratified random sample of images from CivitAI for each month from January 2023 through December 2024 (e.g., 1,000 images per month, 24,000 total). Have the same fixed annotation procedure—either one frozen NSFW classifier run on all sampled images, or a small set of human annotators using a written rubric and blind to `browsingLevel`—assign a binary NSFW label to every image. Then compare the fixed-label monthly NSFW ratio to the platform-label series in Table S1. If the fixed-label series still rises from roughly 41% to roughly 80%, the trend is robust; if it is flat, smaller, or shows step changes at dates when CivitAI changed moderation policies, the headline finding is an artifact of label drift rather than a real content shift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative finding is the two-year NSFW ratio increase from 41% (Jan 2023) to 80% (Dec 2024) in Figure 3 and Table S1. This ratio is computed from CivitAI's `browsingLevel` field, which Section 2.1 says is assigned by Amazon Rekognition, unspecified open-source NSFW classifiers, and manual review. The paper treats `browsingLevel` as a stable measurement of content explicitness across the entire period, but provides no evidence that CivitAI kept the same classifier versions, thresholds, or manual-review policies from January 2023 to December 2024. A moderation-system update, a change in Rekognition model versions, or a shift in what the platform sends to manual review could move the NSFW share by tens of percentage points without any change in the content users upload. The paper acknowledges cross-sectional classifier bias (Section 2.2.2, citing Leu et al. 2024) but does not address temporal drift. Since the title's 'disproportionate rise' and the platform-decay argument in Section 3.1 rest on this trend, this is the load-bearing assumption. The demographic and POI results are real measurements, but they do not carry the temporal claim; the NSFW trend does.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an exploratory sociotechnical analysis of CivitAI, a major open-source text-to-image model-sharing platform, using metadata for 40,630,560 images and 231,252 models from late November 2022 to December 2024. It reports a rise in the platform-assigned NSFW ratio from 41% in January 2023 to 80% in December 2024, a female-to-male ratio of 6.24:1 among depicted subjects with women depicted younger and more often in NSFW contexts, a tag and training-metadata analysis showing widespread use of Danbooru-based captioning, and an estimate that 15.27% (33,804) of model adapters are flagged as person-of-interest, i.e., tailored to replicate real individuals. The authors interpret these findings through feminist and constructivist frameworks and propose interventions in content moderation, tool design, and platform policy.","tokens_in":25483,"tokens_out":3907,"duration_ms":39529,"significance":"If the headline trends hold, this is a substantial empirical contribution to the literature on AI governance and gender-based harm: it is one of the largest descriptive studies of a text-to-image model-sharing platform, it covers a two-year period at a scale no prior study has matched, and it connects image-level, model-level, and training-pipeline evidence. The paper is also transparent about several limitations, including acknowledged bias in NSFW classifiers, MiVOLO's binary gender classification, and potential misclassification in the LLM-based profession inference. The main caveat is that the central temporal claim rests on platform-assigned NSFW labels whose stability over time is not demonstrated, and the demographic analysis rests on an underspecified 0.1% sample. These issues are fixable and do not undermine the value of the dataset or the cross-sectional measurements, but they do affect how strongly the paper's central narrative can be asserted.","major_comments":[{"comment":"The headline NSFW increase from 41% in January 2023 to 80% in December 2024 is computed entirely from CivitAI's browsingLevel field, which the paper states (Section 2.1) is assigned by Amazon Rekognition, unspecified open-source NSFW classifiers, and manual review. The paper does not test whether classifier versions, classification thresholds, or manual-review policies drifted over the two-year period. This is load-bearing because the platform-decay argument in Section 3.1 and the phrase 'disproportionate rise' rest directly on this trend. The paper cites Leu et al. (2024) for cross-sectional classifier bias but does not address temporal drift. I ask the authors to re-annotate a monthly stratified sample of images with one fixed classifier and present both the platform-assigned series and the re-annotated series; at minimum, they should document any known CivitAI moderation-policy changes during 2023-2024 and perform a sensitivity analysis using alternative NSFW thresholds. Without such a check, the temporal trend may partly reflect a change in measurement rather than a change in uploaded content.","section":"Section 2.2.1, Figure 3, Table S1"},{"comment":"The demographic analysis is based on a 0.1% subsample described only as 'temporally balanced and distributionally representative' (Section 2.2.2). No sampling protocol is given: the text does not specify the strata (e.g., month, NSFW level, or image type), the random seed, the exclusion criteria, or the confidence threshold used to discard images with no detected human subjects. Table S2 suggests that monthly sample sizes scale roughly with total image volume, but the reader cannot determine whether images were drawn uniformly within months or whether any NSFW/SFW balancing was applied. Because the female-to-male ratio, the mean ages, and the NSFW co-occurrence rates are all derived from this sample, the authors should specify the exact sampling procedure, release the sample image identifiers, and report the number of images excluded at each step.","section":"Section 2.2.2, Table S2"},{"comment":"The checkpoint inference experiment compares outputs from the ten most downloaded checkpoints using the prompts 'woman' and 'man', but the caption states that 'initial noise seeds were iterated to ensure a representative, uncurated sample' without specifying the iteration rule, the number of seeds tried, or the criteria for stopping. This makes the comparison non-reproducible and raises the possibility of selective presentation. Since this experiment is illustrative rather than central to the paper's main claims, the issue is not blocking, but the text should either state the exact seed-iteration protocol or present results over a fixed set of seeds.","section":"Section 2.3.2, Figure 6"}],"minor_comments":[{"comment":"The sentence 'The dataset covers 40,630,560 images and 231,252 models as of 19 th January' is missing a year; it should read 'as of January 19, 2025' to match the later text.","section":"Section 2.1"},{"comment":"The caption says 'as of 19 th January 2024', while the main text (Section 2.3.2) says the checkpoints were selected 'as of January 19 th, 2025'. The caption year appears to be a typo and should be corrected.","section":"Figure 6 caption"},{"comment":"In the sentence describing safetensors files, 'as described by by321 (2022)' contains a duplicated 'by'; this should be 'as described by321 (2022)' or 'as described by by321 (2022)' should be rephrased.","section":"Section 2.1"},{"comment":"The tag co-occurrence network contains a full-width comma in 'girl，'; this appears to be a typo and should be replaced with a regular comma.","section":"Figure 5"},{"comment":"The profession and country inference using DeepSeek is acknowledged as potentially error-prone, but the paper does not report any validation of the LLM annotations against a human-annotated sample. A short inter-annotator agreement or manual-validation subsection would strengthen confidence in the sunburst-chart results.","section":"Section 2.3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper's central temporal claim is the most fragile part, and it is also the part most likely to be quoted by policymakers and journalists. If the authors cannot obtain a fixed-classifier re-annotation, they should soften the 'rise' language and present the trend as a rise in platform-assigned NSFW labels, not necessarily in content explicitness. The dataset release and the multi-level analysis are valuable regardless, so this is a revision rather than a rejection issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the most complete public quantification of CivitAI to date: 40M images, 231K models, a two-year time window, and a dataset that is publicly available. Second, the central trend—NSFW share rising from 41% to 80%—is credible in direction, but the exact numbers rest on CivitAI's own browsingLevel labels, and the paper never tests whether those labels stayed stable over 2023–2024.\n\nWhat is actually new: the full two-year trend line, the MiVOLO demographic sample (6.24:1 female-to-male ratio, women younger and more often NSFW), the 33,804 \"person of interest\" adapters with LLM-profiled targets, and the training-tag provenance result showing Danbooru-style tags dominate. Wei et al. and Palmini et al. already documented NSFW growth and abusive models, but this extends the window and adds layers they did not. The dataset release is a genuine contribution.\n\nThe soft spots are real but addressable. The headline NSFW trend is load-bearing and depends on platform-assigned labels (Amazon Rekognition plus unspecified open-source classifiers and manual review). The paper acknowledges cross-sectional classifier bias but does not test for temporal drift; a moderation-system update or threshold shift could move the curve by tens of points. Re-annotating a monthly sample with one fixed classifier would settle it. The 0.1% demographic sample is described as \"temporally balanced\" but the sampling procedure is under-specified. The checkpoint inference experiment says seeds were iterated to ensure representativeness, without stating a rule. The POI count also relies on a platform flag, though the paper does a reasonable job of interpreting it. The DeepSeek profiling is explicitly hedged. None of this sinks the paper; it means the precise figures should be cited with caution.\n\nThe interpretive sections on platform decay and misogyny are standard sociotechnical framing, and the authors do not lean on them as if they were measurements.\n\nWho this is for: anyone working on AI governance, content moderation, or gender and AI. It deserves a serious referee. My recommendation: send it to review, and ask for a re-annotation stability check, a clearer sampling description, and a seed-iteration rule.","headline":"The most complete CivitAI audit to date, with a credible but label-dependent NSFW trend; the numbers need a stability check before they are cited as exact.","tokens_in":26055,"tokens_out":3021,"would_cite":true,"duration_ms":28798,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"On the open-source image-model hub CivitAI, explicit content rose from 41% to 80% of posted images in two years, while 15.27% of all model adapters are built to mimic real people.","keywords":["generative AI","text-to-image models","CivitAI","model personalization","NSFW content","deepfakes","gender bias","platform decay"],"falsifier":"Take a random sample of images from each month between January 2023 and December 2024, have annotators apply a single fixed explicitness rubric blind to date, and recompute the monthly share of explicit images. If the fixed-rubric share stays roughly flat while the platform-labeled share climbs from 41% to 80%, the headline trend would be shown to be an artifact of classifier or moderation drift.","tokens_in":1815,"feed_emoji":"🤖","tokens_out":2849,"duration_ms":97439,"temperature":0.7,"pith_summary":"This paper is an empirical study of CivitAI, the largest hub for sharing and downloading open-source text-to-image models and the small adapters that personalize them. Drawing on metadata for more than 40 million images and over 230,000 models, it argues that the platform's content has shifted sharply toward explicit imagery, that female-presenting subjects are over-represented, younger, and more often sexualized, and that a substantial share of adapters are designed to replicate identifiable real people. The authors read these patterns as a self-reinforcing sociotechnical cycle rather than isolated misuse: engagement incentives reward sensational content, subculture-derived auto-tagging systems inject sexualized vocabulary into model training, and policy gaps let deepfake adapters circulate. If the picture is right, open-source image generation is not merely reflecting existing biases but actively normalizing misogynistic and exploitative imagery at scale.","feed_headline":"Explicit AI images jumped from 41% to 80% in two years","feed_subtitle":"A 40-million-image study finds one in seven model adapters is built to copy a real person.","key_machinery":"The machinery that carries the argument has three parts. First, the model adapter—chiefly LoRA (low-rank adaptation, a small trainable module that steers a large image model) and textual inversion—is what lets a user replicate a specific face, body, or style without retraining a foundation model, turning a shared model into a multiplier for unlimited derivative images. Second, the \"person of interest\" flag in CivitAI metadata identifies which adapters are tailored to real people, giving the authors a population-level count of deepfake-ready assets. Third, the auto-tagging pipeline: captioning systems such as CLIP Interrogator and Danbooru-based taggers supply the text used to train adapters, and the paper shows that more than 70% of adapters with extractable captions used Danbooru-derived tags, whose sexualized vocabulary normalizes explicit content, with 5.6% containing \"loli\"/\"shota\" and 2.1% containing \"rape\". These three components together convert individual user choices into systematic, platform-wide bias.","core_discovery":"On the paper's own terms, the central finding is that CivitAI has become an explicit-content pipeline: the share of posted images labeled NSFW rose from 41% in January 2023 to 80% by December 2024, and image subjects inferred as female outnumber male subjects by 6.24 to 1, are younger on average (23.67 versus 31.16 years), and appear in NSFW contexts more often (85.44% versus 73.66%). It further reports that 15.27% of model adapters (33,804) carry a \"person of interest\" flag indicating they are tailored to replicate a real individual, with actors, adult performers, and online personalities among the most common targets. The paper attributes these patterns to a layered mechanism: popularity-based feedback loops favor sensational content, and auto-tagging systems built on anime-imageboard vocabularies—used by more than 70% of adapters with identifiable training captions—embed sexualized, occasionally abusive label sets into model training. The authors conclude that personalization hubs are not neutral distribution points but active amplifiers that normalize gendered harm.","pith_inferences":["An implicit testable implication: the 41-to-80 trend should be re-measured on a fixed, human-labeled sample to separate genuine content change from drift in the platform's own classification thresholds.","The training-tag finding suggests a natural experiment: comparing two otherwise identical LoRA training runs, one captioned with a Danbooru-based tagger and one with a neutral captioner, could isolate how much of the gendered NSFW skew comes from captioning vocabulary rather than from the underlying images.","If the celebrity-creator leaderboard and tag-based discoverability are as influential as the paper suggests, removing popularity-sorted feeds or sexualized default tags would provide a direct test of whether platform design, not user demand alone, drives the rise.","The demographic measurements rely on a binary gender classifier; with more inclusive annotation the 6.24:1 ratio might shift, but the qualitative pattern of younger female subjects in explicit contexts is unlikely to vanish."],"forward_implications":["If the current trajectory holds, the NSFW share on model-sharing hubs will continue to climb because engagement incentives reward sensational content, making the rise a structural feature rather than a removable nuisance.","Deepfake adapters remain available after download and can be redistributed outside the platform, so policy changes that only affect on-platform display or tags will not stop non-consensual imagery from spreading.","Because most adapters are trained with Danbooru-based auto-tagging, changing the default captioning tool used in training interfaces is a concrete intervention point for reducing explicit and abusive training vocabularies.","Existing image-perturbation defenses lose effectiveness against newer architectures and against LoRA fine-tuning, so protective tools must be continuously updated to keep pace with model development.","The biases the paper documents are embedded in models rather than only in individual images, so any downstream consumer product that integrates these open-source models or adapters would inherit the skewed representations."],"supporting_citations":[{"why":"Established the earlier NSFW-ratio measurement on the same platform that this study extends to a two-year window.","marker":"Palmini et al., 2024"},{"why":"Showed that NSFW images receive more social engagement on CivitAI, the load-bearing evidence for the engagement feedback loop.","marker":"Wei et al., 2024"},{"why":"Audited NSFW classifiers and found systematic misclassification of female subjects, the limitation that qualifies the demographic NSFW numbers.","marker":"Leu et al., 2024"},{"why":"Supplies the MiVOLO model used to infer age and gender from generated images.","marker":"Kuprashevich and Tolstykh, 2023"},{"why":"Introduced LoRA, the adapter method that makes personalized model-sharing possible and is the paper's central object.","marker":"Hu et al., 2022"},{"why":"Introduced textual inversion, the other main adapter technique behind model personalization on the platform.","marker":"Gal et al., 2023"},{"why":"Released Stable Diffusion, the open-source foundation model that the CivitAI ecosystem and its checkpoints are built on.","marker":"Rombach et al., 2022"},{"why":"Documented misogyny, pornography, and malignant stereotypes in web-scraped multimodal datasets, grounding the paper's argument about training-data toxicity.","marker":"Birhane et al., 2021"},{"why":"Provided the enshittification/platform-decay concept that frames the paper's interpretation of the NSFW rise.","marker":"Doctorow (2025)"}],"fun_headline_variants":["Explicit images on CivitAI rose from 41% to 80% by 2024","1 in 7 AI adapters built to copy real individuals","Female AI subjects outnumber males 6:1 and are younger","Anime tag vocabularies train sexualized content into AI","Design feedback loops normalize misogyny in AI image models"],"cache_read_input_tokens":28160,"weakest_assumption_plain":"The paper's central trend depends on the assumption that the platform's explicit-content labels were equally strict in 2023 and 2024; if the labeler's thresholds shifted, part of the 41-to-80 surge could be classification drift rather than a change in what users actually post.","fun_headline_variants_meta":{"raw":{"variants":["Explicit images on CivitAI rose from 41% to 80% by 2024","1 in 7 AI adapters built to copy real individuals","Female AI subjects outnumber males 6:1 and are younger","Anime tag vocabularies train sexualized content into AI","Design feedback loops normalize misogyny in AI image models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000307,"raw_usage":{"total_tokens":1792,"prompt_tokens":1013,"completion_tokens":779,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":685}},"tokens_in":629,"tokens_out":779,"duration_ms":6858,"temperature":1.0,"reasoning_tokens":685,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:23:57.775267+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of images from each month between January 2023 and December 2024, have annotators apply a single fixed explicitness rubric blind to date, and recompute the monthly share of explicit images. If the fixed-rubric share stays roughly flat while the platform-labeled share climbs from 41% to 80%, the headline trend would be shown to be an artifact of classifier or moderation drift.","supporting_citations":[],"review_version":1}