{"id":"8a8490f4-5fb6-4809-a7d4-98359b9a2326","arxiv_id":"2506.14018","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"GAI tools' content moderation policies are comprehensive in scope but thin on user reporting and appeals, and Reddit users report frequent frustration with opaque moderation decisions.","lead":"This paper studied the content moderation policies of 14 popular AI chat and image tools, then read Reddit discussions about how users experience moderation in those tools. It finds that the policies are broad but vague, and that users often get blocked without a clear reason or a working appeal.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'pervasively' and 'frequently' claims outrun the self-selected, keyword-filtered Reddit data; the paper's own §5.3 acknowledges this, so the strongest wording should be qualified or tested against an unfiltered sample.","rationale":"The reader's weakest assumption identifies the representativeness of the Reddit sample as the key vulnerability in the abstract's frequency claims. My analysis agrees, and I locate the specific mechanism: the keyword-based corpus construction (§5.1) and the random sample of 130 posts (§5.2) cannot support 'pervasively' or 'frequently' because the sampling frame is conditioned on the presence of moderation-specific vocabulary, which is inherently more likely to appear in posts about negative experiences. The paper's own limitation statement (§5.3) concedes that the study 'may not have assessed all successes and failures,' which directly undercuts the quantitative-sounding adverbs in the abstract. This is not a fatal flaw in the qualitative contribution—the policy analysis is thorough, the Reddit coding is transparent, and the failure themes (false positives, inconsistency, opaque decisions, poor appeals) are credible and well-illustrated. But the central claim as worded overstates the evidentiary basis. The proposed test—a random unfiltered sample coded with the same codebook—would provide the missing denominator and either validate the frequency language or force a hedge. Since the reader already recommends a conditional verdict, my concern does not change that verdict; it strengthens the rationale for qualifying the abstract's strongest wording.","tokens_in":24857,"tokens_out":6205,"duration_ms":62228,"concrete_test":"Draw a random sample of 200 posts from each of the seven subreddits in Table 2 (n=1,400) without applying any keyword filter. Code them with the same codebook for the presence of (a) user-reported moderation successes and (b) user-reported moderation failures. Compute the proportion of posts in each category with confidence intervals and compare to the proportions in the keyword-filtered corpus. If the failure proportion in the unfiltered sample is substantially lower (e.g., less than half) than in the filtered sample, the 'frequently' claim is an artifact of selection; if failure narratives appear in a substantial share of unfiltered posts (e.g., >10%), the claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that 'moderation systems succeeded in blocking malicious generations pervasively, users frequently experienced frustration'—rests entirely on qualitative coding of 130 posts drawn from a keyword-filtered corpus (§5.1, §5.2). The keyword list is skewed toward restriction terms ('ban', 'block', 'suspend', 'flag'), which systematically over-captures negative moderation experiences and under-captures routine successes that users never mention. The paper itself states in §5.3 that 'our study may not have assessed all successes and failures around content moderation policy enforcement,' yet the abstract uses frequency adverbs ('pervasively', 'frequently') that imply quantitative prevalence without any denominator. Because the recommendations in §7 (soft moderation, transparency, appeals) are motivated by the presumed scale of these phenomena, the unsupported frequency language is load-bearing. A random unfiltered sample from the same subreddits would provide the missing baseline and test whether the observed balance of success and failure narratives is an artifact of selection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates content moderation in consumer-facing generative AI (GAI) online tools through two studies. Study 1 analyzes the content moderation policies of 14 US-based GAI tools, finding that policies are comprehensive in covering moderation criteria, methodology, and consequences, but lack detail on user-driven moderation and appeals. Study 2 analyzes 130 randomly sampled Reddit posts (with their comments) from keyword-filtered discussions about moderation experiences in AIGC creative tasks, identifying both successes in blocking malicious generations and failures in moderation decision-making and post-moderation user support. The authors propose policy and product improvements such as unified policy structures, soft moderation, personalized guardrails, and more transparent moderation pipelines.","tokens_in":25028,"tokens_out":3409,"duration_ms":39225,"significance":"If its claims are suitably qualified, this is a useful contribution to HCI and AI-safety research. The paper provides one of the first empirical mappings of GAI product content moderation policies and user experiences, complementing model-level safety auditing work. It makes two datasets publicly available (the policy corpus and the filtered Reddit post dataset) and follows a transparent coding process, with 44 of 51 policy pages double-coded and all 130 Reddit posts double-coded. The qualitative findings—particularly the gap between policy detail and user-perceived enforcement, and the lack of meaningful appeal mechanisms—are valuable and actionable for designers and policymakers. The main weakness is that the abstract and several findings sections use quantitative prevalence language that the sample design cannot support; this is correctable within the manuscript's scope and does not undermine the qualitative core.","major_comments":[{"comment":"The abstract's claim that moderation systems succeeded 'pervasively' and that users 'frequently experienced frustration' is not supported by the study design. The underlying data are 130 posts from a keyword-filtered, self-selected corpus of Reddit discussions, which has no unbiased denominator and systematically over-captures negative experiences. The paper itself acknowledges in §5.3 that the study 'may not have assessed all successes and failures around content moderation policy enforcement.' These frequency adverbs are load-bearing because they define the paper's central takeaway. I recommend rewording the abstract and the corresponding passages in §6 to describe the types of experiences observed (e.g., 'users reported both successes and failures') or, alternatively, adding an explicit within-sample quantitative analysis with clear caveats about selection bias.","section":"Abstract; §5.3"},{"comment":"The keyword list contains 16 terms, most of which are restriction-oriented: 'moderate,' 'censor,' 'ban,' 'block,' 'suspend,' 'restrict,' 'warn,' 'flag,' 'appeal,' 'violate,' 'terminate,' 'remove,' 'content policy,' 'guardrail,' 'filter,' and 'refuse.' This selection strategy will over-capture negative moderation experiences and under-capture routine successes that users never mention. Consequently, the paper's comparative statements—e.g., that successes were 'pervasive' while failures were 'frequent'—may be artifacts of which posts users choose to share and which keywords the collection used. Please either temper such comparative claims or provide a baseline comparison from an unfiltered sample to assess the selection effect.","section":"§5.1 (Keyword List Creation)"},{"comment":"Several findings use fractions such as '5/6 tools' to characterize the reach of a phenomenon (e.g., 'Failure in Mitigating False-Positive Rate (5/6 tools)'). As written, these counts can be read as prevalence estimates across the six GAI tools, but they are actually counts of tools for which the theme appeared somewhere in the 130 sampled posts and their comments. The sample is not representative of any tool's user base, and the keyword filter further biases detection of negative themes. Please clarify in the text that these fractions are descriptive within the coded sample, not population estimates, and adjust any associated frequency wording.","section":"§6.1 and §6.2.1 (counts such as '5/6 tools')"}],"minor_comments":[{"comment":"The text states that the authors recorded '52 PDF files of 51 pages,' which is mildly confusing; please clarify whether one page spanned two PDF files or whether a duplicate was retained.","section":"§3.2"},{"comment":"The abstract and introduction describe the study's scope as 'content moderation in GAI online tools' without consistently noting that Study 2 focuses specifically on AIGC creative tasks. Since the findings about 'frustration' and 'failures' in the abstract are drawn entirely from that narrower activity, the scope should be stated in the abstract to avoid overgeneralization.","section":"Abstract and §1"},{"comment":"The exclusion of r/NovelAI is justified by the authors' observation that NovelAI enforces almost no content moderation, but this exclusion, combined with the limited subreddit list, further narrows the set of tools studied; a sentence acknowledging the potential effect on the diversity of experiences would be helpful.","section":"§5.1 (Subreddits Choice)"}],"recommendation":"major_revision","confidential_remarks":"The use of the Schaffner et al. framework by co-authors who are also authors on that prior work is a shared self-citation, but the empirical coding here is clearly independent and the issue is not circular; I would not mention it as a blocking concern. The main revision needed is to align the abstract and findings with the qualitative nature of the evidence. The paper would also benefit from a brief quantitative summary of the 130 coded posts (e.g., how many posts contained any success narrative vs. any failure narrative) with the selection-bias caveats explicitly stated. The scope limitation to AIGC creative tasks should be more visible, since it materially bounds the 'user experiences' contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you need to know: this is one of the first papers to look at content moderation in consumer GAI products from both the policy side and the user side. The policy study is the stronger half; the Reddit study is suggestive but the abstract oversells it with frequency words.\n\nThe genuinely new pieces are the 14-tool policy dataset and the 1123-post Reddit corpus across six tools. The policy coding is careful: 44 of 51 pages coded by at least two coders, and the codebook is public. The finding that policies are comprehensive yet vague on reporting and appeals is well supported, and the concrete examples (e.g., DreamStudio telling users to just 'try again') are effective. The Reddit analysis uses a random sample of 130 posts with iterative double coding and reports saturation—reasonable qualitative practice.\n\nThe soft spot is the abstract's wording: 'pervasively' succeeded and 'frequently' frustrated. The keyword list leans on restriction terms like 'ban,' 'block,' 'suspend,' 'flag,' so the corpus over-captures complaints and under-captures routine successes that users never post about. There is no denominator, so the frequency claims are not supported by the sample design. The authors acknowledge this in §5.3, but the abstract and discussion use the stronger language anyway. The qualitative findings about transparency and appeal failures stand; the scale claims need an unfiltered sample or explicit hedging.\n\nOther minor issues: the tool list is US-only (acknowledged), and the lack of IRR is common in this kind of iterative qualitative work, so I don't hold it against them.\n\nVerdict: worth a serious referee. It's a useful, timely empirical study with public data. The main revision should be aligning the abstract with the evidence—drop or qualify 'pervasively' and 'frequently,' or add a baseline. I'd cite it for the policy dataset, and if the group works on HCI/AI safety, it's worth reading. Recommendation: send to peer review with a request to fix the frequency claims.","headline":"A solid first map of GAI moderation policies and user experiences; the policy analysis is careful, but the abstract's frequency claims outrun the self-selected Reddit sample.","tokens_in":25573,"tokens_out":2249,"would_cite":true,"duration_ms":22678,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Generative AI content filters block abuse but leave users without an effective appeal.","keywords":["content moderation","generative AI","policy analysis","user experience","Reddit","AIGC creative tasks","appeals","transparency"],"falsifier":"A longitudinal telemetry study inside one or more GAI products—logging every moderation decision, its true or false positive status, and the resolution of each appeal—would settle the central claim; if false-positive moderation and unresolved appeals are rare in system logs, the paper's frequency claims about user frustration would be falsified.","tokens_in":24648,"feed_emoji":"🛡️","tokens_out":6257,"duration_ms":57479,"temperature":0.7,"pith_summary":"This paper examines content moderation in consumer generative AI products, asking how the systems perform for everyday users rather than at the model level. The authors analyze the content moderation policies of 14 web-based text and image generation tools and then read thousands of Reddit discussions about users' moderation experiences with six of those tools. They conclude that moderation systems succeed pervasively in blocking malicious generations, yet users frequently experience frustrating failures: harmless prompts get flagged, similar requests get different treatment, and after moderation users get generic or hallucinated explanations, unclear criteria, and almost no working appeal channel. The paper argues that GAI products inherited the policy structure of online community moderation but not the user-facing machinery of reporting and appeals.","feed_headline":"AI content filters block abuse but frustrate users","feed_subtitle":"A two-part study of 14 generative AI tools' policies and thousands of Reddit posts documents a missing appeals process.","key_machinery":"The central object is the content moderation pipeline of a GAI online tool, analyzed through a three-part policy framework: moderation criteria (what content is forbidden), moderation methodology (how problematic content is detected), and moderation consequences (what happens to content and users). The paper's analytic move is to map user-experienced successes and failures onto that framework, using Reddit discussions of AIGC creative tasks to see where enforcement diverges from lived experience. The comparative lens is the analogy to online communities: GAI products copied the policy structure but not the user-side apparatus of reporting and appeals.","core_discovery":"The paper's central claim is that content moderation in generative AI products is a two-sided story: it blocks clearly malicious content at scale, but fails users precisely at the moments that determine trust—explaining decisions, allowing appeals, and judging context. Based on a qualitative policy analysis of 14 GAI online tools and a thematic analysis of 130 randomly sampled Reddit posts about creative generation tasks, the authors find that policies comprehensively outline what is forbidden, how violations are detected, and what consequences follow, yet omit concrete details on user reporting and appeals. User discussions show the same split: widespread appreciation that harmful requests are blocked, alongside frequent false positives, inconsistent decisions, context-blind censorship of ordinary fiction and art, opaque explanations that are sometimes generated by the model itself, and appeal processes that are slow or silent. The paper treats these failures as a structural gap in the moderation pipeline, not as isolated bugs.","pith_inferences":["Because the Reddit sample is self-selected and frustration-heavy, the paper's 'pervasive success' of moderation is likely conservative for the broader population; a representative user survey would probably show even higher block rates and lower complaint rates than the forums suggest.","The findings imply a testable design claim: providing stage-specific, concrete explanations of moderation decisions should measurably reduce perceived unfairness and improve retention, which an A/B test could verify.","The paper's focus on creative generation tasks leaves open how moderation failures differ in dialogue, search, and coding tasks, where false positives may take different forms.","If regulators require transparency and redress for automated content decisions, the gaps documented here, such as no banned-word list, no general appeal, and opaque grounds, could become legal liabilities for GAI providers."],"forward_implications":["GAI products should consolidate scattered moderation rules into a single dedicated policy page, separating tool-specific rules from rules that govern sharing in associated online communities.","Products should establish clear, step-by-step reporting and appeal channels for all moderation decisions, not only copyright-related takedowns.","Moderation systems should prefer soft moderation, such as warnings, content masking, or modified output, over outright denial when there is a chance of a false positive.","Policies should disclose implementation details such as banned-word lists and explain moderation decisions at the specific stage where they occur, so users can judge whether a decision was justified."],"supporting_citations":[{"why":"Supplies the policy-comprehensiveness framework and the collection method that Study 1 adapts to GAI tools.","marker":"[68]"},{"why":"Provides the three-component model of criteria, methodology, and consequences used to organize the policy analysis.","marker":"[74]"},{"why":"Documents over 120 prohibited behaviors in foundation-model AUPs, the baseline for the paper's findings on scattered and vague criteria.","marker":"[47]"},{"why":"Establishes the user-perspective framework of censorship, suspension, and shadowbanning that the Reddit analysis extends to GAI products.","marker":"[59]"},{"why":"Audits GPT's content moderation guardrails and shows over-censorship of cultural content, aligning with user-reported failures.","marker":"[54]"},{"why":"Shows text-to-image moderation over-censors cultural content, corroborating the paper's user-experience failures.","marker":"[62]"},{"why":"Finds users are most frustrated by direct denials without explanations, supporting the paper's soft-moderation recommendation.","marker":"[88]"},{"why":"Models how to study user reactions to content removal on Reddit, informing the paper's Reddit methodology.","marker":"[37]"}],"fun_headline_variants":["AI filters block abuse but fail on appeals","Content moderation in GAI: strict, opaque, frustrating","14 AI tools' policies miss the appeals gap","AI moderation blocks harm, then frustrates users","Generative AI moderation: wins and silent appeals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that self-selected Reddit discussions give a representative window onto how often moderation succeeds and fails, so that qualitative examples can support frequency claims like 'pervasively' and 'frequently'; if Reddit overrepresents frustrated users, the success rate may be overstated.","fun_headline_variants_meta":{"raw":{"variants":["AI filters block abuse but fail on appeals","Content moderation in GAI: strict, opaque, frustrating","14 AI tools' policies miss the appeals gap","AI moderation blocks harm, then frustrates users","Generative AI moderation: wins and silent appeals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000168,"raw_usage":{"total_tokens":1233,"prompt_tokens":893,"completion_tokens":340,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":268}},"tokens_in":509,"tokens_out":340,"duration_ms":4810,"temperature":1.0,"reasoning_tokens":268,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:23:37.456835+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal telemetry study inside one or more GAI products—logging every moderation decision, its true or false positive status, and the resolution of each appeal—would settle the central claim; if false-positive moderation and unresolved appeals are rare in system logs, the paper's frequency claims about user frustration would be falsified.","supporting_citations":[{"cited_title":"community guidelines make this the best party on the internet","cited_arxiv_id":null,"evidence_quote":"Supplies the policy-comprehensiveness framework and the collection method that Study 1 adapts to GAI tools."},{"cited_title":"Acceptable use policies for foundation models","cited_arxiv_id":null,"evidence_quote":"Documents over 120 prohibited behaviors in foundation-model AUPs, the baseline for the paper's findings on scattered and vague criteria."},{"cited_title":"Censored, suspended, shadow- banned: User interpretations of content moderation on social media platforms","cited_arxiv_id":null,"evidence_quote":"Establishes the user-perspective framework of censorship, suspension, and shadowbanning that the Reddit analysis extends to GAI products."},{"cited_title":"Auditing gpt’s content moderation guardrails: Can chatgpt write your favorite tv show? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 660–686, 2024","cited_arxiv_id":null,"evidence_quote":"Audits GPT's content moderation guardrails and shows over-censorship of cultural content, aligning with user-reported failures."},{"cited_title":"Exploring the Boundaries of Content Moderation in Text-to-Image Generation","cited_arxiv_id":"2409.17155","evidence_quote":"Shows text-to-image moderation over-censors cultural content, corroborating the paper's user-experience failures."},{"cited_title":"Contestability for content moderation","cited_arxiv_id":null,"evidence_quote":"Finds users are most frustrated by direct denials without explanations, supporting the paper's soft-moderation recommendation."},{"cited_title":"did you suspect the post would be removed?","cited_arxiv_id":null,"evidence_quote":"Models how to study user reactions to content removal on Reddit, informing the paper's Reddit methodology."}],"review_version":1}