Pith. sign in

REVIEW 3 major objections 3 minor 96 references

"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products

T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Generative AI content filters block abuse but leave users without an effective appeal.

desk verdict A solid first map of GAI moderation policies and user experiences; the policy analysis is careful, but the abstract's frequency claims outrun the self-selected Reddit sample. read the letter →

arxiv 2506.14018 v1 pith:RRJWAIDO submitted 2025-06-16 cs.HC

classification cs.HC
keywords contentmoderationgenerativeAIpolicyanalysisuserexperienceRedditAIGCcreativetasksappealstransparency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper examines content moderation in consumer generative AI products, asking how the systems perform for everyday users rather than at the model level. The authors analyze the content moderation policies of 14 web-based text and image generation tools and then read thousands of Reddit discussions about users' moderation experiences with six of those tools. They conclude that moderation systems succeed pervasively in blocking malicious generations, yet users frequently experience frustrating failures: harmless prompts get flagged, similar requests get different treatment, and after moderation users get generic or hallucinated explanations, unclear criteria, and almost no working appeal channel. The paper argues that GAI products inherited the policy structure of online community moderation but not the user-facing machinery of reporting and appeals.

What carries the argument

The central object is the content moderation pipeline of a GAI online tool, analyzed through a three-part policy framework: moderation criteria (what content is forbidden), moderation methodology (how problematic content is detected), and moderation consequences (what happens to content and users). The paper's analytic move is to map user-experienced successes and failures onto that framework, using Reddit discussions of AIGC creative tasks to see where enforcement diverges from lived experience. The comparative lens is the analogy to online communities: GAI products copied the policy structure but not the user-side apparatus of reporting and appeals.

What would settle it

A longitudinal telemetry study inside one or more GAI products—logging every moderation decision, its true or false positive status, and the resolution of each appeal—would settle the central claim; if false-positive moderation and unresolved appeals are rare in system logs, the paper's frequency claims about user frustration would be falsified.

Watch

Extended reading notes

Core claim

The paper's central claim is that content moderation in generative AI products is a two-sided story: it blocks clearly malicious content at scale, but fails users precisely at the moments that determine trust—explaining decisions, allowing appeals, and judging context. Based on a qualitative policy analysis of 14 GAI online tools and a thematic analysis of 130 randomly sampled Reddit posts about creative generation tasks, the authors find that policies comprehensively outline what is forbidden, how violations are detected, and what consequences follow, yet omit concrete details on user reporting and appeals. User discussions show the same split: widespread appreciation that harmful requests are blocked, alongside frequent false positives, inconsistent decisions, context-blind censorship of ordinary fiction and art, opaque explanations that are sometimes generated by the model itself, and appeal processes that are slow or silent. The paper treats these failures as a structural gap in the moderation pipeline, not as isolated bugs.

Load-bearing premise

The load-bearing premise is that self-selected Reddit discussions give a representative window onto how often moderation succeeds and fails, so that qualitative examples can support frequency claims like 'pervasively' and 'frequently'; if Reddit overrepresents frustrated users, the success rate may be overstated.

Editorial extensions

If this is right

  • GAI products should consolidate scattered moderation rules into a single dedicated policy page, separating tool-specific rules from rules that govern sharing in associated online communities.
  • Products should establish clear, step-by-step reporting and appeal channels for all moderation decisions, not only copyright-related takedowns.
  • Moderation systems should prefer soft moderation, such as warnings, content masking, or modified output, over outright denial when there is a chance of a false positive.
  • Policies should disclose implementation details such as banned-word lists and explain moderation decisions at the specific stage where they occur, so users can judge whether a decision was justified.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the Reddit sample is self-selected and frustration-heavy, the paper's 'pervasive success' of moderation is likely conservative for the broader population; a representative user survey would probably show even higher block rates and lower complaint rates than the forums suggest.
  • The findings imply a testable design claim: providing stage-specific, concrete explanations of moderation decisions should measurably reduce perceived unfairness and improve retention, which an A/B test could verify.
  • The paper's focus on creative generation tasks leaves open how moderation failures differ in dialogue, search, and coding tasks, where false positives may take different forms.
  • If regulators require transparency and redress for automated content decisions, the gaps documented here, such as no banned-word list, no general appeal, and opaque grounds, could become legal liabilities for GAI providers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper investigates content moderation in consumer-facing generative AI (GAI) online tools through two studies. Study 1 analyzes the content moderation policies of 14 US-based GAI tools, finding that policies are comprehensive in covering moderation criteria, methodology, and consequences, but lack detail on user-driven moderation and appeals. Study 2 analyzes 130 randomly sampled Reddit posts (with their comments) from keyword-filtered discussions about moderation experiences in AIGC creative tasks, identifying both successes in blocking malicious generations and failures in moderation decision-making and post-moderation user support. The authors propose policy and product improvements such as unified policy structures, soft moderation, personalized guardrails, and more transparent moderation pipelines.

Significance. If its claims are suitably qualified, this is a useful contribution to HCI and AI-safety research. The paper provides one of the first empirical mappings of GAI product content moderation policies and user experiences, complementing model-level safety auditing work. It makes two datasets publicly available (the policy corpus and the filtered Reddit post dataset) and follows a transparent coding process, with 44 of 51 policy pages double-coded and all 130 Reddit posts double-coded. The qualitative findings—particularly the gap between policy detail and user-perceived enforcement, and the lack of meaningful appeal mechanisms—are valuable and actionable for designers and policymakers. The main weakness is that the abstract and several findings sections use quantitative prevalence language that the sample design cannot support; this is correctable within the manuscript's scope and does not undermine the qualitative core.

major comments (3)
  1. [Abstract; §5.3] The abstract's claim that moderation systems succeeded 'pervasively' and that users 'frequently experienced frustration' is not supported by the study design. The underlying data are 130 posts from a keyword-filtered, self-selected corpus of Reddit discussions, which has no unbiased denominator and systematically over-captures negative experiences. The paper itself acknowledges in §5.3 that the study 'may not have assessed all successes and failures around content moderation policy enforcement.' These frequency adverbs are load-bearing because they define the paper's central takeaway. I recommend rewording the abstract and the corresponding passages in §6 to describe the types of experiences observed (e.g., 'users reported both successes and failures') or, alternatively, adding an explicit within-sample quantitative analysis with clear caveats about selection bias.
  2. [§5.1 (Keyword List Creation)] The keyword list contains 16 terms, most of which are restriction-oriented: 'moderate,' 'censor,' 'ban,' 'block,' 'suspend,' 'restrict,' 'warn,' 'flag,' 'appeal,' 'violate,' 'terminate,' 'remove,' 'content policy,' 'guardrail,' 'filter,' and 'refuse.' This selection strategy will over-capture negative moderation experiences and under-capture routine successes that users never mention. Consequently, the paper's comparative statements—e.g., that successes were 'pervasive' while failures were 'frequent'—may be artifacts of which posts users choose to share and which keywords the collection used. Please either temper such comparative claims or provide a baseline comparison from an unfiltered sample to assess the selection effect.
  3. [§6.1 and §6.2.1 (counts such as '5/6 tools')] Several findings use fractions such as '5/6 tools' to characterize the reach of a phenomenon (e.g., 'Failure in Mitigating False-Positive Rate (5/6 tools)'). As written, these counts can be read as prevalence estimates across the six GAI tools, but they are actually counts of tools for which the theme appeared somewhere in the 130 sampled posts and their comments. The sample is not representative of any tool's user base, and the keyword filter further biases detection of negative themes. Please clarify in the text that these fractions are descriptive within the coded sample, not population estimates, and adjust any associated frequency wording.
minor comments (3)
  1. [§3.2] The text states that the authors recorded '52 PDF files of 51 pages,' which is mildly confusing; please clarify whether one page spanned two PDF files or whether a duplicate was retained.
  2. [Abstract and §1] The abstract and introduction describe the study's scope as 'content moderation in GAI online tools' without consistently noting that Study 2 focuses specifically on AIGC creative tasks. Since the findings about 'frustration' and 'failures' in the abstract are drawn entirely from that narrower activity, the scope should be stated in the abstract to avoid overgeneralization.
  3. [§5.1 (Subreddits Choice)] The exclusion of r/NovelAI is justified by the authors' observation that NovelAI enforces almost no content moderation, but this exclusion, combined with the limited subreddit list, further narrows the set of tools studied; a sentence acknowledging the potential effect on the diversity of experiences would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the policy and Reddit analyses are newly coded empirical studies, and the shared Schaffner et al. self-citation is a methodological lens, not a load-bearing result.

full rationale

The paper's findings are not derived from their inputs by construction. Study 1 collects and qualitatively codes 52 policy pages from 14 GAI tools, and Study 2 keyword-searches Reddit, manually filters the results, and thematically codes a random sample of 130 posts with 3839 comments. The reuse of Schaffner et al.'s four-element framework (what, why, how, who) is a coding lens, and although three present authors are co-authors of that prior work, the framework is not the target claim and the coded data are newly collected. No fitted parameters, equations, or quantitative predictions are involved. The abstract's frequency wording ('pervasively', 'frequently') goes beyond what a qualitative sample can support, and the paper's own §5.3 acknowledges that 'our study may not have assessed all successes and failures around content moderation policy enforcement.' That is a validity or overclaim concern, not circularity, because the findings would still be meaningful even if prevalence cannot be estimated from the sample. There is also no imported uniqueness theorem or ansatz that forces the conclusions. Under the hard rules, no circular step can be quoted with a specific reduction, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claims rest on qualitative and sampling assumptions rather than numeric parameters. Free parameters and invented entities are absent because the study is observational and does not introduce new theoretical constructs or fitted constants.

assumptions (3)
  • domain assumption Qualitative coding with team consensus is sufficient to support thematic claims without an inter-rater reliability statistic.
    The authors explicitly decline to compute IRR in §3.3 and §5.2, citing iterative consensus practice. This is a methodological assumption about the validity of qualitative findings.
  • domain assumption Reddit posts and comments are authentic reports of user experience with content moderation.
    The study does not verify whether user-recounted moderation events actually occurred as described; it relies on public narratives as evidence in §5 and §6.
  • domain assumption Schaffner et al.'s four-element framework is an appropriate lens for evaluating the comprehensiveness of GAI moderation policies.
    Adopted directly from prior work [68] in §3.3. If the framework misses GAI-specific policy dimensions, the gap analysis may be incomplete.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products." pith.science (2026). https://pith.science/paper/RRJWAIDO

@misc{pith2026250614018,
  author       = {Pith},
  title        = {Pith review of: "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RRJWAIDO}},
  note         = {Machine review of arXiv:2506.14018}
}
read the original abstract

While recent research has focused on developing safeguards for generative AI (GAI) model-level content safety, little is known about how content moderation to prevent malicious content performs for end-users in real-world GAI products. To bridge this gap, we investigated content moderation policies and their enforcement in GAI online tools -- consumer-facing web-based GAI applications. We first analyzed content moderation policies of 14 GAI online tools. While these policies are comprehensive in outlining moderation practices, they usually lack details on practical implementations and are not specific about how users can aid in moderation or appeal moderation decisions. Next, we examined user-experienced content moderation successes and failures through Reddit discussions on GAI online tools. We found that although moderation systems succeeded in blocking malicious generations pervasively, users frequently experienced frustration in failures of both moderation systems and user support after moderation. Based on these findings, we suggest improvements for content moderation policy and user experiences in real-world GAI products.

Figures

Figures reproduced from arXiv: 2506.14018 by the authors.

Figure 1
Figure 1. Findings summary of Study 1: Policy Analysis and Study 2: Reddit Study. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

96 extracted references · 62 canonical work pages

  1. [1]

    Soyun Ahn, Jeeyun Baik, and Clara Sol Krause. Splinter- ing and centralizing platform governance: how facebook adapted its content moderation practices to the politi- cal and legal contexts in the united states, germany, and south korea. Information, Communication & Society, 26(14):2843–2862, 2023

  2. [2]

    Generative AI regulation can learn from social media regulation

    Ruth Elisabeth Appel. Generative ai regulation can learn from social media regulation. arXiv preprint arXiv:2412.11335, 2024

  3. [3]

    The place of inter-rater reliability in qualitative research: An empirical study

    David Armstrong, Ann Gosling, John Weinman, and Theresa Marteau. The place of inter-rater reliability in qualitative research: An empirical study. Sociology, 31(3):597–606, 1997

  4. [4]

    Detecting harmful content on online platforms: what platforms need vs

    Arnav Arora, Preslav Nakov, Momchil Hardalov, Sheikh Muhammad Sarwar, Vibha Nayak, Yoan Dinkov, Dimitrina Zlatkova, Kyle Dent, Ameya Bhatawdekar, Guillaume Bouchard, et al. Detecting harmful content on online platforms: what platforms need vs. where re- search efforts go. ACM Computing Surveys, 56(3):1–17, 2023

  5. [5]

    Training a helpful and harmless assistant with reinforce- ment learning from human feedback

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforce- ment learning from human feedback. arXiv preprint arXiv:2204.05862, 2022

  6. [6]

    Like trainer, like bot? inheritance of bias in algorithmic content moderation

    Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. Like trainer, like bot? inheritance of bias in algorithmic content moderation. In Social In- formatics: 9th International Conference, SocInfo 2017, Oxford, UK, September 13-15, 2017, Proceedings, Part II 9, pages 405–415. Springer, 2017

  7. [7]

    ’censorship- free’platforms: Evaluating content moderation policies and practices of alternative social media

    Nicole Buckley and Joseph S Schafer. ’censorship- free’platforms: Evaluating content moderation policies and practices of alternative social media. 2022

  8. [8]

    A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt

    Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip S Yu, and Lichao Sun. A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt. arXiv preprint arXiv:2303.04226, 2023

Show all 96 references
  1. [9]

    You can’t stay here: The efficacy of reddit’s 2015 ban examined through hate speech

    Eshwar Chandrasekharan, Umashanthi Pavalanathan, Anirudh Srinivasan, Adam Glynn, Jacob Eisenstein, and Eric Gilbert. You can’t stay here: The efficacy of reddit’s 2015 ban examined through hate speech. Proc. ACM Hum.-Comput. Interact., 1(CSCW), December 2017

  2. [10]

    A pathway to- wards responsible ai generated content

    Chen Chen, Jie Fu, and Lingjuan Lyu. A pathway to- wards responsible ai generated content. arXiv preprint arXiv:2303.01325, 2023

  3. [11]

    Chatgpt blocked 250,000 image generations of presidential candidates

    CNBC. Chatgpt blocked 250,000 image generations of presidential candidates. https://www.cnbc.com /2024/11/08/chatgpt-blocked-250000-image-g enerations-of-presidential-candidates.html ,

  4. [12]

    Safesora: Towards safety alignment of text2video generation via a human preference dataset

    Josef Dai, Tianle Chen, Xuyao Wang, Ziran Yang, Taiye Chen, Jiaming Ji, and Yaodong Yang. Safesora: Towards safety alignment of text2video generation via a human preference dataset. arXiv preprint arXiv:2406.14477, 2024

  5. [13]

    Safe rlhf: Safe reinforcement learning from human feedback

    Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023

  6. [14]

    Weaudit: Scaffolding user auditors and ai practitioners in auditing generative ai

    Wesley Hanwen Deng, Claire Wang, Howard Ziyu Han, Jason I Hong, Kenneth Holstein, and Motahhare Eslami. Weaudit: Scaffolding user auditors and ai practitioners in auditing generative ai. arXiv preprint arXiv:2501.01397, 2025

  7. [15]

    First i" like" it, then i hide it: Folk theories of social feeds

    Motahhare Eslami, Karrie Karahalios, Christian Sand- vig, Kristen Vaccaro, Aimee Rickman, Kevin Hamilton, and Alex Kirlik. First i" like" it, then i hide it: Folk theories of social feeds. In Proceedings of the 2016 cHI conference on human factors in computing systems, pages 2...

  8. [16]

    User- driven value alignment: Understanding users’ percep- tions and strategies for addressing biased and discrim- inatory statements in ai companions

    Xianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei, Maarten Sap, Zhicong Lu, and Hong Shen. User- driven value alignment: Understanding users’ percep- tions and strategies for addressing biased and discrim- inatory statements in ai companions. arXiv preprint arXiv:2409.00862, 2024

  9. [17]

    Feuston, and Amy S

    Casey Fiesler, Jessica L. Feuston, and Amy S. Bruck- man. Understanding copyright law in online creative communities. In Proceedings of the 18th ACM Confer- ence on Computer Supported Cooperative Work & So- cial Computing, CSCW ’15, page 116–129, New York, NY , USA, 2015. Asso...

  10. [18]

    Reddit rules! characterizing an ecosystem of governance

    Casey Fiesler, Jialun Jiang, Joshua McCann, Kyle Frye, and Jed Brubaker. Reddit rules! characterizing an ecosystem of governance. In Proceedings of the In- ternational AAAI Conference on Web and Social Media, volume 12, 2018

  11. [19]

    Bruckman

    Casey Fiesler, Cliff Lampe, and Amy S. Bruckman. Re- ality and perception of copyright terms of service for online content creation. In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing, CSCW ’16, page 1450–1461, New York, NY , US...

  12. [20]

    Remember the human: A systematic review of ethical considerations in reddit research

    Casey Fiesler, Michael Zimmer, Nicholas Proferes, Sarah Gilbert, and Naiyan Jones. Remember the human: A systematic review of ethical considerations in reddit research. Proceedings of the ACM on Human-Computer Interaction, 8(GROUP):1–33, 2024

  13. [21]

    [report] generative ai top 150: The world’s most used ai tools (feb 2024)

    FlexOS. [report] generative ai top 150: The world’s most used ai tools (feb 2024). https://www.flexos.wor k/learn/generative-ai-top-150 , 2025. Accessed: 2025-01-08

  14. [22]

    New research shows chatgpt reigns supreme in ai tool sector

    Forbes. New research shows chatgpt reigns supreme in ai tool sector. https://www.forbes.com/sites/c hriswestfall/2023/11/16/new-research-shows -chatgpt-reigns-supreme-in-ai-tool-sector/ ,

  15. [23]

    Generative artificial intelligence and education: An analysis from multi- ple perspectives

    Francisco José García-Peñalvo. Generative artificial intelligence and education: An analysis from multi- ple perspectives. Education in the Knowledge Society, 25:e31942, 2024

  16. [24]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language mod- els. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguis- tics: EMNLP 202...

  17. [25]

    Governance of and by platforms

    Tarleton Gillespie. Governance of and by platforms. SAGE handbook of social media, pages 254–278, 2017

  18. [26]

    Custodians of the Internet: Plat- forms, Content Moder ation, and the Hidden Decisions that Shape Social Media

    Tarleton Gillespie. Custodians of the Internet: Plat- forms, Content Moder ation, and the Hidden Decisions that Shape Social Media. Yale University Press, 2018

  19. [27]

    Generative ai and the politics of visi- bility

    Tarleton Gillespie. Generative ai and the politics of visi- bility. Big Data & Society, 11(2):20539517241252131, 2024

  20. [28]

    Expanding the debate about content moder- ation: Scholarly research agendas for the coming policy debates

    Tarleton Gillespie, Patricia Aufderheide, Elinor Carmi, Ysabel Gerrard, Robert Gorwa, Ariadna Matamoros- Fernández, Sarah T Roberts, Aram Sinnreich, and Sarah Myers West. Expanding the debate about content moder- ation: Scholarly research agendas for the coming policy debates....

  21. [29]

    Awesome generative ai

    Github. Awesome generative ai. https://github.com /steven2358/awesome-generative-ai, 2025. Ac- cessed: 2025-01-08

  22. [30]

    Content moderation remedies

    Eric Goldman. Content moderation remedies. Mich. Tech. L. Rev., 28:1, 2021

  23. [31]

    Algorithmic arbitrariness in content moderation

    Juan Felipe Gomez, Caio Machado, Lucas Monteiro Paes, and Flavio Calmon. Algorithmic arbitrariness in content moderation. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2234–2253, 2024

  24. [32]

    Reg- ulating chatgpt and other large generative ai models

    Philipp Hacker, Andreas Engel, and Marco Mauer. Reg- ulating chatgpt and other large generative ai models. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1112–1123, 2023

  25. [33]

    Disproportionate removals and differ- ing content moderation experiences for conservative, transgender, and black social media users: Marginaliza- tion and moderation gray areas

    Oliver L Haimson, Daniel Delmonaco, Peipei Nie, and Andrea Wegner. Disproportionate removals and differ- ing content moderation experiences for conservative, transgender, and black social media users: Marginaliza- tion and moderation gray areas. Proceedings of the ACM on Human...

  26. [34]

    Aligning artificial intelligence with human values: reflections from a phenomenological perspective

    Shengnan Han, Eugene Kelly, Shahrokh Nikou, and Eric- Oluf Svee. Aligning artificial intelligence with human values: reflections from a phenomenological perspective. AI & SOCIETY, pages 1–13, 2022

  27. [35]

    Generative ai in user- generated content

    Yiqing Hua, Shuo Niu, Jie Cai, Lydia B Chilton, Hendrik Heuer, and Donghee Yvette Wohn. Generative ai in user- generated content. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–7, 2024

  28. [36]

    Llama guard: Llm-based input-output safeguard for human-ai conversations

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023

  29. [37]

    did you suspect the post would be removed?

    Shagun Jhaver, Darren Scott Appling, Eric Gilbert, and Amy Bruckman. "did you suspect the post would be removed?": Understanding user reactions to content re- movals on reddit. Proc. ACM Hum.-Comput. Interact., 3(CSCW), November 2019

  30. [38]

    Does transparency in moderation really matter? user behavior after content removal explanations on reddit.Proc

    Shagun Jhaver, Amy Bruckman, and Eric Gilbert. Does transparency in moderation really matter? user behavior after content removal explanations on reddit.Proc. ACM Hum.-Comput. Interact., 3(CSCW), November 2019

  31. [39]

    Personalizing content moderation on social media: User perspectives on moderation choices, interface de- sign, and labor

    Shagun Jhaver, Alice Qian Zhang, Quan Ze Chen, Nikhila Natarajan, Ruotong Wang, and Amy X Zhang. Personalizing content moderation on social media: User perspectives on moderation choices, interface de- sign, and labor. Proceedings of the ACM on Human- Computer Interaction, 7(C...

  32. [40]

    Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset

    Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36, 2024

  33. [41]

    Ai alignment: A com- prehensive survey

    Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. Ai alignment: A com- prehensive survey. arXiv preprint arXiv:2310.19852, 2023

  34. [42]

    Brubaker, and Casey Fiesler

    Jialun ’Aaron’ Jiang, Skyler Middler, Jed R. Brubaker, and Casey Fiesler. Characterizing community guide- lines on social media platforms. In Companion Publi- cation of the 2020 Conference on Computer Supported Cooperative Work and Social Computing, CSCW ’20 Companion, page 28...

  35. [43]

    Jailbreaking large language models against modera- tion guardrails via cipher characters

    Haibo Jin, Andy Zhou, Joe D Menke, and Haohan Wang. Jailbreaking large language models against modera- tion guardrails via cipher characters. arXiv preprint arXiv:2405.20413, 2024

  36. [44]

    Through the looking glass: Study of transparency in reddit’s moderation practices

    Prerna Juneja, Deepika Rama Subramanian, and Tanushree Mitra. Through the looking glass: Study of transparency in reddit’s moderation practices. Pro- ceedings of the ACM on Human-Computer Interaction, 4(GROUP):1–35, 2020

  37. [45]

    Regulating behavior in online communities

    Sara Kiesler, Robert Kraut, Paul Resnick, and Aniket Kit- tur. Regulating behavior in online communities. Build- ing successful online communities: Evidence-based so- cial design, 1:4–2, 2012

  38. [46]

    The new governors: The people, rules, and processes governing online speech

    Kate Klonick. The new governors: The people, rules, and processes governing online speech. Harv. L. Rev., 131:1598, 2017

  39. [47]

    Acceptable use policies for foundation models

    Kevin Klyman. Acceptable use policies for foundation models. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume 7, pages 752–767, 2024

  40. [48]

    Community begins where moderation ends: Peer support and its implications for community- based rehabilitation

    Yubo Kou, Renkai Ma, Zinan Zhang, Yingfan Zhou, and Xinning Gui. Community begins where moderation ends: Peer support and its implications for community- based rehabilitation. In Proceedings of the CHI Confer- ence on Human Factors in Computing Systems, pages 1–18, 2024

  41. [49]

    Regulating online content moderation

    Kyle Langvardt. Regulating online content moderation. Geo. LJ, 106:1353, 2017

  42. [50]

    Safetydpo: Scalable safety align- ment for text-to-image generation

    Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, and Fabio Pizzati. Safetydpo: Scalable safety align- ment for text-to-image generation. arXiv preprint arXiv:2412.10493, 2024

  43. [51]

    What’s the appeal? perceptions of review processes for algorithmic decisions

    Henrietta Lyons, Senuri Wijenayake, Tim Miller, and Eduardo Velloso. What’s the appeal? perceptions of review processes for algorithmic decisions. In Proceed- ings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–15, 2022

  44. [52]

    how advertiser-friendly is my video?

    Renkai Ma and Yubo Kou. " how advertiser-friendly is my video?": Youtuber’s socioeconomic interactions with algorithmic content moderation. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1– 25, 2021

  45. [53]

    How do users experience moderation?: A systematic literature review

    Renkai Ma, Yue You, Xinning Gui, and Yubo Kou. How do users experience moderation?: A systematic literature review. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2):1–30, 2023

  46. [54]

    Auditing gpt’s content moderation guardrails: Can chatgpt write your favorite tv show? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 660–686, 2024

    Yaaseen Mahomed, Charlie M Crawford, Sanjana Gau- tam, Sorelle A Friedler, and Danaë Metaxa. Auditing gpt’s content moderation guardrails: Can chatgpt write your favorite tv show? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 660–686, 2024

  47. [55]

    A holistic approach to undesired content detection in the real world

    Todor Markov, Chong Zhang, Sandhini Agarwal, Flo- rentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. A holistic approach to undesired content detection in the real world. In Proceedings of the AAAI Conference on Artificial Intel- ligence, volum...

  48. [56]

    Nathan Matias, Austin Hounsel, and Nick Feamster

    J. Nathan Matias, Austin Hounsel, and Nick Feamster. Software-supported audits of decision-making systems: Testing google and facebook’s political advertising poli- cies. Proc. ACM Hum.-Comput. Interact., 6(CSCW1), April 2022

  49. [57]

    Reliability and inter-rater reliability in qualitative re- search: Norms and guidelines for cscw and hci practice

    Nora McDonald, Sarita Schoenebeck, and Andrea Forte. Reliability and inter-rater reliability in qualitative re- search: Norms and guidelines for cscw and hci practice. Proc. ACM Hum.-Comput. Interact., 3(CSCW), Novem- ber 2019

  50. [58]

    Folk theories of avoiding content moderation: How vaccine-opposed influencers amplify vaccine oppo- sition on instagram

    Rachel E Moran, Izzi Grasso, and Kolina Koltai. Folk theories of avoiding content moderation: How vaccine-opposed influencers amplify vaccine oppo- sition on instagram. Social Media+ Society , 8(4):20563051221144252, 2022

  51. [59]

    Censored, suspended, shadow- banned: User interpretations of content moderation on social media platforms

    Sarah Myers West. Censored, suspended, shadow- banned: User interpretations of content moderation on social media platforms. New Media & Society , 20(11):4366–4383, 2018

  52. [60]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information pro- cessing systems, 35:27...

  53. [61]

    Red-teaming the stable dif- fusion safety filter

    Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr. Red-teaming the stable dif- fusion safety filter. arXiv preprint arXiv:2210.04610, 2022

  54. [62]

    Ex- ploring the boundaries of content moderation in text- to-image generation

    Piera Riccio, Georgina Curto, and Nuria Oliver. Ex- ploring the boundaries of content moderation in text- to-image generation. arXiv preprint arXiv:2409.17155, 2024

  55. [63]

    Ex- posed or erased: Algorithmic censorship of nudity in art

    Piera Riccio, Thomas Hofmann, and Nuria Oliver. Ex- posed or erased: Algorithmic censorship of nudity in art. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–17, 2024

  56. [64]

    Behind the screen: Content Modera- tion in the Shadows of Social Media

    Sarah T Roberts. Behind the screen: Content Modera- tion in the Shadows of Social Media . Yale University Press, 2019

  57. [65]

    Generative ai meets copyright

    Pamela Samuelson. Generative ai meets copyright. Sci- ence, 381(6654):158–161, 2023

  58. [66]

    Ai industry analysis: 50 most visited ai tools and their 24b+ traffic behavior

    Sujan Sarkar. Ai industry analysis: 50 most visited ai tools and their 24b+ traffic behavior. https://writ erbuddy.ai/blog/ai-industry-analysis/, 2023. Accessed: 2025-01-08

  59. [67]

    Saturation in qualitative re- search: exploring its conceptualization and operational- ization

    Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. Saturation in qualitative re- search: exploring its conceptualization and operational- ization. Quality & quantity, 52:1893–1907, 2018

  60. [68]

    community guidelines make this the best party on the internet

    Brennan Schaffner, Arjun Nitin Bhagoji, Siyuan Cheng, Jacqueline Mei, Jay L Shen, Grace Wang, Marshini Chetty, Nick Feamster, Genevieve Lakier, and Chenhao Tan. " community guidelines make this the best party on the internet": An in-depth study of online platforms’ content mod...

  61. [69]

    Implications of regulations on large generative ai models in the super-election year and the impact on disinformation

    Vera Schmitt, Jakob Tesch, Eva Lopez, Tim Polzehl, Aljoscha Burchardt, Konstanze Neumann, Salar Mohtaj, and Sebastian Möller. Implications of regulations on large generative ai models in the super-election year and the impact on disinformation. In Proceedings of the Workshop o...

  62. [70]

    Safe latent diffusion: Mitigat- ing inappropriate degeneration in diffusion models

    Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting. Safe latent diffusion: Mitigat- ing inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 22522–22531, 2023

  63. [71]

    Proximal policy optimiza- tion algorithms

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347, 2017

  64. [72]

    Reconsidering self-moderation: the role of research in supporting community-based models for online content moderation

    Joseph Seering. Reconsidering self-moderation: the role of research in supporting community-based models for online content moderation. Proceedings of the ACM on Human-Computer Interaction, 4(CSCW2):1–28, 2020

  65. [73]

    Why so toxic? measuring and triggering toxic behavior in open-domain chatbots

    Wai Man Si, Michael Backes, Jeremy Blackburn, Emil- iano De Cristofaro, Gianluca Stringhini, Savvas Zannet- tou, and Yang Zhang. Why so toxic? measuring and triggering toxic behavior in open-domain chatbots. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Comm...

  66. [74]

    Sok: Content moderation in social media, from guidelines to enforcement, and research to practice

    Mohit Singhal, Chen Ling, Pujan Paudel, Poojitha Thota, Nihal Kumarswamy, Gianluca Stringhini, and Shirin Nilizadeh. Sok: Content moderation in social media, from guidelines to enforcement, and research to practice. In 2023 IEEE 8th European Symposium on Security and Privacy (...

  67. [75]

    Intention of generative artificial in- telligence (gai) usage by adults in the united states as of august 2023, by type

    Statista. Intention of generative artificial in- telligence (gai) usage by adults in the united states as of august 2023, by type. https: //www.statista.com/statistics/1461998/us a-generative-ai-usage-intention-by-type/ ,

  68. [76]

    Use of generative artificial intelligence (ai) programs in the united states in 2023, by use case

    Statista. Use of generative artificial intelligence (ai) programs in the united states in 2023, by use case. https://www.statista.com/statistics/ 1413836/use-of-generative-ai-us/ , 2024. Ac- cessed: 2025-01-08

  69. [77]

    Lawless: The secret rules that govern our digital lives

    Nicolas P Suzor. Lawless: The secret rules that govern our digital lives. Cambridge University Press, 2019

  70. [79]

    Tiktok adds more generative ai features

    Social Media Today. Tiktok adds more generative ai features. https://www.socialmediatoday.c om/news/tiktok-adds-gen-ai-image-tools-c aption-suggestions/736361/, 2025. Accessed: 2025-01-19

  71. [80]

    The digital services act and the eu as the global regulator of the internet

    Ioanna Tourkochoriti. The digital services act and the eu as the global regulator of the internet. Chi. J. Int’l L., 24:129, 2023

  72. [81]

    Investigating moderation challenges to com- bating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on reddit

    Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. Investigating moderation challenges to com- bating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on reddit. In 33rd USENIX Security Symposium (USENIX Security 24) , pages 37–54, Phila...

  73. [82]

    House of Representatives

    U.S. House of Representatives. United states code: Title 15, section 9401 (preliminary edition). https://uscode.house.gov/view.xhtml?req= (title:15%20section:9401%20edition:prelim),

  74. [83]

    at the end of the day facebook does what itwants

    Kristen Vaccaro, Christian Sandvig, and Karrie Kara- halios. " at the end of the day facebook does what itwants" how users experience contesting algorithmic content moderation. Proceedings of the ACM on human- computer interaction, 4(CSCW2):1–22, 2020

  75. [84]

    Llama 2: Open foundation and fine-tuned chat models

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023

  76. [85]

    Moderator: Moderating text-to- image diffusion models through fine-grained context- based policies

    Peiran Wang, Qiyu Li, Longxuan Yu, Ziyao Wang, Ang Li, and Haojian Jin. Moderator: Moderating text-to- image diffusion models through fine-grained context- based policies. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 1181–1...

  77. [86]

    Finetuned language models are zero-shot learners

    Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021

  78. [87]

    Understanding the im- pact of ai-generated content on social media: The pixiv case

    Yiluo Wei and Gareth Tyson. Understanding the im- pact of ai-generated content on social media: The pixiv case. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 6813–6822, 2024

  79. [88]

    Contestability for content moderation

    Kristen Vaccaro, Ziang Xiao, Kevin Hamilton, and Kar- rie Karahalios. Contestability for content moderation. Proc. ACM Hum.-Comput. Interact., 5(CSCW2), Octo- ber 2021

  80. [89]

    Ai-generated content (aigc): A survey

    Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and Hong Lin. Ai-generated content (aigc): A survey. arXiv preprint arXiv:2304.06632, 2023

  81. [90]

    Fine-grained hu- man feedback gives better rewards for language model training

    Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. Fine-grained hu- man feedback gives better rewards for language model training. Advances in Neural Information Processing Systems, 36:59008–5...

  82. [91]

    Don’t listen to me: Understanding and exploring jailbreak prompts of large language models

    Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang. Don’t listen to me: Understanding and exploring jailbreak prompts of large language models. In 33rd USENIX Security Symposium (USENIX Security 24), pages 4675–4692, Philadelphia, PA, August 2...

  83. [92]

    as an ai language model, i cannot

    Joel Wester, Tim Schrills, Henning Pohl, and Niels van Berkel. “as an ai language model, i cannot”: Investi- gating llm denials of user requests. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–14, 2024

  84. [93]

    Con- trollable safety alignment: Inference-time adaptation to diverse safety requirements

    Jingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, and Benjamin Van Durme. Con- trollable safety alignment: Inference-time adaptation to diverse safety requirements. arXiv preprint arXiv:2410.08968, 2024

  85. [94]

    Siren’s song in the ai ocean: a survey on hallucination in large language models

    Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023

  86. [95]

    the 2 characters’ foreheads are touching in a display of tenderness and affection

    Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. Synthetic lies: Un- derstanding ai-generated misinformation and evaluating algorithmic and human solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1...

  87. [96]

    i won the election!

    Savvas Zannettou. " i won the election!": an empirical analysis of soft moderation interventions on twitter. In Proceedings of the international AAAI conference on web and social media, volume 15, pages 865–876, 2021

  88. [2025]

    Accessed: 2025-01-08

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.