REVIEW 3 major objections 3 minor 96 references
"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products
T0 review · 3 major / 3 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Generative AI content filters block abuse but leave users without an effective appeal.
desk verdict A solid first map of GAI moderation policies and user experiences; the policy analysis is careful, but the abstract's frequency claims outrun the self-selected Reddit sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the content moderation pipeline of a GAI online tool, analyzed through a three-part policy framework: moderation criteria (what content is forbidden), moderation methodology (how problematic content is detected), and moderation consequences (what happens to content and users). The paper's analytic move is to map user-experienced successes and failures onto that framework, using Reddit discussions of AIGC creative tasks to see where enforcement diverges from lived experience. The comparative lens is the analogy to online communities: GAI products copied the policy structure but not the user-side apparatus of reporting and appeals.
What would settle it
A longitudinal telemetry study inside one or more GAI products—logging every moderation decision, its true or false positive status, and the resolution of each appeal—would settle the central claim; if false-positive moderation and unresolved appeals are rare in system logs, the paper's frequency claims about user frustration would be falsified.
Extended reading notes
Core claim
The paper's central claim is that content moderation in generative AI products is a two-sided story: it blocks clearly malicious content at scale, but fails users precisely at the moments that determine trust—explaining decisions, allowing appeals, and judging context. Based on a qualitative policy analysis of 14 GAI online tools and a thematic analysis of 130 randomly sampled Reddit posts about creative generation tasks, the authors find that policies comprehensively outline what is forbidden, how violations are detected, and what consequences follow, yet omit concrete details on user reporting and appeals. User discussions show the same split: widespread appreciation that harmful requests are blocked, alongside frequent false positives, inconsistent decisions, context-blind censorship of ordinary fiction and art, opaque explanations that are sometimes generated by the model itself, and appeal processes that are slow or silent. The paper treats these failures as a structural gap in the moderation pipeline, not as isolated bugs.
Load-bearing premise
The load-bearing premise is that self-selected Reddit discussions give a representative window onto how often moderation succeeds and fails, so that qualitative examples can support frequency claims like 'pervasively' and 'frequently'; if Reddit overrepresents frustrated users, the success rate may be overstated.
Editorial extensions
If this is right
- GAI products should consolidate scattered moderation rules into a single dedicated policy page, separating tool-specific rules from rules that govern sharing in associated online communities.
- Products should establish clear, step-by-step reporting and appeal channels for all moderation decisions, not only copyright-related takedowns.
- Moderation systems should prefer soft moderation, such as warnings, content masking, or modified output, over outright denial when there is a chance of a false positive.
- Policies should disclose implementation details such as banned-word lists and explain moderation decisions at the specific stage where they occur, so users can judge whether a decision was justified.
Reading between the lines
- Because the Reddit sample is self-selected and frustration-heavy, the paper's 'pervasive success' of moderation is likely conservative for the broader population; a representative user survey would probably show even higher block rates and lower complaint rates than the forums suggest.
- The findings imply a testable design claim: providing stage-specific, concrete explanations of moderation decisions should measurably reduce perceived unfairness and improve retention, which an A/B test could verify.
- The paper's focus on creative generation tasks leaves open how moderation failures differ in dialogue, search, and coding tasks, where false positives may take different forms.
- If regulators require transparency and redress for automated content decisions, the gaps documented here, such as no banned-word list, no general appeal, and opaque grounds, could become legal liabilities for GAI providers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates content moderation in consumer-facing generative AI (GAI) online tools through two studies. Study 1 analyzes the content moderation policies of 14 US-based GAI tools, finding that policies are comprehensive in covering moderation criteria, methodology, and consequences, but lack detail on user-driven moderation and appeals. Study 2 analyzes 130 randomly sampled Reddit posts (with their comments) from keyword-filtered discussions about moderation experiences in AIGC creative tasks, identifying both successes in blocking malicious generations and failures in moderation decision-making and post-moderation user support. The authors propose policy and product improvements such as unified policy structures, soft moderation, personalized guardrails, and more transparent moderation pipelines.
Significance. If its claims are suitably qualified, this is a useful contribution to HCI and AI-safety research. The paper provides one of the first empirical mappings of GAI product content moderation policies and user experiences, complementing model-level safety auditing work. It makes two datasets publicly available (the policy corpus and the filtered Reddit post dataset) and follows a transparent coding process, with 44 of 51 policy pages double-coded and all 130 Reddit posts double-coded. The qualitative findings—particularly the gap between policy detail and user-perceived enforcement, and the lack of meaningful appeal mechanisms—are valuable and actionable for designers and policymakers. The main weakness is that the abstract and several findings sections use quantitative prevalence language that the sample design cannot support; this is correctable within the manuscript's scope and does not undermine the qualitative core.
major comments (3)
- [Abstract; §5.3] The abstract's claim that moderation systems succeeded 'pervasively' and that users 'frequently experienced frustration' is not supported by the study design. The underlying data are 130 posts from a keyword-filtered, self-selected corpus of Reddit discussions, which has no unbiased denominator and systematically over-captures negative experiences. The paper itself acknowledges in §5.3 that the study 'may not have assessed all successes and failures around content moderation policy enforcement.' These frequency adverbs are load-bearing because they define the paper's central takeaway. I recommend rewording the abstract and the corresponding passages in §6 to describe the types of experiences observed (e.g., 'users reported both successes and failures') or, alternatively, adding an explicit within-sample quantitative analysis with clear caveats about selection bias.
- [§5.1 (Keyword List Creation)] The keyword list contains 16 terms, most of which are restriction-oriented: 'moderate,' 'censor,' 'ban,' 'block,' 'suspend,' 'restrict,' 'warn,' 'flag,' 'appeal,' 'violate,' 'terminate,' 'remove,' 'content policy,' 'guardrail,' 'filter,' and 'refuse.' This selection strategy will over-capture negative moderation experiences and under-capture routine successes that users never mention. Consequently, the paper's comparative statements—e.g., that successes were 'pervasive' while failures were 'frequent'—may be artifacts of which posts users choose to share and which keywords the collection used. Please either temper such comparative claims or provide a baseline comparison from an unfiltered sample to assess the selection effect.
- [§6.1 and §6.2.1 (counts such as '5/6 tools')] Several findings use fractions such as '5/6 tools' to characterize the reach of a phenomenon (e.g., 'Failure in Mitigating False-Positive Rate (5/6 tools)'). As written, these counts can be read as prevalence estimates across the six GAI tools, but they are actually counts of tools for which the theme appeared somewhere in the 130 sampled posts and their comments. The sample is not representative of any tool's user base, and the keyword filter further biases detection of negative themes. Please clarify in the text that these fractions are descriptive within the coded sample, not population estimates, and adjust any associated frequency wording.
minor comments (3)
- [§3.2] The text states that the authors recorded '52 PDF files of 51 pages,' which is mildly confusing; please clarify whether one page spanned two PDF files or whether a duplicate was retained.
- [Abstract and §1] The abstract and introduction describe the study's scope as 'content moderation in GAI online tools' without consistently noting that Study 2 focuses specifically on AIGC creative tasks. Since the findings about 'frustration' and 'failures' in the abstract are drawn entirely from that narrower activity, the scope should be stated in the abstract to avoid overgeneralization.
- [§5.1 (Subreddits Choice)] The exclusion of r/NovelAI is justified by the authors' observation that NovelAI enforces almost no content moderation, but this exclusion, combined with the limited subreddit list, further narrows the set of tools studied; a sentence acknowledging the potential effect on the diversity of experiences would be helpful.
Circularity Check
No significant circularity: the policy and Reddit analyses are newly coded empirical studies, and the shared Schaffner et al. self-citation is a methodological lens, not a load-bearing result.
full rationale
The paper's findings are not derived from their inputs by construction. Study 1 collects and qualitatively codes 52 policy pages from 14 GAI tools, and Study 2 keyword-searches Reddit, manually filters the results, and thematically codes a random sample of 130 posts with 3839 comments. The reuse of Schaffner et al.'s four-element framework (what, why, how, who) is a coding lens, and although three present authors are co-authors of that prior work, the framework is not the target claim and the coded data are newly collected. No fitted parameters, equations, or quantitative predictions are involved. The abstract's frequency wording ('pervasively', 'frequently') goes beyond what a qualitative sample can support, and the paper's own §5.3 acknowledges that 'our study may not have assessed all successes and failures around content moderation policy enforcement.' That is a validity or overclaim concern, not circularity, because the findings would still be meaningful even if prevalence cannot be estimated from the sample. There is also no imported uniqueness theorem or ansatz that forces the conclusions. Under the hard rules, no circular step can be quoted with a specific reduction, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Qualitative coding with team consensus is sufficient to support thematic claims without an inter-rater reliability statistic.
- domain assumption Reddit posts and comments are authentic reports of user experience with content moderation.
- domain assumption Schaffner et al.'s four-element framework is an appropriate lens for evaluating the comprehensiveness of GAI moderation policies.
Cite this review
Pith. "Pith review of "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products." pith.science (2026). https://pith.science/paper/RRJWAIDO
@misc{pith2026250614018,
author = {Pith},
title = {Pith review of: "I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products},
year = {2026},
howpublished = {\url{https://pith.science/paper/RRJWAIDO}},
note = {Machine review of arXiv:2506.14018}
}
read the original abstract
While recent research has focused on developing safeguards for generative AI (GAI) model-level content safety, little is known about how content moderation to prevent malicious content performs for end-users in real-world GAI products. To bridge this gap, we investigated content moderation policies and their enforcement in GAI online tools -- consumer-facing web-based GAI applications. We first analyzed content moderation policies of 14 GAI online tools. While these policies are comprehensive in outlining moderation practices, they usually lack details on practical implementations and are not specific about how users can aid in moderation or appeal moderation decisions. Next, we examined user-experienced content moderation successes and failures through Reddit discussions on GAI online tools. We found that although moderation systems succeeded in blocking malicious generations pervasively, users frequently experienced frustration in failures of both moderation systems and user support after moderation. Based on these findings, we suggest improvements for content moderation policy and user experiences in real-world GAI products.
Figures
Reference graph
Works this paper leans on
-
[1]
Soyun Ahn, Jeeyun Baik, and Clara Sol Krause. Splinter- ing and centralizing platform governance: how facebook adapted its content moderation practices to the politi- cal and legal contexts in the united states, germany, and south korea. Information, Communication & Society, 26(14):2843–2862, 2023
2023
-
[2]
Generative AI regulation can learn from social media regulation
Ruth Elisabeth Appel. Generative ai regulation can learn from social media regulation. arXiv preprint arXiv:2412.11335, 2024
work page Pith review arXiv 2024
-
[3]
The place of inter-rater reliability in qualitative research: An empirical study
David Armstrong, Ann Gosling, John Weinman, and Theresa Marteau. The place of inter-rater reliability in qualitative research: An empirical study. Sociology, 31(3):597–606, 1997
1997
-
[4]
Detecting harmful content on online platforms: what platforms need vs
Arnav Arora, Preslav Nakov, Momchil Hardalov, Sheikh Muhammad Sarwar, Vibha Nayak, Yoan Dinkov, Dimitrina Zlatkova, Kyle Dent, Ameya Bhatawdekar, Guillaume Bouchard, et al. Detecting harmful content on online platforms: what platforms need vs. where re- search efforts go. ACM Computing Surveys, 56(3):1–17, 2023
2023
-
[5]
Training a helpful and harmless assistant with reinforce- ment learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforce- ment learning from human feedback. arXiv preprint arXiv:2204.05862, 2022
arXiv 2022
-
[6]
Like trainer, like bot? inheritance of bias in algorithmic content moderation
Reuben Binns, Michael Veale, Max Van Kleek, and Nigel Shadbolt. Like trainer, like bot? inheritance of bias in algorithmic content moderation. In Social In- formatics: 9th International Conference, SocInfo 2017, Oxford, UK, September 13-15, 2017, Proceedings, Part II 9, pages 405–415. Springer, 2017
2017
-
[7]
’censorship- free’platforms: Evaluating content moderation policies and practices of alternative social media
Nicole Buckley and Joseph S Schafer. ’censorship- free’platforms: Evaluating content moderation policies and practices of alternative social media. 2022
2022
-
[8]
Yihan Cao, Siyu Li, Yixin Liu, Zhiling Yan, Yutong Dai, Philip S Yu, and Lichao Sun. A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt. arXiv preprint arXiv:2303.04226, 2023
arXiv 2023
Show all 96 references
-
[9]
You can’t stay here: The efficacy of reddit’s 2015 ban examined through hate speech
Eshwar Chandrasekharan, Umashanthi Pavalanathan, Anirudh Srinivasan, Adam Glynn, Jacob Eisenstein, and Eric Gilbert. You can’t stay here: The efficacy of reddit’s 2015 ban examined through hate speech. Proc. ACM Hum.-Comput. Interact., 1(CSCW), December 2017
2015
-
[10]
A pathway to- wards responsible ai generated content
Chen Chen, Jie Fu, and Lingjuan Lyu. A pathway to- wards responsible ai generated content. arXiv preprint arXiv:2303.01325, 2023
2023 arXiv
-
[11]
Chatgpt blocked 250,000 image generations of presidential candidates
CNBC. Chatgpt blocked 250,000 image generations of presidential candidates. https://www.cnbc.com /2024/11/08/chatgpt-blocked-250000-image-g enerations-of-presidential-candidates.html ,
2024
-
[12]
Safesora: Towards safety alignment of text2video generation via a human preference dataset
Josef Dai, Tianle Chen, Xuyao Wang, Ziran Yang, Taiye Chen, Jiaming Ji, and Yaodong Yang. Safesora: Towards safety alignment of text2video generation via a human preference dataset. arXiv preprint arXiv:2406.14477, 2024
2024 arXiv
-
[13]
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. Safe rlhf: Safe reinforcement learning from human feedback. arXiv preprint arXiv:2310.12773, 2023
2023 arXiv
-
[14]
Weaudit: Scaffolding user auditors and ai practitioners in auditing generative ai
Wesley Hanwen Deng, Claire Wang, Howard Ziyu Han, Jason I Hong, Kenneth Holstein, and Motahhare Eslami. Weaudit: Scaffolding user auditors and ai practitioners in auditing generative ai. arXiv preprint arXiv:2501.01397, 2025
2025 arXiv
-
[15]
First i" like" it, then i hide it: Folk theories of social feeds
Motahhare Eslami, Karrie Karahalios, Christian Sand- vig, Kristen Vaccaro, Aimee Rickman, Kevin Hamilton, and Alex Kirlik. First i" like" it, then i hide it: Folk theories of social feeds. In Proceedings of the 2016 cHI conference on human factors in computing systems, pages 2...
2016
-
[16]
User- driven value alignment: Understanding users’ percep- tions and strategies for addressing biased and discrim- inatory statements in ai companions
Xianzhe Fan, Qing Xiao, Xuhui Zhou, Jiaxin Pei, Maarten Sap, Zhicong Lu, and Hong Shen. User- driven value alignment: Understanding users’ percep- tions and strategies for addressing biased and discrim- inatory statements in ai companions. arXiv preprint arXiv:2409.00862, 2024
2024 arXiv
-
[17]
Feuston, and Amy S
Casey Fiesler, Jessica L. Feuston, and Amy S. Bruck- man. Understanding copyright law in online creative communities. In Proceedings of the 18th ACM Confer- ence on Computer Supported Cooperative Work & So- cial Computing, CSCW ’15, page 116–129, New York, NY , USA, 2015. Asso...
2015
-
[18]
Reddit rules! characterizing an ecosystem of governance
Casey Fiesler, Jialun Jiang, Joshua McCann, Kyle Frye, and Jed Brubaker. Reddit rules! characterizing an ecosystem of governance. In Proceedings of the In- ternational AAAI Conference on Web and Social Media, volume 12, 2018
2018
-
[19]
Bruckman
Casey Fiesler, Cliff Lampe, and Amy S. Bruckman. Re- ality and perception of copyright terms of service for online content creation. In Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work & Social Computing, CSCW ’16, page 1450–1461, New York, NY , US...
2016
-
[20]
Remember the human: A systematic review of ethical considerations in reddit research
Casey Fiesler, Michael Zimmer, Nicholas Proferes, Sarah Gilbert, and Naiyan Jones. Remember the human: A systematic review of ethical considerations in reddit research. Proceedings of the ACM on Human-Computer Interaction, 8(GROUP):1–33, 2024
2024
-
[21]
[report] generative ai top 150: The world’s most used ai tools (feb 2024)
FlexOS. [report] generative ai top 150: The world’s most used ai tools (feb 2024). https://www.flexos.wor k/learn/generative-ai-top-150 , 2025. Accessed: 2025-01-08
2024
-
[22]
New research shows chatgpt reigns supreme in ai tool sector
Forbes. New research shows chatgpt reigns supreme in ai tool sector. https://www.forbes.com/sites/c hriswestfall/2023/11/16/new-research-shows -chatgpt-reigns-supreme-in-ai-tool-sector/ ,
2023
-
[23]
Generative artificial intelligence and education: An analysis from multi- ple perspectives
Francisco José García-Peñalvo. Generative artificial intelligence and education: An analysis from multi- ple perspectives. Education in the Knowledge Society, 25:e31942, 2024
2024
-
[24]
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. RealToxicityPrompts: Evaluating neural toxic degeneration in language mod- els. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Association for Computational Linguis- tics: EMNLP 202...
2020
-
[25]
Governance of and by platforms
Tarleton Gillespie. Governance of and by platforms. SAGE handbook of social media, pages 254–278, 2017
2017
-
[26]
Custodians of the Internet: Plat- forms, Content Moder ation, and the Hidden Decisions that Shape Social Media
Tarleton Gillespie. Custodians of the Internet: Plat- forms, Content Moder ation, and the Hidden Decisions that Shape Social Media. Yale University Press, 2018
2018
-
[27]
Generative ai and the politics of visi- bility
Tarleton Gillespie. Generative ai and the politics of visi- bility. Big Data & Society, 11(2):20539517241252131, 2024
2024
-
[28]
Expanding the debate about content moder- ation: Scholarly research agendas for the coming policy debates
Tarleton Gillespie, Patricia Aufderheide, Elinor Carmi, Ysabel Gerrard, Robert Gorwa, Ariadna Matamoros- Fernández, Sarah T Roberts, Aram Sinnreich, and Sarah Myers West. Expanding the debate about content moder- ation: Scholarly research agendas for the coming policy debates....
2020
-
[29]
Awesome generative ai
Github. Awesome generative ai. https://github.com /steven2358/awesome-generative-ai, 2025. Ac- cessed: 2025-01-08
2025
-
[30]
Content moderation remedies
Eric Goldman. Content moderation remedies. Mich. Tech. L. Rev., 28:1, 2021
2021
-
[31]
Algorithmic arbitrariness in content moderation
Juan Felipe Gomez, Caio Machado, Lucas Monteiro Paes, and Flavio Calmon. Algorithmic arbitrariness in content moderation. In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 2234–2253, 2024
2024
-
[32]
Reg- ulating chatgpt and other large generative ai models
Philipp Hacker, Andreas Engel, and Marco Mauer. Reg- ulating chatgpt and other large generative ai models. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, pages 1112–1123, 2023
2023
-
[33]
Disproportionate removals and differ- ing content moderation experiences for conservative, transgender, and black social media users: Marginaliza- tion and moderation gray areas
Oliver L Haimson, Daniel Delmonaco, Peipei Nie, and Andrea Wegner. Disproportionate removals and differ- ing content moderation experiences for conservative, transgender, and black social media users: Marginaliza- tion and moderation gray areas. Proceedings of the ACM on Human...
2021
-
[34]
Aligning artificial intelligence with human values: reflections from a phenomenological perspective
Shengnan Han, Eugene Kelly, Shahrokh Nikou, and Eric- Oluf Svee. Aligning artificial intelligence with human values: reflections from a phenomenological perspective. AI & SOCIETY, pages 1–13, 2022
2022
-
[35]
Generative ai in user- generated content
Yiqing Hua, Shuo Niu, Jie Cai, Lydia B Chilton, Hendrik Heuer, and Donghee Yvette Wohn. Generative ai in user- generated content. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–7, 2024
2024
-
[36]
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023
2023 arXiv
-
[37]
did you suspect the post would be removed?
Shagun Jhaver, Darren Scott Appling, Eric Gilbert, and Amy Bruckman. "did you suspect the post would be removed?": Understanding user reactions to content re- movals on reddit. Proc. ACM Hum.-Comput. Interact., 3(CSCW), November 2019
2019
-
[38]
Does transparency in moderation really matter? user behavior after content removal explanations on reddit.Proc
Shagun Jhaver, Amy Bruckman, and Eric Gilbert. Does transparency in moderation really matter? user behavior after content removal explanations on reddit.Proc. ACM Hum.-Comput. Interact., 3(CSCW), November 2019
2019
-
[39]
Personalizing content moderation on social media: User perspectives on moderation choices, interface de- sign, and labor
Shagun Jhaver, Alice Qian Zhang, Quan Ze Chen, Nikhila Natarajan, Ruotong Wang, and Amy X Zhang. Personalizing content moderation on social media: User perspectives on moderation choices, interface de- sign, and labor. Proceedings of the ACM on Human- Computer Interaction, 7(C...
2023
-
[40]
Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards im- proved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[41]
Ai alignment: A com- prehensive survey
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Jiayi Zhou, Zhaowei Zhang, et al. Ai alignment: A com- prehensive survey. arXiv preprint arXiv:2310.19852, 2023
2023 arXiv
-
[42]
Brubaker, and Casey Fiesler
Jialun ’Aaron’ Jiang, Skyler Middler, Jed R. Brubaker, and Casey Fiesler. Characterizing community guide- lines on social media platforms. In Companion Publi- cation of the 2020 Conference on Computer Supported Cooperative Work and Social Computing, CSCW ’20 Companion, page 28...
2020
-
[43]
Jailbreaking large language models against modera- tion guardrails via cipher characters
Haibo Jin, Andy Zhou, Joe D Menke, and Haohan Wang. Jailbreaking large language models against modera- tion guardrails via cipher characters. arXiv preprint arXiv:2405.20413, 2024
2024 arXiv
-
[44]
Through the looking glass: Study of transparency in reddit’s moderation practices
Prerna Juneja, Deepika Rama Subramanian, and Tanushree Mitra. Through the looking glass: Study of transparency in reddit’s moderation practices. Pro- ceedings of the ACM on Human-Computer Interaction, 4(GROUP):1–35, 2020
2020
-
[45]
Regulating behavior in online communities
Sara Kiesler, Robert Kraut, Paul Resnick, and Aniket Kit- tur. Regulating behavior in online communities. Build- ing successful online communities: Evidence-based so- cial design, 1:4–2, 2012
2012
-
[46]
The new governors: The people, rules, and processes governing online speech
Kate Klonick. The new governors: The people, rules, and processes governing online speech. Harv. L. Rev., 131:1598, 2017
2017
-
[47]
Acceptable use policies for foundation models
Kevin Klyman. Acceptable use policies for foundation models. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume 7, pages 752–767, 2024
2024
-
[48]
Community begins where moderation ends: Peer support and its implications for community- based rehabilitation
Yubo Kou, Renkai Ma, Zinan Zhang, Yingfan Zhou, and Xinning Gui. Community begins where moderation ends: Peer support and its implications for community- based rehabilitation. In Proceedings of the CHI Confer- ence on Human Factors in Computing Systems, pages 1–18, 2024
2024
-
[49]
Regulating online content moderation
Kyle Langvardt. Regulating online content moderation. Geo. LJ, 106:1353, 2017
2017
-
[50]
Safetydpo: Scalable safety align- ment for text-to-image generation
Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, and Fabio Pizzati. Safetydpo: Scalable safety align- ment for text-to-image generation. arXiv preprint arXiv:2412.10493, 2024
2024 arXiv
-
[51]
What’s the appeal? perceptions of review processes for algorithmic decisions
Henrietta Lyons, Senuri Wijenayake, Tim Miller, and Eduardo Velloso. What’s the appeal? perceptions of review processes for algorithmic decisions. In Proceed- ings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–15, 2022
2022
-
[52]
how advertiser-friendly is my video?
Renkai Ma and Yubo Kou. " how advertiser-friendly is my video?": Youtuber’s socioeconomic interactions with algorithmic content moderation. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW2):1– 25, 2021
2021
-
[53]
How do users experience moderation?: A systematic literature review
Renkai Ma, Yue You, Xinning Gui, and Yubo Kou. How do users experience moderation?: A systematic literature review. Proceedings of the ACM on Human-Computer Interaction, 7(CSCW2):1–30, 2023
2023
-
[54]
Auditing gpt’s content moderation guardrails: Can chatgpt write your favorite tv show? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 660–686, 2024
Yaaseen Mahomed, Charlie M Crawford, Sanjana Gau- tam, Sorelle A Friedler, and Danaë Metaxa. Auditing gpt’s content moderation guardrails: Can chatgpt write your favorite tv show? In The 2024 ACM Conference on Fairness, Accountability, and Transparency, pages 660–686, 2024
2024
-
[55]
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Flo- rentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. A holistic approach to undesired content detection in the real world. In Proceedings of the AAAI Conference on Artificial Intel- ligence, volum...
2023
-
[56]
Nathan Matias, Austin Hounsel, and Nick Feamster
J. Nathan Matias, Austin Hounsel, and Nick Feamster. Software-supported audits of decision-making systems: Testing google and facebook’s political advertising poli- cies. Proc. ACM Hum.-Comput. Interact., 6(CSCW1), April 2022
2022
-
[57]
Reliability and inter-rater reliability in qualitative re- search: Norms and guidelines for cscw and hci practice
Nora McDonald, Sarita Schoenebeck, and Andrea Forte. Reliability and inter-rater reliability in qualitative re- search: Norms and guidelines for cscw and hci practice. Proc. ACM Hum.-Comput. Interact., 3(CSCW), Novem- ber 2019
2019
-
[58]
Folk theories of avoiding content moderation: How vaccine-opposed influencers amplify vaccine oppo- sition on instagram
Rachel E Moran, Izzi Grasso, and Kolina Koltai. Folk theories of avoiding content moderation: How vaccine-opposed influencers amplify vaccine oppo- sition on instagram. Social Media+ Society , 8(4):20563051221144252, 2022
2022
-
[59]
Censored, suspended, shadow- banned: User interpretations of content moderation on social media platforms
Sarah Myers West. Censored, suspended, shadow- banned: User interpretations of content moderation on social media platforms. New Media & Society , 20(11):4366–4383, 2018
2018
-
[60]
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information pro- cessing systems, 35:27...
2022
-
[61]
Red-teaming the stable dif- fusion safety filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr. Red-teaming the stable dif- fusion safety filter. arXiv preprint arXiv:2210.04610, 2022
2022 arXiv
-
[62]
Ex- ploring the boundaries of content moderation in text- to-image generation
Piera Riccio, Georgina Curto, and Nuria Oliver. Ex- ploring the boundaries of content moderation in text- to-image generation. arXiv preprint arXiv:2409.17155, 2024
2024 arXiv
-
[63]
Ex- posed or erased: Algorithmic censorship of nudity in art
Piera Riccio, Thomas Hofmann, and Nuria Oliver. Ex- posed or erased: Algorithmic censorship of nudity in art. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–17, 2024
2024
-
[64]
Behind the screen: Content Modera- tion in the Shadows of Social Media
Sarah T Roberts. Behind the screen: Content Modera- tion in the Shadows of Social Media . Yale University Press, 2019
2019
-
[65]
Generative ai meets copyright
Pamela Samuelson. Generative ai meets copyright. Sci- ence, 381(6654):158–161, 2023
2023
-
[66]
Ai industry analysis: 50 most visited ai tools and their 24b+ traffic behavior
Sujan Sarkar. Ai industry analysis: 50 most visited ai tools and their 24b+ traffic behavior. https://writ erbuddy.ai/blog/ai-industry-analysis/, 2023. Accessed: 2025-01-08
2023
-
[67]
Saturation in qualitative re- search: exploring its conceptualization and operational- ization
Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. Saturation in qualitative re- search: exploring its conceptualization and operational- ization. Quality & quantity, 52:1893–1907, 2018
1907
-
[68]
community guidelines make this the best party on the internet
Brennan Schaffner, Arjun Nitin Bhagoji, Siyuan Cheng, Jacqueline Mei, Jay L Shen, Grace Wang, Marshini Chetty, Nick Feamster, Genevieve Lakier, and Chenhao Tan. " community guidelines make this the best party on the internet": An in-depth study of online platforms’ content mod...
2024
-
[69]
Implications of regulations on large generative ai models in the super-election year and the impact on disinformation
Vera Schmitt, Jakob Tesch, Eva Lopez, Tim Polzehl, Aljoscha Burchardt, Konstanze Neumann, Salar Mohtaj, and Sebastian Möller. Implications of regulations on large generative ai models in the super-election year and the impact on disinformation. In Proceedings of the Workshop o...
2024
-
[70]
Safe latent diffusion: Mitigat- ing inappropriate degeneration in diffusion models
Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting. Safe latent diffusion: Mitigat- ing inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 22522–22531, 2023
2023
-
[71]
Proximal policy optimiza- tion algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[72]
Reconsidering self-moderation: the role of research in supporting community-based models for online content moderation
Joseph Seering. Reconsidering self-moderation: the role of research in supporting community-based models for online content moderation. Proceedings of the ACM on Human-Computer Interaction, 4(CSCW2):1–28, 2020
2020
-
[73]
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots
Wai Man Si, Michael Backes, Jeremy Blackburn, Emil- iano De Cristofaro, Gianluca Stringhini, Savvas Zannet- tou, and Yang Zhang. Why so toxic? measuring and triggering toxic behavior in open-domain chatbots. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Comm...
2022
-
[74]
Sok: Content moderation in social media, from guidelines to enforcement, and research to practice
Mohit Singhal, Chen Ling, Pujan Paudel, Poojitha Thota, Nihal Kumarswamy, Gianluca Stringhini, and Shirin Nilizadeh. Sok: Content moderation in social media, from guidelines to enforcement, and research to practice. In 2023 IEEE 8th European Symposium on Security and Privacy (...
2023
-
[75]
Intention of generative artificial in- telligence (gai) usage by adults in the united states as of august 2023, by type
Statista. Intention of generative artificial in- telligence (gai) usage by adults in the united states as of august 2023, by type. https: //www.statista.com/statistics/1461998/us a-generative-ai-usage-intention-by-type/ ,
2023
-
[76]
Use of generative artificial intelligence (ai) programs in the united states in 2023, by use case
Statista. Use of generative artificial intelligence (ai) programs in the united states in 2023, by use case. https://www.statista.com/statistics/ 1413836/use-of-generative-ai-us/ , 2024. Ac- cessed: 2025-01-08
2023
-
[77]
Lawless: The secret rules that govern our digital lives
Nicolas P Suzor. Lawless: The secret rules that govern our digital lives. Cambridge University Press, 2019
2019
-
[79]
Tiktok adds more generative ai features
Social Media Today. Tiktok adds more generative ai features. https://www.socialmediatoday.c om/news/tiktok-adds-gen-ai-image-tools-c aption-suggestions/736361/, 2025. Accessed: 2025-01-19
2025
-
[80]
The digital services act and the eu as the global regulator of the internet
Ioanna Tourkochoriti. The digital services act and the eu as the global regulator of the internet. Chi. J. Int’l L., 24:129, 2023
2023
-
[81]
Investigating moderation challenges to com- bating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on reddit
Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. Investigating moderation challenges to com- bating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on reddit. In 33rd USENIX Security Symposium (USENIX Security 24) , pages 37–54, Phila...
2024
-
[82]
House of Representatives
U.S. House of Representatives. United states code: Title 15, section 9401 (preliminary edition). https://uscode.house.gov/view.xhtml?req= (title:15%20section:9401%20edition:prelim),
-
[83]
at the end of the day facebook does what itwants
Kristen Vaccaro, Christian Sandvig, and Karrie Kara- halios. " at the end of the day facebook does what itwants" how users experience contesting algorithmic content moderation. Proceedings of the ACM on human- computer interaction, 4(CSCW2):1–22, 2020
2020
-
[84]
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[85]
Moderator: Moderating text-to- image diffusion models through fine-grained context- based policies
Peiran Wang, Qiyu Li, Longxuan Yu, Ziyao Wang, Ang Li, and Haojian Jin. Moderator: Moderating text-to- image diffusion models through fine-grained context- based policies. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 1181–1...
2024
-
[86]
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021
2021 arXiv
-
[87]
Understanding the im- pact of ai-generated content on social media: The pixiv case
Yiluo Wei and Gareth Tyson. Understanding the im- pact of ai-generated content on social media: The pixiv case. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 6813–6822, 2024
2024
-
[88]
Contestability for content moderation
Kristen Vaccaro, Ziang Xiao, Kevin Hamilton, and Kar- rie Karahalios. Contestability for content moderation. Proc. ACM Hum.-Comput. Interact., 5(CSCW2), Octo- ber 2021
2021
-
[89]
Ai-generated content (aigc): A survey
Jiayang Wu, Wensheng Gan, Zefeng Chen, Shicheng Wan, and Hong Lin. Ai-generated content (aigc): A survey. arXiv preprint arXiv:2304.06632, 2023
2023 arXiv
-
[90]
Fine-grained hu- man feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. Fine-grained hu- man feedback gives better rewards for language model training. Advances in Neural Information Processing Systems, 36:59008–5...
2023
-
[91]
Don’t listen to me: Understanding and exploring jailbreak prompts of large language models
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron, Chaowei Xiao, and Ning Zhang. Don’t listen to me: Understanding and exploring jailbreak prompts of large language models. In 33rd USENIX Security Symposium (USENIX Security 24), pages 4675–4692, Philadelphia, PA, August 2...
2024
-
[92]
as an ai language model, i cannot
Joel Wester, Tim Schrills, Henning Pohl, and Niels van Berkel. “as an ai language model, i cannot”: Investi- gating llm denials of user requests. In Proceedings of the CHI Conference on Human Factors in Computing Systems, pages 1–14, 2024
2024
-
[93]
Con- trollable safety alignment: Inference-time adaptation to diverse safety requirements
Jingyu Zhang, Ahmed Elgohary, Ahmed Magooda, Daniel Khashabi, and Benjamin Van Durme. Con- trollable safety alignment: Inference-time adaptation to diverse safety requirements. arXiv preprint arXiv:2410.08968, 2024
2024 arXiv
-
[94]
Siren’s song in the ai ocean: a survey on hallucination in large language models
Yue Zhang, Yafu Li, Leyang Cui, Deng Cai, Lemao Liu, Tingchen Fu, Xinting Huang, Enbo Zhao, Yu Zhang, Yulong Chen, et al. Siren’s song in the ai ocean: a survey on hallucination in large language models. arXiv preprint arXiv:2309.01219, 2023
2023 arXiv
-
[95]
the 2 characters’ foreheads are touching in a display of tenderness and affection
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G Parker, and Munmun De Choudhury. Synthetic lies: Un- derstanding ai-generated misinformation and evaluating algorithmic and human solutions. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1...
2023
-
[96]
i won the election!
Savvas Zannettou. " i won the election!": an empirical analysis of soft moderation interventions on twitter. In Proceedings of the international AAAI conference on web and social media, volume 15, pages 865–876, 2021
2021
-
[2025]
Accessed: 2025-01-08
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.