REVIEW 3 major objections 4 minor 43 references
The paper argues that delivering a moderation policy to an LLM as a natural-language prompt cannot, by itself, guarantee reliable or accountable community governance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 06:39 UTC pith:TMRDEZY6
load-bearing objection A coherent position paper that applies the known limits of prompt-based moderation to community governance; the claim is plausible but the idealization of human moderation and heavy reliance on the authors' prior work temper it. the 3 major comments →
It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that policy-as-prompt approaches on their own are not stable enough to deliver on the envisioned guarantees of alignment, performance, or robustness. Because prompts live within a hierarchical 'prompt stack' and can be overridden or circumvented, writing rules into a prompt is not equivalent to enforcing them. Moreover, moderation has never been just rule application: it is a community governance practice where moderators deliberate, interpret context, and provide recourse through appeals. Outsourcing these sense-making operations to an LLM risks disempowering communities and degrading the feedback loops that sustain self-governance. The paper concludes that writing prom
What carries the argument
The central object is the 'policy-as-prompt' configuration, in which a moderation policy is encoded as natural-language system instructions supplied to a general-purpose LLM. The argument is carried by two mechanisms: the 'prompt stack,' a hierarchical ordering that prioritizes instructions from foundation-model developers over downstream users, and 'prompt governance,' the idea that system prompts are not hard rules but unstable, contestable artifacts within a larger technical and institutional context.
Load-bearing premise
The argument rests on the premise that human moderation currently performs genuinely deliberative, contextual, appealable governance, and that outsourcing to an LLM necessarily degrades these practices; if human moderation often fails to deliver these goods, the claim that prompts alone are insufficient loses much of its force.
What would settle it
A longitudinal field study in which a community runs entirely on policy-as-prompt moderation and shows that the model's decisions track the community's evolving norms without human appeals, prompt updates, or oversight—while community members report the same or higher levels of procedural fairness—would undercut the paper's central claim.
If this is right
- Organizations adopting policy-as-prompt must supplement prompts with rigorous evaluation, sensitivity analysis, and broader governance mechanisms; prompts alone cannot guarantee alignment or robustness.
- Communities that outsource rule application to an LLM risk losing the situated expertise and feedback loops that make self-governance work, so moderation should remain human-assisted and appealable.
- Because LLMs cannot take responsibility for their outputs, moderation actions must remain contestable and subject to human appeal; AI should assist rather than replace human moderators.
- In centralized moderation, compressing deliberated policy into a machine-readable prompt shifts the burden of interpretation from an organization-community dialogue to an organization-AI translation, with unavoidable loss of nuance.
- Community guidelines may be altered to fit what an AI can operationalize, meaning that norm evolution becomes partially governed by model affordances and technological change.
Where Pith is reading between the lines
- This reader infers that the argument generalizes beyond moderation: any delegated governance carried out through natural-language instructions—such as automated dispute resolution or algorithmic enforcement of workplace rules—faces the same instability and accountability deficits.
- A testable extension is suggested by the paper's logic: in matched communities using identical written guidelines, those relying on policy-as-prompt moderation should show faster norm drift and lower perceived recourse than those with human moderation, a prediction that could be studied in a comparative field trial.
- The paper's 'warped mirror of past language' point implies that moderation of emergent language—new slang, memes, or dialect shifts—will systematically lag under LLM-only regimes; this lag could be measured by tracking accuracy on novel expressions over time.
- If accepted, the argument reframes 'prompt engineering' for moderation as a governance problem rather than a purely technical one, suggesting that regulators and platform operators should treat moderation prompts as public governance artifacts subject to transparency and audit requirements.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that content moderation based on 'policy-as-prompt' — encoding community or platform policy as a natural-language system prompt for an LLM — is technically and governance-wise insufficient. It reviews known limitations of prompt-based control (prompt injection, the instruction hierarchy, evaluation gaps), traces the risks for both centralized and decentralized moderation, and concludes that writing prompts alone cannot ensure meaningful community governance. The paper is explicitly a position/argumentative piece, not an empirical study: it states no new data, experiments, or measurements, and builds its case from prior literature and the authors' related work.
Significance. The manuscript fills a useful gap by translating recent technical findings about LLM prompt fragility into the vocabulary of community governance, and it gives concrete, actionable design considerations: sensitivity analyses, embedding prompts in broader governance structures, and requiring human-in-the-loop accountability. The paper is clearly written and appropriately hedged in Section 5, where it recommends that LLMs should only assist or supplement human moderation. If the core argument stands, it is a valuable caution against the 'easy edit the prompt' narrative in deployed moderation pipelines. The contribution is conceptual and policy-oriented rather than empirical; its strength lies in synthesis and in framing the problem for future evaluation and governance research.
major comments (3)
- [§4.2 and §6] The conclusion that policy-as-prompt is 'not appropriate for ensuring meaningful community governance' is framed as a loss or degradation from a baseline in which human moderation supplies deliberation, appeals, and contextual interpretation. Section 4.2 asserts that volunteer moderators 'write their own guidelines, deliberate over them in (mostly) transparent ways, and local norms evolve with the community and their discussions', supported only by one citation [5]. This is a substantive empirical claim, and the paper offers no evidence for it; prior work on volunteer moderation often documents under-resourcing, inconsistent enforcement, and absent or informal appeals. Because the argument's force depends on this baseline, the authors should either provide supporting evidence or explicitly reframe the claim as conditional: wherever such governance practices are valued and realized, promp
- [§3 and §5] The central premise that system prompts 'are not reliable enough to give governance guarantees' (Section 3) is attributed mainly to the authors' own prior work [23], which is a non-independent source. The paper should make the external evidence explicit and summarized: for example, prompt injection [10,29], the instruction hierarchy [36], and guardrail effectiveness studies [4,6]. In addition, the term 'governance guarantees' is never defined. If it means perfect robustness against all adversarial inputs, then no human or automated moderation system provides it, and the bar is unfair. If it means some weaker level of predictable reliability, the paper should specify that level and the evidence threshold. This matters because the phrase carries the main technical load.
- [§3 and §5] The target position is underspecified. 'Policy-as-prompt' could mean a single system prompt on a general-purpose LLM, a prompt stack with evaluation and guardrails, or a full pipeline with human appeal mechanisms. Section 5's own recommendations narrow the target substantially. The paper should state what exactly it criticizes: does it reject the claim that a prompt alone is sufficient, or does it reject any use of LLMs in moderation even with human oversight? The conclusion says 'writing prompts alone is not enough', but the body sometimes reads as objecting to the entire approach. Clarifying the scope of the claim would make the argument easier to evaluate.
minor comments (4)
- [Abstract] 'ease moderation burdens of time, mental health, and accuracy' — accuracy is not a burden; suggest 'ease time and mental-health burdens while improving accuracy'.
- [§2] 'risks disempowering communities by destructing their influence' — 'destroying' or 'undermining' is the natural phrasing.
- [§3] 'adding one instruction to a possibly unaligned hierarchy' — 'unaligned' is ambiguous; 'misaligned' or 'not aligned with the downstream policy' would be clearer.
- [References] Reference [19] appears to have a malformed author list ('Zoe McMahon, Liv and Zoe Kleinman and Courtney Subramanian'); the entry should be cleaned up.
Circularity Check
Central technical premise is imported from the authors' own prior work [23], but external corroboration and a largely definitional governance conclusion keep circularity moderate.
specific steps
-
self citation load bearing
[Section 3, 'Moderation through Policy-by-Prompt', paragraph 2]
"However, as outlined in previous research [23], system instructions are not reliable enough to give governance guarantees towards goals such as alignment, performance, or robustness of the system."
The paper's core technical premise — that policy-as-prompt cannot deliver governance guarantees — is framed as 'outlined in previous research [23]', where [23] is authored by the present first author (Neumann et al., FAccT '26). The current paper supplies no independent derivation, experiment, or machine-checked result for this premise; it imports the conclusion from its own prior work. The circular force is bounded because the same section also invokes external evidence (e.g., prompt injection [10, 29], guardrail failures [4, 6]), so the claim does not reduce solely to the self-citation. Still, the 'governance guarantees' framing itself is borrowed from the authors' own conclusions rather than demonstrated in this text.
full rationale
The paper is a position/argument paper, not an empirical derivation: it contains no fitted parameters, no 'prediction' computed from data, and no equation that could reduce to its inputs by construction. The central conclusion ('writing prompts alone is not appropriate for ensuring meaningful community governance') is a normative/analytical claim, and its main technical support is external literature on prompt injection, guardrail failures, and LLM moderation limitations, in addition to the self-cited [23]. The self-citation to [23] is load-bearing for the 'governance guarantees' framing, which raises the score above the purely independent baseline; however, because the argument is externally corroborated and not a statistical fit, the circularity is moderate (score 4), not severe. A related weakness is that the paper assumes without empirical evidence that human moderation actually realizes the deliberation, appeals, and contextual judgment it posits (Sections 2, 4.2, 6); this is an evidentiary gap or correctness risk, not a circular derivation, since current human moderation performance is not used as a fitted input. Overall, the paper does not reduce its conclusion to its inputs by definition or by self-citation alone.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Human moderation provides meaningful community governance through deliberation, contextual interpretation, and appeals.
- domain assumption System-prompt instructions are not reliable enough to give governance guarantees (alignment, performance, robustness).
- domain assumption Prompt stacks prioritize instructions by origin, so adding a moderation policy as a lower-priority prompt cannot reliably override higher-priority instructions.
- domain assumption Generative AI systems cannot take responsibility for their outputs, so they cannot provide accountability for moderation actions.
read the original abstract
Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.
Reference graph
Works this paper leans on
-
[1]
[n. d.]. Scroll. Click. Suffer: The Hidden Human Cost of Content Moderation and Data Labelling – Equidem. https://equidem.org/reports/scroll-click-suffer-the- hidden-human-cost-of-content-moderation-and-data-labelling/
-
[2]
Mark Adkins. 2024. What should I say? – Interacting with AI and Natural Language Interfaces. https://doi.org/10.48550/arXiv.2401.06382 arXiv:2401.06382 [cs.HC]
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2401.06382 2024
-
[3]
Jonathan Klüser, Maël Kubli, and Nahema Marchal
Meysam Alizadeh, Fabrizio Gilardi, Emma Hoes, K. Jonathan Klüser, Maël Kubli, and Nahema Marchal. 2022. Content Moderation As a Political Issue: The Twitter Discourse Around Trump’s Ban.Journal of Quantitative Description: Digital Media2 (Oct. 2022). https://doi.org/10.51685/jqd.2022.023
-
[4]
Giacomo Bertollo, Naz Bodemir, and Jonah Burgess. 2025. Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers. https: //doi.org/10.48550/arXiv.2510.16005 arXiv:2510.16005 [cs.CR]
-
[5]
Cook, Aashka Patel, and Donghee Yvette Wohn
Christine L. Cook, Aashka Patel, and Donghee Yvette Wohn. 2021. Commercial Versus Volunteer: Comparing User Perceptions of Toxicity and Transparency in Content Moderation Across Social Media Platforms.Frontiers in Human Dynamics 3 (Feb. 2021), 626409. https://doi.org/10.3389/fhumd.2021.626409
arXiv 2021
-
[6]
Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, Saddek Bensalem, and Xiaowei Huang
-
[7]
Mirko Franco, Ombretta Gaggi, and Claudio E. Palazzi. 2025. Integrating Content Moderation Systems with Large Language Models.ACM Transactions on the Web 19, 2 (May 2025), 1–21. https://doi.org/10.1145/3700789
doi:10.1145/3700789 2025
-
[8]
Juan Felipe Gomez, Caio Machado, Lucas Monteiro Paes, and Flavio Calmon. 2024. Algorithmic Arbitrariness in Content Moderation. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 2234–2253. https://doi.org/10.1145/3630106.3659036
arXiv 2024
-
[9]
Robert Gorwa, Reuben Binns, and Christian Katzenbach. 2020. Algorithmic con- tent moderation: Technical and political challenges in the automation of platform governance. https://journals.sagepub.com/doi/full/10.1177/2053951719897945
-
[10]
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. InPro- ceedings of the 16th ACM Workshop on Artificial Intelligence and Security. ACM, Copenhagen Denmark, 79–90. https://doi.org/10.1145/...
arXiv 2023
-
[11]
David Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann, Simon Munz- ert, and Hendrik Heuer. 2025. Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, Yokohama Japan, 1–26. https:/...
arXiv 2025
-
[12]
Alice Hunsberger. [n. d.]. LLM Content Moderation: Implementation Guide for Trust & Safety Teams. https://musubilabs.ai/post/the-top-challenges-of-using- llms-for-content-moderation-and-how-to-overcome-them
-
[13]
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv:2312.06674 [cs.CL] https://arxiv.org/abs/2312.06674
Pith/arXiv arXiv 2023
-
[14]
Hyunwoo Kim, Hanau Yi, Jaehee Bae, and Yumin Kim. 2026. Natural Lan- guage Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift. https://doi.org/10.48550/arXiv.2602.22790 arXiv:2602.22790 [cs.CL]
-
[15]
Klonick and K
K. Klonick and K. Klonick. 2017. The New Governors: The Peo- ple, Rules, and Processes Governing Online Speech.Harvard Law Review(March 2017). https://www.semanticscholar.org/paper/The- New-Governors%3A-The-People%2C-Rules%2C-and-Processes-Klonick- Klonick/cb52e32499d15ad3624a228a926416a3db14deb7
2017
-
[16]
Deepak Kumar, Yousef AbuHashem, and Zakir Durumeric. 2023. Watch Your Language: Investigating Content Moderation with Large Language Models. https: //doi.org/10.48550/ARXIV.2309.14517 Version Number: 2
-
[17]
Tina Kuo, Alicia Hernani, and Jens Grossklags. 2023. The Unsung Heroes of Facebook Groups Moderation: A Case Study of Moderation Practices and Tools. Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (April 2023), 1–38. https://doi.org/10.1145/3579530
doi:10.1145/3579530 2023
-
[18]
Artificial Intel- ligence as a Service
Kornel Lewicki, Michelle Seng Ah Lee, Jennifer Cobbe, and Jatinder Singh. 2023. Out of Context: Investigating the Bias and Fairness Concerns of “Artificial Intel- ligence as a Service”. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–17. https://doi.org/10....
arXiv 2023
-
[19]
Zoe McMahon, Liv andZoe Kleinman and Courtney Subramanian. 2025. Meta to replace ’biased’ fact-checkers with moderation by users — bbc.com. https: //www.bbc.com/news/articles/cly74mpy8klo [Accessed 22-06-2026]
2025
-
[20]
M. Nelson. 2015.The Argonauts. Graywolf Press. https://books.google.de/books? id=F4QkBQAAQBAJ
2015
-
[21]
Anna Neumann, Elisabeth Kirsten, Muhammad Bilal Zafar, and Jatinder Singh
-
[22]
Anna Neumann, Yulu Pi, and Jatinder Singh. 2026. Who Controls the Conver- sation? User Perspectives On Generative AI (LLM) System Prompts. https: //doi.org/10.1145/3772318.3791726 arXiv:2603.00089 [cs]
arXiv 2026
-
[23]
InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs). InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. ACM, Athens Greece, 573–598. https://doi. org/10.1145/3715275.3732038
arXiv 2025
-
[24]
Anna Neumann and Jatinder Singh. 2026. AI Safety Evaluations Need To Consider Cascading Effects. https://doi.org/10.48550/arXiv.2603.00088 arXiv:2603.00088 [cs]
-
[25]
Anna Neumann, Holli Sargeant, and Jatinder Singh. 2026. Prompt Governance? On Governing Technologies Governed by Natural Language. InThe 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’26)(Montreal, QC, Canada)(FAccT ’26). Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3805689.3806763
arXiv 2026
-
[26]
Chris Norval, Kristin Cornelius, Jennifer Cobbe, and Jatinder Singh. 2022. Disclo- sure by Design: Designing information disclosures to support meaningful trans- parency and accountability. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machin- ery, New York, NY, USA, 679–690. ...
arXiv 2022
-
[27]
Casey Newton. [n. d.]. The secret lives of Facebook moderators in America. ([n. d.])
-
[28]
Taber, Andreas Damianou, and Mounia Lalmas
Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri, Henrik Lindström, Daniel R. Taber, Andreas Damianou, and Mounia Lalmas. 2025. Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models. https://doi.org/10.48550/arXiv.2502.18695 arXiv:2502.18695 [cs]
-
[29]
OpenAI. 2025. Introducing gpt-oss-safeguard. https://openai.com/index/ introducing-gpt-oss-safeguard/
2025
-
[30]
Minna Ruckenstein and Linda Lisa Maria Turunen. 2020. Re-humanizing the platform: Content moderators and the logic of care.New Media & Society22, 6 (June 2020), 1026–1042. https://doi.org/10.1177/1461444819875990
-
[31]
Fábio Perez and Ian Ribeiro. 2022. Ignore Previous Prompt: Attack Techniques For Language Models. https://doi.org/10.48550/arXiv.2211.09527 arXiv:2211.09527 [cs.CL]
-
[32]
Community Guidelines Make this the Best Party on the Internet
Brennan Schaffner, Arjun Nitin Bhagoji, Siyuan Cheng, Jacqueline Mei, Jay L Shen, Grace Wang, Marshini Chetty, Nick Feamster, Genevieve Lakier, and Chenhao Tan. 2024. "Community Guidelines Make this the Best Party on the Internet": An In-Depth Study of Online Platforms’ Content Moderation Policies. InProceedings of the CHI Conference on Human Factors in C...
arXiv 2024
- [33]
-
[34]
Suzor, Sarah Myers West, Andrew Quodling, and Jillian York
Nicolas P. Suzor, Sarah Myers West, Andrew Quodling, and Jillian York. 2019. What Do We Mean When We Talk About Transparency? Toward Meaningful Transparency in Commercial Content Moderation.International Journal of Com- munication13 (March 2019), 18–18. https://ijoc.org/index.php/ijoc/article/view/ 9736
2019
-
[35]
Farhana Shahid, Mona Elswah, and Aditya Vashistha. 2025. Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines MuC’26, 30. August - 02. September 2026, Duisburg, Germany Neumann et al. for Low-Resource Languages. https://doi.org/10.48550/arXiv.2501.13836 arXiv:2501.13836 [cs.CL]
-
[36]
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. https://doi.org/10.48550/arXiv.2404.13208 arXiv:2404.13208 [cs]
-
[37]
Jacob van de Kerkhof. 2025. Musk, Techbrocracy, and Free Speech — verfas- sungsblog.de. https://verfassungsblog.de/musk-techbrocracy-and-free-speech/ [Accessed 22-06-2026]
2025
-
[38]
Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio Serna, Prateek Gupta, Ivan Soraperra, and Iyad Rahwan. 2025. Empirical evidence of Large Language Model’s influence on human spoken communication. https://doi.org/ 10.48550/arXiv.2409.01754 arXiv:2409.01754 [cs.CY]
-
[39]
Sarah Myers West. 2017. Raging Against the Machine: Network Gatekeeping and Collective Action on Social Media Platforms.Media and Communication5, 3 (Sept. 2017), 28–36. https://doi.org/10.17645/mac.v5i3.989
-
[40]
Jing Zeng, Qinghao Guan, Ariadna Matamoros-Fernández, and Xiran Liu. 2025. How do multi-modal large language models understand non-English visual hate? Insights from studying hate speech in Chinese-speaking communities on Instagram.Platforms & Society2 (Sept. 2025), 29768624251383735. https: //doi.org/10.1177/29768624251383735
-
[41]
Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan, Hiok Hian Ong, David Yuchen Wang, Kanchan Sarkar, Zhenheng Yang, and Danhui Guan. 2026. Dynamic Content Moderation in Livestreams: Combining Supervised Classi- fication with MLLM-Boosted Similarity Matching. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’...
arXiv 2026
-
[43]
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. https://doi.org/10.48550/arXiv.2403.02691 arXiv:2403.02691 [cs.CL]
-
[2025]
Safeguarding large language models: a survey.Artificial Intelligence Review 58, 12 (Oct. 2025), 382. https://doi.org/10.1007/s10462-025-11389-2
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.