Pith. sign in

REVIEW 3 major objections 4 minor 43 references

The paper argues that delivering a moderation policy to an LLM as a natural-language prompt cannot, by itself, guarantee reliable or accountable community governance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 06:39 UTC pith:TMRDEZY6

load-bearing objection A coherent position paper that applies the known limits of prompt-based moderation to community governance; the claim is plausible but the idealization of human moderation and heavy reliance on the authors' prior work temper it. the 3 major comments →

arxiv 2607.12149 v2 pith:TMRDEZY6 submitted 2026-07-13 cs.CY

It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance

classification cs.CY
keywords policy-as-promptcontent moderationLLM moderationcommunity governancesystem promptsprompt governanceprompt injectionaccountability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper examines 'policy-as-prompt' moderation, where a community's rules are written as a natural-language instruction and given to a large language model to apply. It claims that this approach alone is not stable or robust enough to deliver the governance guarantees that content moderation requires. Moderation, the paper argues, is fundamentally a human sense-making practice involving deliberation, contextual interpretation, and appeals. Therefore, even as AI is integrated into moderation workflows, it should assist rather than replace human decisions, and must be embedded in broader governance structures.

Core claim

The central claim is that policy-as-prompt approaches on their own are not stable enough to deliver on the envisioned guarantees of alignment, performance, or robustness. Because prompts live within a hierarchical 'prompt stack' and can be overridden or circumvented, writing rules into a prompt is not equivalent to enforcing them. Moreover, moderation has never been just rule application: it is a community governance practice where moderators deliberate, interpret context, and provide recourse through appeals. Outsourcing these sense-making operations to an LLM risks disempowering communities and degrading the feedback loops that sustain self-governance. The paper concludes that writing prom

What carries the argument

The central object is the 'policy-as-prompt' configuration, in which a moderation policy is encoded as natural-language system instructions supplied to a general-purpose LLM. The argument is carried by two mechanisms: the 'prompt stack,' a hierarchical ordering that prioritizes instructions from foundation-model developers over downstream users, and 'prompt governance,' the idea that system prompts are not hard rules but unstable, contestable artifacts within a larger technical and institutional context.

Load-bearing premise

The argument rests on the premise that human moderation currently performs genuinely deliberative, contextual, appealable governance, and that outsourcing to an LLM necessarily degrades these practices; if human moderation often fails to deliver these goods, the claim that prompts alone are insufficient loses much of its force.

What would settle it

A longitudinal field study in which a community runs entirely on policy-as-prompt moderation and shows that the model's decisions track the community's evolving norms without human appeals, prompt updates, or oversight—while community members report the same or higher levels of procedural fairness—would undercut the paper's central claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Organizations adopting policy-as-prompt must supplement prompts with rigorous evaluation, sensitivity analysis, and broader governance mechanisms; prompts alone cannot guarantee alignment or robustness.
  • Communities that outsource rule application to an LLM risk losing the situated expertise and feedback loops that make self-governance work, so moderation should remain human-assisted and appealable.
  • Because LLMs cannot take responsibility for their outputs, moderation actions must remain contestable and subject to human appeal; AI should assist rather than replace human moderators.
  • In centralized moderation, compressing deliberated policy into a machine-readable prompt shifts the burden of interpretation from an organization-community dialogue to an organization-AI translation, with unavoidable loss of nuance.
  • Community guidelines may be altered to fit what an AI can operationalize, meaning that norm evolution becomes partially governed by model affordances and technological change.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • This reader infers that the argument generalizes beyond moderation: any delegated governance carried out through natural-language instructions—such as automated dispute resolution or algorithmic enforcement of workplace rules—faces the same instability and accountability deficits.
  • A testable extension is suggested by the paper's logic: in matched communities using identical written guidelines, those relying on policy-as-prompt moderation should show faster norm drift and lower perceived recourse than those with human moderation, a prediction that could be studied in a comparative field trial.
  • The paper's 'warped mirror of past language' point implies that moderation of emergent language—new slang, memes, or dialect shifts—will systematically lag under LLM-only regimes; this lag could be measured by tracking accuracy on novel expressions over time.
  • If accepted, the argument reframes 'prompt engineering' for moderation as a governance problem rather than a purely technical one, suggesting that regulators and platform operators should treat moderation prompts as public governance artifacts subject to transparency and audit requirements.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper argues that content moderation based on 'policy-as-prompt' — encoding community or platform policy as a natural-language system prompt for an LLM — is technically and governance-wise insufficient. It reviews known limitations of prompt-based control (prompt injection, the instruction hierarchy, evaluation gaps), traces the risks for both centralized and decentralized moderation, and concludes that writing prompts alone cannot ensure meaningful community governance. The paper is explicitly a position/argumentative piece, not an empirical study: it states no new data, experiments, or measurements, and builds its case from prior literature and the authors' related work.

Significance. The manuscript fills a useful gap by translating recent technical findings about LLM prompt fragility into the vocabulary of community governance, and it gives concrete, actionable design considerations: sensitivity analyses, embedding prompts in broader governance structures, and requiring human-in-the-loop accountability. The paper is clearly written and appropriately hedged in Section 5, where it recommends that LLMs should only assist or supplement human moderation. If the core argument stands, it is a valuable caution against the 'easy edit the prompt' narrative in deployed moderation pipelines. The contribution is conceptual and policy-oriented rather than empirical; its strength lies in synthesis and in framing the problem for future evaluation and governance research.

major comments (3)
  1. [§4.2 and §6] The conclusion that policy-as-prompt is 'not appropriate for ensuring meaningful community governance' is framed as a loss or degradation from a baseline in which human moderation supplies deliberation, appeals, and contextual interpretation. Section 4.2 asserts that volunteer moderators 'write their own guidelines, deliberate over them in (mostly) transparent ways, and local norms evolve with the community and their discussions', supported only by one citation [5]. This is a substantive empirical claim, and the paper offers no evidence for it; prior work on volunteer moderation often documents under-resourcing, inconsistent enforcement, and absent or informal appeals. Because the argument's force depends on this baseline, the authors should either provide supporting evidence or explicitly reframe the claim as conditional: wherever such governance practices are valued and realized, promp
  2. [§3 and §5] The central premise that system prompts 'are not reliable enough to give governance guarantees' (Section 3) is attributed mainly to the authors' own prior work [23], which is a non-independent source. The paper should make the external evidence explicit and summarized: for example, prompt injection [10,29], the instruction hierarchy [36], and guardrail effectiveness studies [4,6]. In addition, the term 'governance guarantees' is never defined. If it means perfect robustness against all adversarial inputs, then no human or automated moderation system provides it, and the bar is unfair. If it means some weaker level of predictable reliability, the paper should specify that level and the evidence threshold. This matters because the phrase carries the main technical load.
  3. [§3 and §5] The target position is underspecified. 'Policy-as-prompt' could mean a single system prompt on a general-purpose LLM, a prompt stack with evaluation and guardrails, or a full pipeline with human appeal mechanisms. Section 5's own recommendations narrow the target substantially. The paper should state what exactly it criticizes: does it reject the claim that a prompt alone is sufficient, or does it reject any use of LLMs in moderation even with human oversight? The conclusion says 'writing prompts alone is not enough', but the body sometimes reads as objecting to the entire approach. Clarifying the scope of the claim would make the argument easier to evaluate.
minor comments (4)
  1. [Abstract] 'ease moderation burdens of time, mental health, and accuracy' — accuracy is not a burden; suggest 'ease time and mental-health burdens while improving accuracy'.
  2. [§2] 'risks disempowering communities by destructing their influence' — 'destroying' or 'undermining' is the natural phrasing.
  3. [§3] 'adding one instruction to a possibly unaligned hierarchy' — 'unaligned' is ambiguous; 'misaligned' or 'not aligned with the downstream policy' would be clearer.
  4. [References] Reference [19] appears to have a malformed author list ('Zoe McMahon, Liv and Zoe Kleinman and Courtney Subramanian'); the entry should be cleaned up.

Circularity Check

1 steps flagged

Central technical premise is imported from the authors' own prior work [23], but external corroboration and a largely definitional governance conclusion keep circularity moderate.

specific steps
  1. self citation load bearing [Section 3, 'Moderation through Policy-by-Prompt', paragraph 2]
    "However, as outlined in previous research [23], system instructions are not reliable enough to give governance guarantees towards goals such as alignment, performance, or robustness of the system."

    The paper's core technical premise — that policy-as-prompt cannot deliver governance guarantees — is framed as 'outlined in previous research [23]', where [23] is authored by the present first author (Neumann et al., FAccT '26). The current paper supplies no independent derivation, experiment, or machine-checked result for this premise; it imports the conclusion from its own prior work. The circular force is bounded because the same section also invokes external evidence (e.g., prompt injection [10, 29], guardrail failures [4, 6]), so the claim does not reduce solely to the self-citation. Still, the 'governance guarantees' framing itself is borrowed from the authors' own conclusions rather than demonstrated in this text.

full rationale

The paper is a position/argument paper, not an empirical derivation: it contains no fitted parameters, no 'prediction' computed from data, and no equation that could reduce to its inputs by construction. The central conclusion ('writing prompts alone is not appropriate for ensuring meaningful community governance') is a normative/analytical claim, and its main technical support is external literature on prompt injection, guardrail failures, and LLM moderation limitations, in addition to the self-cited [23]. The self-citation to [23] is load-bearing for the 'governance guarantees' framing, which raises the score above the purely independent baseline; however, because the argument is externally corroborated and not a statistical fit, the circularity is moderate (score 4), not severe. A related weakness is that the paper assumes without empirical evidence that human moderation actually realizes the deliberation, appeals, and contextual judgment it posits (Sections 2, 4.2, 6); this is an evidentiary gap or correctness risk, not a circular derivation, since current human moderation performance is not used as a fitted input. Overall, the paper does not reduce its conclusion to its inputs by definition or by self-citation alone.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

This is a non-quantitative argumentative paper: it introduces no fitted parameters and no new entities. Its load-bearing content is a set of domain assumptions about the value of human moderation, the limits of prompt-based control, and the accountability of AI systems. The central claim is supported primarily by these assumptions plus the authors' prior self-cited work.

axioms (4)
  • domain assumption Human moderation provides meaningful community governance through deliberation, contextual interpretation, and appeals.
    Stated in §2 and used in §4.2 to argue that outsourcing to LLMs would destroy these governance mechanisms; not empirically demonstrated in the paper.
  • domain assumption System-prompt instructions are not reliable enough to give governance guarantees (alignment, performance, robustness).
    Taken from the authors' prior work [23] in §3; the paper does not reproduce this result.
  • domain assumption Prompt stacks prioritize instructions by origin, so adding a moderation policy as a lower-priority prompt cannot reliably override higher-priority instructions.
    Invoked in §3 via [21, 36]; treated as a technical constraint.
  • domain assumption Generative AI systems cannot take responsibility for their outputs, so they cannot provide accountability for moderation actions.
    Used in §5 to argue LLMs should only assist human decisions; a philosophical/legal premise not argued in detail.

pith-pipeline@v1.3.0-alltime-deepseek · 6771 in / 12505 out tokens · 123001 ms · 2026-08-02T06:39:14.455316+00:00 · methodology

0 comments
read the original abstract

Content moderation practices and governance paradigms are changing rapidly, as fewer human moderators are deployed as `experts' by social media companies in a centralized manner. Instead, the companies are focusing more on community approaches, relying on volunteers to provide accurate information and make correct decisions. In decentralized moderation, communities have always relied on volunteers, updated community guidelines, and internal discussions thereof. For both content moderation paradigms, Artificial Intelligence (AI) seems like it could help ease moderation burdens of time, mental health, and accuracy. One possible way to operationalize AI in content moderation is a `policy-as-prompt'' approach, where the policy is formulated as a natural-language prompt and then passed to a large language model (LLM). This model then aids in moderation tasks. In this paper, we briefly lay out the technical and governance properties of this approach, and argue that its limitations lead to specific risks and harms that have to be addressed. Towards alleviating them, we lay out multiple considerations towards more effective prompt governance, but ultimately find that writing prompts alone is not appropriate for ensuring meaningful community governance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

43 extracted references · 7 canonical work pages · 1 internal anchor

  1. [1]

    [n. d.]. Scroll. Click. Suffer: The Hidden Human Cost of Content Moderation and Data Labelling – Equidem. https://equidem.org/reports/scroll-click-suffer-the- hidden-human-cost-of-content-moderation-and-data-labelling/

  2. [2]

    Mark Adkins. 2024. What should I say? – Interacting with AI and Natural Language Interfaces. https://doi.org/10.48550/arXiv.2401.06382 arXiv:2401.06382 [cs.HC]

  3. [3]

    Jonathan Klüser, Maël Kubli, and Nahema Marchal

    Meysam Alizadeh, Fabrizio Gilardi, Emma Hoes, K. Jonathan Klüser, Maël Kubli, and Nahema Marchal. 2022. Content Moderation As a Political Issue: The Twitter Discourse Around Trump’s Ban.Journal of Quantitative Description: Digital Media2 (Oct. 2022). https://doi.org/10.51685/jqd.2022.023

  4. [4]

    Giacomo Bertollo, Naz Bodemir, and Jonah Burgess. 2025. Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers. https: //doi.org/10.48550/arXiv.2510.16005 arXiv:2510.16005 [cs.CR]

  5. [5]

    Cook, Aashka Patel, and Donghee Yvette Wohn

    Christine L. Cook, Aashka Patel, and Donghee Yvette Wohn. 2021. Commercial Versus Volunteer: Comparing User Perceptions of Toxicity and Transparency in Content Moderation Across Social Media Platforms.Frontiers in Human Dynamics 3 (Feb. 2021), 626409. https://doi.org/10.3389/fhumd.2021.626409

  6. [6]

    Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, Saddek Bensalem, and Xiaowei Huang

  7. [7]

    Mirko Franco, Ombretta Gaggi, and Claudio E. Palazzi. 2025. Integrating Content Moderation Systems with Large Language Models.ACM Transactions on the Web 19, 2 (May 2025), 1–21. https://doi.org/10.1145/3700789

  8. [8]

    Juan Felipe Gomez, Caio Machado, Lucas Monteiro Paes, and Flavio Calmon. 2024. Algorithmic Arbitrariness in Content Moderation. InThe 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 2234–2253. https://doi.org/10.1145/3630106.3659036

  9. [9]

    Robert Gorwa, Reuben Binns, and Christian Katzenbach. 2020. Algorithmic con- tent moderation: Technical and political challenges in the automation of platform governance. https://journals.sagepub.com/doi/full/10.1177/2053951719897945

  10. [10]

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. InPro- ceedings of the 16th ACM Workshop on Artificial Intelligence and Security. ACM, Copenhagen Denmark, 79–90. https://doi.org/10.1145/...

  11. [11]

    David Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann, Simon Munz- ert, and Hendrik Heuer. 2025. Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, Yokohama Japan, 1–26. https:/...

  12. [12]

    Alice Hunsberger. [n. d.]. LLM Content Moderation: Implementation Guide for Trust & Safety Teams. https://musubilabs.ai/post/the-top-challenges-of-using- llms-for-content-moderation-and-how-to-overcome-them

  13. [13]

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, and Madian Khabsa. 2023. Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations. arXiv:2312.06674 [cs.CL] https://arxiv.org/abs/2312.06674

  14. [14]

    Hyunwoo Kim, Hanau Yi, Jaehee Bae, and Yumin Kim. 2026. Natural Lan- guage Declarative Prompting (NLD-P): A Modular Governance Method for Prompt Design Under Model Drift. https://doi.org/10.48550/arXiv.2602.22790 arXiv:2602.22790 [cs.CL]

  15. [15]

    Klonick and K

    K. Klonick and K. Klonick. 2017. The New Governors: The Peo- ple, Rules, and Processes Governing Online Speech.Harvard Law Review(March 2017). https://www.semanticscholar.org/paper/The- New-Governors%3A-The-People%2C-Rules%2C-and-Processes-Klonick- Klonick/cb52e32499d15ad3624a228a926416a3db14deb7

  16. [16]

    Deepak Kumar, Yousef AbuHashem, and Zakir Durumeric. 2023. Watch Your Language: Investigating Content Moderation with Large Language Models. https: //doi.org/10.48550/ARXIV.2309.14517 Version Number: 2

  17. [17]

    Tina Kuo, Alicia Hernani, and Jens Grossklags. 2023. The Unsung Heroes of Facebook Groups Moderation: A Case Study of Moderation Practices and Tools. Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (April 2023), 1–38. https://doi.org/10.1145/3579530

  18. [18]

    Artificial Intel- ligence as a Service

    Kornel Lewicki, Michelle Seng Ah Lee, Jennifer Cobbe, and Jatinder Singh. 2023. Out of Context: Investigating the Bias and Fairness Concerns of “Artificial Intel- ligence as a Service”. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Machinery, New York, NY, USA, 1–17. https://doi.org/10....

  19. [19]

    Zoe McMahon, Liv andZoe Kleinman and Courtney Subramanian. 2025. Meta to replace ’biased’ fact-checkers with moderation by users — bbc.com. https: //www.bbc.com/news/articles/cly74mpy8klo [Accessed 22-06-2026]

  20. [20]

    M. Nelson. 2015.The Argonauts. Graywolf Press. https://books.google.de/books? id=F4QkBQAAQBAJ

  21. [21]

    Anna Neumann, Elisabeth Kirsten, Muhammad Bilal Zafar, and Jatinder Singh

  22. [22]

    Anna Neumann, Yulu Pi, and Jatinder Singh. 2026. Who Controls the Conver- sation? User Perspectives On Generative AI (LLM) System Prompts. https: //doi.org/10.1145/3772318.3791726 arXiv:2603.00089 [cs]

  23. [23]

    InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency

    Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs). InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. ACM, Athens Greece, 573–598. https://doi. org/10.1145/3715275.3732038

  24. [24]

    Anna Neumann and Jatinder Singh. 2026. AI Safety Evaluations Need To Consider Cascading Effects. https://doi.org/10.48550/arXiv.2603.00088 arXiv:2603.00088 [cs]

  25. [25]

    Anna Neumann, Holli Sargeant, and Jatinder Singh. 2026. Prompt Governance? On Governing Technologies Governed by Natural Language. InThe 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’26)(Montreal, QC, Canada)(FAccT ’26). Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3805689.3806763

  26. [26]

    Chris Norval, Kristin Cornelius, Jennifer Cobbe, and Jatinder Singh. 2022. Disclo- sure by Design: Designing information disclosures to support meaningful trans- parency and accountability. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22). Association for Computing Machin- ery, New York, NY, USA, 679–690. ...

  27. [27]

    Casey Newton. [n. d.]. The secret lives of Facebook moderators in America. ([n. d.])

  28. [28]

    Taber, Andreas Damianou, and Mounia Lalmas

    Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri, Henrik Lindström, Daniel R. Taber, Andreas Damianou, and Mounia Lalmas. 2025. Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models. https://doi.org/10.48550/arXiv.2502.18695 arXiv:2502.18695 [cs]

  29. [29]

    OpenAI. 2025. Introducing gpt-oss-safeguard. https://openai.com/index/ introducing-gpt-oss-safeguard/

  30. [30]

    Minna Ruckenstein and Linda Lisa Maria Turunen. 2020. Re-humanizing the platform: Content moderators and the logic of care.New Media & Society22, 6 (June 2020), 1026–1042. https://doi.org/10.1177/1461444819875990

  31. [31]

    Fábio Perez and Ian Ribeiro. 2022. Ignore Previous Prompt: Attack Techniques For Language Models. https://doi.org/10.48550/arXiv.2211.09527 arXiv:2211.09527 [cs.CL]

  32. [32]

    Community Guidelines Make this the Best Party on the Internet

    Brennan Schaffner, Arjun Nitin Bhagoji, Siyuan Cheng, Jacqueline Mei, Jay L Shen, Grace Wang, Marshini Chetty, Nick Feamster, Genevieve Lakier, and Chenhao Tan. 2024. "Community Guidelines Make this the Best Party on the Internet": An In-Depth Study of Online Platforms’ Content Moderation Policies. InProceedings of the CHI Conference on Human Factors in C...

  33. [33]

    HariKiran

    Yalamanchili Salini and J. HariKiran. 2023. Sarcasm Detection: A Systematic Review of Methods and Approaches. In2023 3rd International Conference on Smart Data Intelligence (ICSMDI). IEEE, Trichy, India, 15–22. https://doi.org/10.1109/ ICSMDI57622.2023.00012

  34. [34]

    Suzor, Sarah Myers West, Andrew Quodling, and Jillian York

    Nicolas P. Suzor, Sarah Myers West, Andrew Quodling, and Jillian York. 2019. What Do We Mean When We Talk About Transparency? Toward Meaningful Transparency in Commercial Content Moderation.International Journal of Com- munication13 (March 2019), 18–18. https://ijoc.org/index.php/ijoc/article/view/ 9736

  35. [35]

    Farhana Shahid, Mona Elswah, and Aditya Vashistha. 2025. Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines MuC’26, 30. August - 02. September 2026, Duisburg, Germany Neumann et al. for Low-Resource Languages. https://doi.org/10.48550/arXiv.2501.13836 arXiv:2501.13836 [cs.CL]

  36. [36]

    Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. https://doi.org/10.48550/arXiv.2404.13208 arXiv:2404.13208 [cs]

  37. [37]

    Jacob van de Kerkhof. 2025. Musk, Techbrocracy, and Free Speech — verfas- sungsblog.de. https://verfassungsblog.de/musk-techbrocracy-and-free-speech/ [Accessed 22-06-2026]

  38. [38]

    Hiromu Yakura, Ezequiel Lopez-Lopez, Levin Brinkmann, Ignacio Serna, Prateek Gupta, Ivan Soraperra, and Iyad Rahwan. 2025. Empirical evidence of Large Language Model’s influence on human spoken communication. https://doi.org/ 10.48550/arXiv.2409.01754 arXiv:2409.01754 [cs.CY]

  39. [39]

    Sarah Myers West. 2017. Raging Against the Machine: Network Gatekeeping and Collective Action on Social Media Platforms.Media and Communication5, 3 (Sept. 2017), 28–36. https://doi.org/10.17645/mac.v5i3.989

  40. [40]

    Jing Zeng, Qinghao Guan, Ariadna Matamoros-Fernández, and Xiran Liu. 2025. How do multi-modal large language models understand non-English visual hate? Insights from studying hate speech in Chinese-speaking communities on Instagram.Platforms & Society2 (Sept. 2025), 29768624251383735. https: //doi.org/10.1177/29768624251383735

  41. [41]

    Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan, Hiok Hian Ong, David Yuchen Wang, Kanchan Sarkar, Zhenheng Yang, and Danhui Guan. 2026. Dynamic Content Moderation in Livestreams: Combining Supervised Classi- fication with MLLM-Boosted Similarity Matching. InProceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’...

  42. [43]

    Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents. https://doi.org/10.48550/arXiv.2403.02691 arXiv:2403.02691 [cs.CL]

  43. [2025]

    2025), 382

    Safeguarding large language models: a survey.Artificial Intelligence Review 58, 12 (Oct. 2025), 382. https://doi.org/10.1007/s10462-025-11389-2