{"id":"1854cbc5-ecc3-48ea-bf0d-1f0a47546fb0","arxiv_id":"2607.12149","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Writing a moderation policy as an LLM prompt cannot by itself ensure meaningful community governance.","lead":"This paper argues that using an LLM's system prompt to encode a community's moderation rules is not enough to govern content. It outlines risks to community governance and recommends that prompts be supported by evaluations, accountability structures, and human deliberation.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's conclusion presupposes that human moderation actually delivers deliberation, appeals, and contextual judgment; without evidence, the claimed degradation is not established.","rationale":"The reader identified the weakest assumption as the paper's reliance on human moderation actually delivering meaningful community governance. I agree. The paper's argument proceeds from a normative and empirical claim about human moderation that is presented as self-evident but is not substantiated. This is load-bearing because the conclusion that policy-as-prompt alone is 'not appropriate' relies on the premise that what is lost (deliberation, appeals, contextual interpretation) is valuable and present in current human moderation. If human moderation is often flawed or lacking these elements, the contrast is less stark, and the policy recommendation may need qualification. The paper does acknowledge some shortcomings of human moderation in passing (e.g., references to biases and performance drops), but these are framed as problems AI could 'ease' without questioning the existence of the governance goods. The concrete test I propose would provide empirical grounding for this premise by measuring the actual existence of appeals and deliberation in real communities. Depending on the result, the paper's critique would either be reinforced or weakened. Since the paper is a position paper and the concern is about an unexamined assumption rather than a logical contradiction, the appropriate verdict remains CONDITIONAL, as the reader set. No change is needed.","tokens_in":7061,"tokens_out":7112,"duration_ms":77939,"concrete_test":"Conduct a secondary analysis of existing moderation datasets (e.g., Reddit moderator logs, appeals data) to measure the actual prevalence of formal appeals channels and moderator deliberation. For instance, analyze a random sample of 1000 active subreddits to determine what fraction publicly document an appeals process or use mod discussion threads for ambiguous cases. If the fraction is low (e.g., <20%), the paper's idealization of human moderation is unsupported; if high, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that policy-as-prompt alone is not enough for meaningful community governance—depends on a sharp contrast with human moderation. Section 2 asserts that 'content moderation has never been just about the application of a set of textual rules: it is also a community governance practice,' citing ideal processes: moderator deliberation, appeals, and contextual interpretation. However, this paper provides no empirical evidence that these practices are realized in typical communities. Existing research (e.g., Cook et al. 2021; Kuo et al. 2023) suggests volunteer moderation is often informal, under-resourced, and without formal appeals. If human moderation is already inconsistent and lacks transparent deliberation, then replacing it with policy-as-prompt (especially with human oversight as recommended in Section 5) may not be a degradation. The paper's dichotomy between 'sense-making operations' performed by humans and 'warped mirror' LLM decisions (Section 6) assumes a level of human capability that is not demonstrated. This assumption is load-bearing: if human moderation does not deliver the asserted goods, the paper's conclusion that writing prompts alone is 'not appropriate' loses much of its practical force.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that content moderation based on 'policy-as-prompt' — encoding community or platform policy as a natural-language system prompt for an LLM — is technically and governance-wise insufficient. It reviews known limitations of prompt-based control (prompt injection, the instruction hierarchy, evaluation gaps), traces the risks for both centralized and decentralized moderation, and concludes that writing prompts alone cannot ensure meaningful community governance. The paper is explicitly a position/argumentative piece, not an empirical study: it states no new data, experiments, or measurements, and builds its case from prior literature and the authors' related work.","tokens_in":7290,"tokens_out":5983,"duration_ms":66514,"significance":"The manuscript fills a useful gap by translating recent technical findings about LLM prompt fragility into the vocabulary of community governance, and it gives concrete, actionable design considerations: sensitivity analyses, embedding prompts in broader governance structures, and requiring human-in-the-loop accountability. The paper is clearly written and appropriately hedged in Section 5, where it recommends that LLMs should only assist or supplement human moderation. If the core argument stands, it is a valuable caution against the 'easy edit the prompt' narrative in deployed moderation pipelines. The contribution is conceptual and policy-oriented rather than empirical; its strength lies in synthesis and in framing the problem for future evaluation and governance research.","major_comments":[{"comment":"The conclusion that policy-as-prompt is 'not appropriate for ensuring meaningful community governance' is framed as a loss or degradation from a baseline in which human moderation supplies deliberation, appeals, and contextual interpretation. Section 4.2 asserts that volunteer moderators 'write their own guidelines, deliberate over them in (mostly) transparent ways, and local norms evolve with the community and their discussions', supported only by one citation [5]. This is a substantive empirical claim, and the paper offers no evidence for it; prior work on volunteer moderation often documents under-resourcing, inconsistent enforcement, and absent or informal appeals. Because the argument's force depends on this baseline, the authors should either provide supporting evidence or explicitly reframe the claim as conditional: wherever such governance practices are valued and realized, promp","section":"§4.2 and §6"},{"comment":"The central premise that system prompts 'are not reliable enough to give governance guarantees' (Section 3) is attributed mainly to the authors' own prior work [23], which is a non-independent source. The paper should make the external evidence explicit and summarized: for example, prompt injection [10,29], the instruction hierarchy [36], and guardrail effectiveness studies [4,6]. In addition, the term 'governance guarantees' is never defined. If it means perfect robustness against all adversarial inputs, then no human or automated moderation system provides it, and the bar is unfair. If it means some weaker level of predictable reliability, the paper should specify that level and the evidence threshold. This matters because the phrase carries the main technical load.","section":"§3 and §5"},{"comment":"The target position is underspecified. 'Policy-as-prompt' could mean a single system prompt on a general-purpose LLM, a prompt stack with evaluation and guardrails, or a full pipeline with human appeal mechanisms. Section 5's own recommendations narrow the target substantially. The paper should state what exactly it criticizes: does it reject the claim that a prompt alone is sufficient, or does it reject any use of LLMs in moderation even with human oversight? The conclusion says 'writing prompts alone is not enough', but the body sometimes reads as objecting to the entire approach. Clarifying the scope of the claim would make the argument easier to evaluate.","section":"§3 and §5"}],"minor_comments":[{"comment":"'ease moderation burdens of time, mental health, and accuracy' — accuracy is not a burden; suggest 'ease time and mental-health burdens while improving accuracy'.","section":"Abstract"},{"comment":"'risks disempowering communities by destructing their influence' — 'destroying' or 'undermining' is the natural phrasing.","section":"§2"},{"comment":"'adding one instruction to a possibly unaligned hierarchy' — 'unaligned' is ambiguous; 'misaligned' or 'not aligned with the downstream policy' would be clearer.","section":"§3"},{"comment":"Reference [19] appears to have a malformed author list ('Zoe McMahon, Liv and Zoe Kleinman and Courtney Subramanian'); the entry should be cleaned up.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a workshop-style position paper rather than an empirical study. Its central argument is defensible, but the unexamined baseline of human moderation is load-bearing. I would not reject on that basis; a revision that either supplies evidence for the baseline or restates the claim in conditional terms would make it publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper argues that encoding moderation policy as a system prompt and handing it to an LLM cannot by itself deliver meaningful community governance. I think the argument holds as far as it goes, but it is not a new technical or empirical result. The core claim has already been made by the authors in earlier work and by Palla et al.'s 'policy-as-prompt' critique. What's new here is the focus on decentralized community moderation—the idea that volunteer communities lose more than rule enforcement when a prompt replaces human moderators, because the community loses the sense-making work that came with it. That is a genuinely useful extension.\n\nThe paper also does a few things well. It connects prompt stacks and instruction hierarchies to moderation practice, which is not a trivial link. It spells out why the ease of editing a prompt creates an incentive to use policy-as-prompt even when it isn't robust, and its final recommendations—evaluations, human oversight, appeals—are sensible and not overblown. There's no overclaiming: the authors explicitly call it a position piece, and they say 'writing prompts alone is not appropriate,' not 'LLM moderation is impossible.'\n\nThe soft spots are the usual ones for this genre. The paper has no data, so everything rests on the plausibility of the synthesis. The more specific issue is the dichotomy between human moderation and policy-as-prompt. The authors assume, in sections 2 and 4.2, that human moderation is a community governance practice with deliberation, appeals, and contextual judgment. That's an ideal. The literature they themselves cite (e.g., Cook et al., Kuo et al.) suggests volunteer moderation is often informal, under-resourced, and without formal appeals. If so, the claim that replacing it with an LLM degrades governance needs to be tested rather than assumed. Without that, the 'warped mirror' language is rhetorical.\n\nAlso, the central technical premise—system prompts cannot give governance guarantees—is largely drawn from the authors' own prior work. It is echoed by external sources, so I don't see it as a circularity flaw, but it does mean this paper is a synthesis, not an independent proof.\n\nBottom line: a useful and honest problem statement for a workshop audience, but it doesn't rise to an empirical contribution. I'd referee it if I were the editor—it's coherent and timely, and the recommendations are actionable—but I'd send it back with a request to either narrow the claims about human moderation or bring in evidence.","headline":"A coherent position paper that applies the known limits of prompt-based moderation to community governance; the claim is plausible but the idealization of human moderation and heavy reliance on the authors' prior work temper it.","tokens_in":7785,"tokens_out":2576,"would_cite":false,"duration_ms":27875,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that delivering a moderation policy to an LLM as a natural-language prompt cannot, by itself, guarantee reliable or accountable community governance.","keywords":["policy-as-prompt","content moderation","LLM moderation","community governance","system prompts","prompt governance","prompt injection","accountability"],"falsifier":"A longitudinal field study in which a community runs entirely on policy-as-prompt moderation and shows that the model's decisions track the community's evolving norms without human appeals, prompt updates, or oversight—while community members report the same or higher levels of procedural fairness—would undercut the paper's central claim.","tokens_in":6939,"feed_emoji":"🤖","tokens_out":2466,"duration_ms":28830,"temperature":0.7,"pith_summary":"This paper examines 'policy-as-prompt' moderation, where a community's rules are written as a natural-language instruction and given to a large language model to apply. It claims that this approach alone is not stable or robust enough to deliver the governance guarantees that content moderation requires. Moderation, the paper argues, is fundamentally a human sense-making practice involving deliberation, contextual interpretation, and appeals. Therefore, even as AI is integrated into moderation workflows, it should assist rather than replace human decisions, and must be embedded in broader governance structures.","feed_headline":"Policy-as-prompt alone can't govern communities","feed_subtitle":"LLM prompts lack the stability to replace human deliberation, appeals, and contextual judgment in moderation.","key_machinery":"The central object is the 'policy-as-prompt' configuration, in which a moderation policy is encoded as natural-language system instructions supplied to a general-purpose LLM. The argument is carried by two mechanisms: the 'prompt stack,' a hierarchical ordering that prioritizes instructions from foundation-model developers over downstream users, and 'prompt governance,' the idea that system prompts are not hard rules but unstable, contestable artifacts within a larger technical and institutional context.","core_discovery":"The central claim is that policy-as-prompt approaches on their own are not stable enough to deliver on the envisioned guarantees of alignment, performance, or robustness. Because prompts live within a hierarchical 'prompt stack' and can be overridden or circumvented, writing rules into a prompt is not equivalent to enforcing them. Moreover, moderation has never been just rule application: it is a community governance practice where moderators deliberate, interpret context, and provide recourse through appeals. Outsourcing these sense-making operations to an LLM risks disempowering communities and degrading the feedback loops that sustain self-governance. The paper concludes that writing prom","pith_inferences":["This reader infers that the argument generalizes beyond moderation: any delegated governance carried out through natural-language instructions—such as automated dispute resolution or algorithmic enforcement of workplace rules—faces the same instability and accountability deficits.","A testable extension is suggested by the paper's logic: in matched communities using identical written guidelines, those relying on policy-as-prompt moderation should show faster norm drift and lower perceived recourse than those with human moderation, a prediction that could be studied in a comparative field trial.","The paper's 'warped mirror of past language' point implies that moderation of emergent language—new slang, memes, or dialect shifts—will systematically lag under LLM-only regimes; this lag could be measured by tracking accuracy on novel expressions over time.","If accepted, the argument reframes 'prompt engineering' for moderation as a governance problem rather than a purely technical one, suggesting that regulators and platform operators should treat moderation prompts as public governance artifacts subject to transparency and audit requirements."],"forward_implications":["Organizations adopting policy-as-prompt must supplement prompts with rigorous evaluation, sensitivity analysis, and broader governance mechanisms; prompts alone cannot guarantee alignment or robustness.","Communities that outsource rule application to an LLM risk losing the situated expertise and feedback loops that make self-governance work, so moderation should remain human-assisted and appealable.","Because LLMs cannot take responsibility for their outputs, moderation actions must remain contestable and subject to human appeal; AI should assist rather than replace human moderators.","In centralized moderation, compressing deliberated policy into a machine-readable prompt shifts the burden of interpretation from an organization-community dialogue to an organization-AI translation, with unavoidable loss of nuance.","Community guidelines may be altered to fit what an AI can operationalize, meaning that norm evolution becomes partially governed by model affordances and technological change."],"fun_headline_variants":["LLM prompts can't replace human moderation governance","Policy-as-prompt alone undermines community governance","Prompt-based moderation lacks stability for true governance","Prompts alone can't deliver community self-governance","LLM-based moderation can't substitute for community governance"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument rests on the premise that human moderation currently performs genuinely deliberative, contextual, appealable governance, and that outsourcing to an LLM necessarily degrades these practices; if human moderation often fails to deliver these goods, the claim that prompts alone are insufficient loses much of its force.","fun_headline_variants_meta":{"raw":{"variants":["LLM prompts can't replace human moderation governance","Policy-as-prompt alone undermines community governance","Prompt-based moderation lacks stability for true governance","Prompts alone can't deliver community self-governance","LLM-based moderation can't substitute for community governance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2285,"prompt_tokens":708,"completion_tokens":1577,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":452,"completion_tokens_details":{"reasoning_tokens":1518}},"tokens_in":452,"tokens_out":1577,"duration_ms":10568,"temperature":1.0,"reasoning_tokens":1518,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:39:14.455316+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal field study in which a community runs entirely on policy-as-prompt moderation and shows that the model's decisions track the community's evolving norms without human appeals, prompt updates, or oversight—while community members report the same or higher levels of procedural fairness—would undercut the paper's central claim.","supporting_citations":[],"review_version":2}