Pith. sign in

REVIEW 3 major objections 5 minor

Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor

T0 review · 3 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read This paper reports two randomized controlled trials on Nextdoor showing that filtering offensive content reduced its visibility by up to 95% while leaving platform visitation, content consumption, and content production unchanged.

desk verdict Two huge, clean field RCTs show opaque content filtering changes visibility but not behavior—worth engaging, though the abstract oversells one null and the Study 2 manipulation check is tighter than the headline claims. read the letter →

arxiv 2607.21853 v2 pith:AQCDB7GD submitted 2026-07-23 cs.HC cs.SI

classification cs.HCcs.SI
keywords contentmoderationoffensivefilteringtoxicrandomizedcontrolledtrialNextdoorPerspectiveAPIuserengagementplatformgovernance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to establish that reducing the visibility of offensive content—hiding or downranking posts that brush against a platform's rules—changes what users see but not what they do. Two randomized controlled trials on Nextdoor, each with 100,000 users, produced the same pattern: filtering cut views of offensive content by 12% in a report-triggered trial and by 95% in a proactive classifier-based trial, yet neither trial moved platform visitation, content consumption, or content production, including the creation of offensive content. The convergent nulls, read across manipulation strengths that differed by an order of magnitude, are meant to rule out the explanation that the first trial's nulls came from a weak intervention. A sympathetic reader takes away that opaque visibility-only moderation is effective at hiding specific content but is not, on its own, a lever for changing behavior or norms.

What carries the argument

The argument is carried by a paired design: each participant is enrolled upon an eligible platform visit, randomized to filtered or visible conditions, and followed 30 or 60 days before and after enrollment; a two-way mixed ANOVA with condition×time interaction isolates the intervention effect while absorbing the enrollment-day activity spike that both groups share. Study 2's manipulation is powered by proactive classification, with content scored by the Perspective API's 'Toxicity' attribute at a threshold of 0.70 and filtered from the newsfeed before accruing views. The escalation from a 12% report-triggered reduction to a 95% proactive reduction is the load-bearing axis that turns two nul

What would settle it

A falsifying observation would come from re-running Study 2's design with the outcome scored by an independent toxicity detector instead of the one used for filtering, or with an added condition that notifies authors when their content is filtered: if offensive-content creation drops in either variant, the nulls are specific to this implementation rather than to visibility reduction generally. Concretely, compare new-offensive-post counts in (a) opaque filtering, (b) filtering plus author notice, and (c) no filtering; a significant drop in (b) but not (a) would settle the paper's claim.

Watch

Extended reading notes

Core claim

The central claim, on the paper's own terms, is that filtering offensive content reliably achieves its direct target—reduced visibility—and has no detectable effect on any measured downstream behavior. Study 1 filtered comments in post threads after member reports, reducing views of those comments by 12%. Study 2 scored posts and comments at creation with Google Jigsaw's Perspective API and filtered content with Toxicity ≥ 0.70 from the newsfeed, suppressing related notifications, and reduced offensive-post views by 95%. Across the two trials, the time×condition interaction was null for platform sessions, content consumption, content production, and offensive-content creation, with Bayesian

Load-bearing premise

The load-bearing premise is that the operational definitions of 'offensive' (member reports in Study 1, Perspective Toxicity ≥ 0.70 in Study 2) capture what users actually experience as offensive and that the filter—not some unmeasured part of the treatment such as altered notifications—is the only difference between conditions; the paper itself notes in Section 6.1 that its measures capture behavior rather than attitudes and that both interventions were opaque by design.

Editorial extensions

If this is right

  • Filtering can reduce the visibility of offensive content by an order of magnitude without measurably reducing platform engagement, so platforms can treat it as a low-risk visibility tool.
  • The behavioral nulls replicate when manipulation strength rises from 12% to 95%, ruling out weak treatment as the explanation for the first study's nulls.
  • Opaque visibility reduction does not reduce future creation of offensive content, in contrast to removal-with-notice effects reported in prior moderation research.
  • Combining filtering with author-facing feedback—such as a notification that content was filtered—is the paper's suggested next step for turning visibility reduction into behavior change.
  • Regulatory regimes requiring notice of demotion (e.g., the EU Digital Services Act) may push platforms toward the transparent variants this paper identifies as untested.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if mere exposure to offensive content were a major driver of users' own offensive posting, a 95% reduction in such exposure should have moved production; its absence hints that authors' behavior is shaped more by awareness of moderation or by habits formed elsewhere than by what appears in the feed.
  • Editorial inference: the 95% manipulation check is partly mechanical because eligibility and outcome share the same Perspective threshold; a replication that scores outcomes with an independent detector would test whether the behavioral nulls survive classifier error.
  • Editorial inference: Nextdoor's hyperlocal, real-name setting may dampen effects that would appear in anonymous, interest-based communities, so the generalizability of the nulls is an open empirical question.
  • Editorial inference: the one significant secondary effect—lower average toxicity of comments viewed—suggests feed quality moved even though volume did not; future trials measuring exposure quality rather than volume may be more sensitive to filtering's effects.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports two large randomized controlled trials on Nextdoor, each with 100,000 users, testing interventions that reduce the visibility of offensive content without removing it or notifying authors. Study 1 (2022) used a report-triggered filter on comments in post threads, achieving a 12% reduction in views of offensive comments and finding no significant effects on eleven other platform-behavior measures. Study 2 (2023–2024) proactively scored posts and comments at creation with the Perspective API and filtered threshold-scored content from the newsfeed, achieving a 95% reduction in views of offensive posts, with no significant effects on most of thirteen further measures; the one exception was a small significant reduction in the average toxicity of comments viewed. The authors argue that the convergent null results across different content types, classifiers, and countries show that opaque visibility-reduction filters do not change downstream user behavior, while also demonstrating that such filters do not suppress engagement.

Significance. If the results hold, this is one of the first large-scale field experimental evaluations of a ubiquitous but understudied moderation practice. The design has important strengths: preregistration-like clarity in design comparison, broad outcome batteries, manipulation checks showing the filters operated as intended, and complementary frequentist and Bayesian analyses. The two trials are complementary in mechanism and manipulation strength, and the paper's framing as an integrated pair is a useful contribution. The central null result—that even a near-total reduction in views of classifier-defined offensive content did not move most engagement, consumption, or production outcomes—is a valuable addition to the evidence base on content moderation. The main weaknesses are internal inconsistencies in the abstract, the partly mechanical manipulation check in Study 2, and missing tables that prevent full verification of the reported analyses.

major comments (3)
  1. [Abstract and §5.2] The abstract claims 'across thirteen further measures we again found no significant effects' for Study 2, but §5.2 reports a significant time×condition interaction for average toxicity of comments viewed (F(1,99998)=7.779, p=.005, η²=4.07e−6). While the effect size is tiny, the statement as written is factually incorrect and the paper's cross-study claim of 'same behavioral nulls' is overstated. This inconsistency must be fixed, and the abstract/conclusions should acknowledge the one exception or justify why it is treated as negligible.
  2. [§5.1–5.2, Table 6] The Study 2 manipulation check is partly mechanical: offensive content is defined as Perspective Toxicity≥0.70 in §5.1, the filter removes exactly that content, and the manipulation check in §5.2 counts views of the same threshold-defined content. The 95% reduction therefore verifies that the filter ran, but it does not by itself establish that users experienced 95% less subjectively offensive content. If the classifier misses or mislabels content that users actually find offensive—e.g., content just below 0.70—the behavioral nulls may reflect the specific classifier and threshold rather than a general property of visibility-only moderation. The paper should either temper the generalization to 'offensive content' or provide evidence (e.g., validation of the threshold against user perceptions, or secondary analyses using alternative thresholds/definitions) that the filtered set correspond
  3. [Tables 5–7] Tables 5, 6, and 7 are not included in the manuscript; they are placeholders reading '(Carried over from the second manuscript.)' These tables contain the pre/post means, ANOVA F/p values for all fifteen Study 2 measures, and Bayes factors. Without them, the reader cannot verify the central Study 2 results. The tables must be inserted before the manuscript can be properly evaluated.
minor comments (5)
  1. [§5.2] The sentence reporting the significant interaction on average toxicity of comments viewed should be carried through to the abstract, the general discussion, and the limitations section; currently the abstract and the conclusion present the results as entirely null.
  2. [§4.1 and §5.1] The eligibility criteria for Study 1 (reported by at least one member for a 'hurtful or harmful' reason and scored above a Nextdoor model threshold) and Study 2 (Perspective Toxicity≥0.70) are presented as cutoffs without documenting how these thresholds were chosen. A sentence on the rationale or sensitivity of results to the specific threshold would strengthen the interpretation.
  3. [§3.2] The sensitivity analysis is described as indicating that effects of d=.02 can be detected, but no formal justification or simulation details are given. A brief note on the assumed variance structure would help the reader judge the claim.
  4. [Various tables] Some table headers use inconsistent capitalization and spacing (e.g., 'Manip. Check', 'Consump.', 'Prod.') while the text spells out the full names. Also, the effect sizes (η²) are omitted from Table 3 and only partially reported in text; the authors say they are in the supplement, but including them in the main tables would be more transparent.
  5. [§6.1] The limitations section does not discuss the validity of the Perspective API threshold as an operationalization of 'offensive' for the study population. Given that the manipulation check is defined by the same classifier, a few sentences on this limitation would be appropriate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the RCT design measures behavioral outcomes independently of the filter definition.

full rationale

This paper is an empirical randomized controlled trial, not a derivation, so the circularity burden is low. The only near-tautological element is the manipulation check: Study 2 defines filter eligibility by Perspective API Toxicity ≥ 0.70 (Section 5.1) and then measures views of threshold-eligible offensive posts, so the 95% drop partly verifies that the filter removed the content it was designed to remove. That is a fidelity check on the intervention, not a prediction from a fitted parameter, and it is not used as evidence for the behavioral nulls. All thirteen downstream outcomes — sessions, consumption, production, and offensive-content creation — are measured from platform records independently of the filtering rule. The creation-outcome threshold (≥0.60) also differs from the filter threshold, further separating the main outcomes from the intervention definition. The paper contains self-citations (e.g., the note that it supersedes two earlier preprints by the same authors), but these are disclosure statements or contextual related-work citations, not load-bearing justifications of the experimental results. There is no fitted parameter renamed as a prediction, no imported uniqueness theorem, and no ansatz smuggled in by citation. Potential construct-validity concerns — such as whether the Perspective threshold captures user-perceived offensiveness or whether the opaque treatment includes additional components — are threats to generalizability, not circularity.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

No new theoretical entities are introduced; the intervention (filtering) is an existing platform feature, not an invented entity. The key choices are classifier thresholds and the assumption that the classifier captures user-relevant offensiveness.

free parameters (3)
  • Perspective Toxicity filter threshold (Study 2) = 0.70
    Content scoring at or above 0.70 is hidden from the newsfeed; chosen by Nextdoor/Jigsaw, not derived in the paper, and directly defines which content is treated as offensive.
  • Offensive-created-content threshold (Study 2) = 0.60
    The outcome counts 'offensive posts/comments created' using a lower Perspective threshold than the filter threshold, a measurement choice that affects the observed null.
  • Study 1 removal-prediction model threshold = not specified
    Nextdoor's proprietary model scores reported comments for filter eligibility; the threshold is not disclosed in the paper.
assumptions (4)
  • domain assumption Perspective API Toxicity score is a valid operationalization of offensive content for both the intervention and the outcomes.
    The intervention and several outcomes rely on Perspective; if the classifier does not align with user-perceived offensiveness, the nulls may not generalize.
  • domain assumption No interference between treatment and control users (SUTVA) within neighborhoods.
    Randomization is at the member level, but members interact in shared neighborhoods; the design assumes filtering one user's feed does not change another user's exposure or behavior.
  • domain assumption Two-way mixed ANOVA inference is valid for highly skewed count outcomes at N=100,000.
    Outcomes such as comment views have SDs much larger than means (e.g., 2362 vs 330); the paper relies on large-N asymptotics without robust alternatives.
  • domain assumption Pre/post observation windows (30 and 60 days) capture the behavioral effects of interest.
    Effects that emerge only after longer exposure (or decay immediately) would be missed; the supplement extends to 80/200 days but only on a subsample.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor." pith.science (2026). https://pith.science/paper/AQCDB7GD

@misc{pith2026260721853,
  author       = {Pith},
  title        = {Pith review of: Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AQCDB7GD}},
  note         = {Machine review of arXiv:2607.21853}
}
read the original abstract

We investigate the effectiveness of interventions that reduce the visibility of offensive content on the local social platform Nextdoor. Content filtering -- hiding or downranking offensive content that brushes against a platform's rules without clearly breaking them -- is deployed across virtually every major platform, yet almost no field evidence exists on whether it changes user behavior. We report two large-scale randomized controlled trials, each involving 100,000 users. Study 1 (2022) tested a report-triggered filter applied to comments in post threads and produced a modest 12% reduction in views of offensive comments; across eleven further measures of platform behavior we found no significant effects. Study 2 (2023-2024) remedied Study 1's central limitation -- a weak manipulation driven by slow, report-based eligibility -- by proactively scoring posts and comments at creation with Google Jigsaw's Perspective API and filtering them from the newsfeed. This produced a near-complete (95%) reduction in views of offensive posts, yet across thirteen further measures we again found no significant effects. Across two independent trials spanning different content types, filtering mechanisms, classifiers, and countries -- and despite manipulation strength rising from 12% to 95% -- filtering reliably reduced the visibility of offensive content without altering platform visitation, content consumption, or content production. These convergent null results provide rare field evidence on a ubiquitous intervention and underscore the complexity of effectively moderating online platforms.

Figures

Figures reproduced from arXiv: 2607.21853 by the authors.

Figure 1
Figure 1. The comment-filter interface shown to test [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Study 1: daily means (±95% CI) for platform sessions, posts and comments created, posts and comments viewed, and offensive-comment views, 30 days before and after enrollment. Control in red, treatment in teal. Only offensive-comment views differs significantly between conditions. The single exception was the average toxicity of comments viewed (𝐹 (1, 99998) = 7.779, 𝑝 = .005, 𝜂 2 = 4.07e−6)—statistically signif￾ican… view at source ↗
Figure 3
Figure 3. Study 2: daily means (±95% CI) for platform sessions, posts created, posts viewed, and offensive-post views, 60 days before and after enrollment. Control in red, treatment in teal. Only offensive-post views differs significantly between condi￾tions [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.