{"id":"8b3e322c-115c-461f-a437-d415aa458dec","arxiv_id":"2510.04748","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across five mock-feed experiments, participants flagged ingroup-directed hate and abuse more than outgroup-directed abuse, while still flagging roughly half of abusive comments.","lead":"Five pre-registered experiments with US participants show that people flag abusive online comments more often when the target shares their political, vaccination, climate, or abortion stance. The findings suggest user-generated reports are not neutral signals, so platforms should treat flags as partially biased when moderating content.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 17–63% effect-size range is not derivable from the reported proportions; Study 4b shows only ~7% relative increase.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but the reader's weakest_assumption focuses on the absence of a manipulation check. While that is a valid internal-validity concern, the explicit group labels in the stimuli make it less likely that the target was not encoded. A more concrete and internally inconsistent issue is the abstract's effect-size claim, which the reader also flagged in passing. My stress-test identifies this numerical mismatch as the most load-bearing concern because it is directly falsifiable from the paper's own tables and figures. It does not overturn the qualitative conclusion of ingroup bias — the target main effect is significant in all five studies — but it does mean the paper's headline quantitative claim is currently unsupported. The CONDITIONAL verdict remains correct: the manuscript needs correction or clarification of the abstract, and possibly the General Discussion, before it can be accepted as-is. I do not see grounds to reject the core finding.","tokens_in":15469,"tokens_out":7037,"duration_ms":56787,"concrete_test":"Recompute the relative percentage increase in flagging for abusive comments directed at ingroup vs outgroup from the proportions reported in each study's Results (or from the OSF raw data). Then compare the resulting range to the abstract's '17% to 63%'. Also convert the reported GLMM odds ratios to relative risk using the outgroup base rates; if neither method yields the stated range, the abstract's numerical claim is erroneous and must be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central qualitative claim of ingroup bias may survive, but the quantitative headline in the abstract is contradicted by the paper's own results. From the reported proportions for abusive comments: Study 1 (50.5% vs 43.8%) gives a 15.3% relative increase; Study 2 (55.3% vs 49.4%) gives 11.9%; Study 3 (56.5% vs 48.6%) gives 16.3%; Study 4a (65.4% vs 53.7%) gives 21.8%; Study 4b (60.4% vs 56.6%) gives 6.7%. Three studies fall below the abstract's claimed lower bound of 17%, and none approach the claimed 63% upper bound. If the authors instead intended the GLMM odds ratios (1.9, 2.73, 3.25, 2.66, 1.51), those correspond to 90%, 173%, 225%, 166%, and 51% increases in odds — still not the stated 17–63% range. The abstract's precise numerical claim is therefore unsupported by the reported statistics. This matters because the abstract is the most visible summary of the find; an inaccurate effect-size range overstates the robustness and magnitude of the bias described in the General Discussion. The paper should either report the correct range or specify exactly how the 17–63% figure was derived.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports five pre-registered, US-based online experiments (total N ≈ 1,581 after exclusions) testing whether users' flagging of online hate/abuse is biased by group identity. Using mock social media feeds, participants saw abusive and mildly critical comments directed at Republicans/Democrats, pro-/anti-vaccination groups, climate-change believers/non-believers, and pro-/anti-abortion-rights groups. In all studies, participants flagged a majority of abusive comments (roughly 50–65%) and very few critical comments. The key claim is a main effect of target: comments directed at the ingroup are flagged significantly more than comments directed at the outgroup, with the effect replicated in all four social contexts and in both between-subjects (Studies 1–4a) and within-subjects (Study 4b) designs. The authors argue that user flags are therefore not a pure severity signal but are systematically skewed by social identity, with implications for platform moderation and intervention design.","tokens_in":15801,"tokens_out":7465,"duration_ms":58514,"significance":"If the central claim holds, the paper makes an important empirical contribution to the emerging literature on user flagging and online content moderation. Its strengths are substantial: all five studies are pre-registered, adequately powered, use balanced samples recruited on both sides of each issue, and combine a pre-registered ANOVA with a preregistered or complementary GLMM. The paradigm—using a mock feed with matched abusive/critical items—is a useful tool for measuring reporting bias experimentally. The paper also reports open data and materials at an OSF link. However, the quantitative headline in the abstract (a 17%–63% increase in flagging ingroup-directed abuse) is not supported by the reported proportions or odds ratios. This is a load-bearing error because it materially overstates the effect size and robustness that the data actually demonstrate. The qualitative pattern of ingroup bias is nonetheless consistent across all studies, which is the core message of the paper.","major_comments":[{"comment":"The abstract states that 'participants were between 17% and 63% more likely to flag abuse directed at the ingroup than at the outgroup.' This range is not derivable from any statistics reported in the manuscript. For abusive comments, the relative increases are: Study 1, (50.5−43.8)/43.8 = 15.3%; Study 2, (55.3−49.4)/49.4 = 11.9%; Study 3, (56.5−48.6)/48.6 = 16.3%; Study 4a, (65.4−53.7)/53.7 = 21.8%; Study 4b, (60.4−56.6)/56.6 = 6.7%. None of these reach 63%, and three fall below 17%. If the intended metric is the GLMM odds ratios (1.9, 2.73, 3.25, 2.66, 1.51), those correspond to 90%–225% increases in odds, not 17–63%. Please either correct the abstract to the actual range (≈7%–22% relative increase) or specify exactly how the 17–63% figure was computed. This is a central issue because the abstract is the most visible statement of the paper's findings and the quoted range overstates the","section":"Abstract"},{"comment":"No manipulation check is reported for the target-group manipulation. The design relies on participants perceiving whether each comment is directed at the ingroup or outgroup, but the only checks mentioned are two attention checks. If participants did not encode the target, the observed target main effects could be attenuated or driven by item-level differences despite the matching of items. Given that the central claim is that flagging is biased by group identity, the authors should report whether participants correctly identified the target group (e.g., from a recall check in the OSF materials) or acknowledge this as a limitation. Without such evidence, the internal validity of the key independent variable is not fully established.","section":"Study 1 Methods, pp. 9–14"},{"comment":"The General Discussion states that 'in all four social contexts that we examined, we found a robust ingroup bias effect in flagging.' However, the effect is asymmetrical in the political context: Study 1 shows a significant ingroup bias only among Republicans; Democrats flagged ingroup- and outgroup-directed abuse at similar rates (M_IG = 46.1%, M_OG = 45.6%, p = .901). The same asymmetry appears in Study 4b for abortion stance (pro-abortion participants show bias for abuse but not criticism; anti-abortion participants show bias for criticism but not abuse). The phrase 'robust ingroup bias effect' as a blanket summary is therefore too strong. Please qualify the claim to reflect the group-dependent nature of the effect, or provide a meta-analytic justification for treating the overall effect as robust.","section":"General Discussion, pp. 34–36"}],"minor_comments":[{"comment":"The submission title as provided is 'Ingroup bias is prevalent in user reports of hate and abuse online,' but the full-text manuscript title is 'Social bias is prevalent in user reports of hate and abuse online.' Please ensure the metadata and full text are consistent.","section":"Title / Metadata"},{"comment":"The phrase 'much online conversation may polarised in this way' contains a grammatical error; should be 'may be polarized.' Check for similar typos throughout (e.g., Table 1: 'supprters' and 'thieving').","section":"Study 4b, p. 29"},{"comment":"The claim that exposing participants to both ingroup- and outgroup-directed comments 'narrowed the social bias effect' is a post hoc comparison between Studies 4a and 4b. This cross-experiment comparison is not formally tested; the effect sizes (ηp² = .06 vs .054) are not directly comparable without an inferential statistic. Please present this as a tentative observation, not a finding.","section":"General Discussion, p. 36"},{"comment":"The abstract says 'approximately half of the abusive comments in each study reported.' In Studies 4a and 4b, overall abuse flagging was 59.6% and 58.5%, respectively. 'Approximately half' is acceptable, but 'approximately 50–60%' would be more precise.","section":"Abstract"},{"comment":"For Study 2, the footnote clarifies that the GLMM was the preregistered primary analysis but the ANOVA is reported first. For consistency, consider making the preregistration status of the GLMM explicit in all study Methods sections, not just as a footnote, so readers know which analysis corresponds to the pre-registration.","section":"Methods/Results"}],"recommendation":"major_revision","confidential_remarks":"The paper's core qualitative finding—that users flag ingroup-directed abuse more than outgroup-directed abuse—is well supported by significant main effects in all five studies. However, the abstract's quantitative claim (17–63%) is demonstrably inconsistent with the reported statistics and cannot be left as is. The missing manipulation check is a correctable but important gap. Given these issues, I recommend major revision, not rejection: the design is sound and the data are appropriate for the central claim, but the public-facing summary and internal-validity evidence need work before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core result here is solid: across five pre-registered, well-powered experiments, participants flag abuse directed at their ingroup more than abuse directed at an outgroup. That pattern holds in four different intergroup contexts, with a between-subjects design in Studies 1–4a and a within-subjects replication in 4b. The paradigm is a reasonable mock-feed setup, the same items are used across conditions with only the target label swapped, and the authors report both ANOVAs and mixed-effects logistic regressions. They also put data on OSF. That's a genuine contribution: previous work looked at misinformation flagging and offline intervention, not hate-speech flagging. The paper deserves credit for the scope and the honesty of the general discussion, which does a decent job acknowledging the artificial setting and the US-only sample.\n\nThe soft spots are real but not fatal. The biggest issue is the abstract's claim that participants were between 17% and 63% more likely to flag ingroup-directed abuse. That range is not derivable from the reported proportions. The relative increases are roughly 7%, 12%, 15%, 16%, and 22% across the five studies (for abusive comments). The GLMM odds ratios give larger-sounding effects, but those are odds ratios, not percentages, and they still don't produce a 17–63% range. So the abstract overstates both the floor and the ceiling. The qualitative claim is fine; the quantitative summary is not.\n\nThe second soft spot is the missing manipulation check. Participants' self-reported affiliation is used to define ingroup/outgroup, but there's no check that they actually noticed who the comments were directed at. If a participant skimmed the target word, the effect would be diluted, not inflated, so this is more likely a conservative bias—but it does mean the 'ingroup bias' label rests on an assumption. A quick post-task check would settle it.\n\nThird, the subgroup patterns are messier than the abstract suggests. Study 1 shows the effect only among Republicans, and Study 4b shows each side biased only for one language type (pro-choice for abuse, anti-abortion for criticism). The main effects are there, but the generality of 'pervasive ingroup bias' needs a more careful framing than the abstract gives it. The general discussion actually acknowledges some of this, which I appreciate.\n\nI'd send this to peer review. The core finding is a useful addition to the moderation literature, and the issues are fixable. I'd ask for a corrected abstract, a manipulation check or an explicit discussion of its absence, and a more nuanced summary of the subgroup effects. The paper is clearly the work of a serious research group, and the limitations they list are plausible rather than cosmetic.","headline":"The central finding is real and worth publishing, but the abstract's 17–63% range doesn't match the reported numbers, and the missing manipulation check leaves a small but real gap.","tokens_in":16228,"tokens_out":1317,"would_cite":true,"duration_ms":15711,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"People flag abuse aimed at their own group 17–63% more often than the same abuse aimed at the other side.","keywords":["ingroup bias","flagging","hate speech","online abuse","content moderation","user reporting","political polarization","online safety"],"falsifier":"Re-run Study 1 with a manipulation check asking each participant to identify the target group of each comment immediately after flagging. If the ingroup bias vanishes in those who fail the check, the effect is tied to attention to target; a real-world test could use platform logs of who flags which content and compare against the group identity of the comment's target.","tokens_in":15392,"feed_emoji":"🚩","tokens_out":6053,"duration_ms":43849,"temperature":0.7,"pith_summary":"The paper asks whether online 'flag' buttons produce an unbiased report of harmful content or whether group identity distorts what users choose to report. Across five pre-registered online experiments in the US, participants saw mock social-media comments aimed at either their own group or an opposing group and flagged roughly half of abusive comments overall. In every tested context—political affiliation, vaccination stance, climate-change belief, and abortion rights—abuse aimed at the ingroup was flagged more often than identical abuse aimed at the outgroup, by 17% to 63%. This matters because platforms and regulators increasingly rely on user flags as a crowdsourced moderation signal; if flags systematically under-report outgroup-directed abuse, they cannot be treated as a simple severity measure.","feed_headline":"People flag abuse of their own group up to 63% more","feed_subtitle":"User flags are a core moderation signal, so biased reporting skews what platforms see.","key_machinery":"The core instrument is a mock flagging task: each participant reads a neutral news excerpt and then 48 randomized comments (24 abusive, 24 mildly critical) aimed at a named group, and clicks a flag icon to report any they would normally report. The design is pre-registered; equal numbers of participants come from both sides of each issue, and the same comment wording is directed at the ingroup for half of participants and the outgroup for the other half, isolating the effect of target identity. Study 4b presents both targets in a single feed, adding a within-subjects test of whether seeing both sides' abuse reduces bias.","core_discovery":"On the paper's own terms: user flagging is a real but biased signal. In a mock social media feed, participants flagged about half of abusive comments and less than 5% of mild criticism, but they consistently flagged more when the comment attacked their own group. The effect held across all four intergroup contexts and ranged from a 17% to a 63% increase in flagging for ingroup-directed abuse. Seeing both groups abused in the same feed reduced but did not eliminate the bias. The authors conclude that flags cannot be treated as a neutral severity measure; they are shaped by the reporter's group membership.","pith_inferences":["On a real platform, this bias would compound the vulnerability of marginalized groups: outgroup-directed abuse is exactly the category users are least motivated to report, so the people most targeted also get the least crowdsourced protection.","The mock-feed setting with explicit instructions and attention checks likely yields higher overall flagging than natural browsing; in a noisy real feed, the relative weight of ingroup bias could be even larger or could dilute as users skim.","A direct design implication the authors do not draw: platforms could experimentally test whether showing users abuse aimed at both sides in a single view—as in Study 4b—reduces bias in live flagging data.","Because no manipulation check confirms that participants perceived the target group, this bias estimate may be an upper bound for attentive users; skimmers who do not encode the target might show no such bias, which itself would change what the finding means for moderation."],"forward_implications":["When users are explicitly invited to flag, reporting rates are high: roughly half of abusive comments are flagged, so low engagement may be partly a design problem rather than user indifference.","Outgroup-directed abuse is systematically under-flagged, meaning moderation queues built from user reports will under-represent attacks on outgroups relative to ingroups.","The bias is not limited to hate speech: mild criticism of one's own group is more likely to be misclassified as reportable than the same criticism of an outgroup.","Simultaneously showing abuse against both groups weakens the bias, suggesting that design choices about what users see can partially correct it."],"fun_headline_variants":["Flagging abuse skews toward own group, up to 63%","In-group bias: 17-63% more flags for own group","User flagging not neutral: up to 63% in-group bias","Online flaggers favor own group by up to 63%","Hate reports biased: 17-63% more for ingroup"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Participants' subjective group membership is self-reported and the target group is not checked after the task, so the whole ingroup-vs-outgroup comparison rests on respondents both encoding who the comment attacks and feeling the declared affiliation as their own.","fun_headline_variants_meta":{"raw":{"variants":["Flagging abuse skews toward own group, up to 63%","In-group bias: 17-63% more flags for own group","User flagging not neutral: up to 63% in-group bias","Online flaggers favor own group by up to 63%","Hate reports biased: 17-63% more for ingroup"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000125,"raw_usage":{"total_tokens":931,"prompt_tokens":719,"completion_tokens":212,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":118}},"tokens_in":463,"tokens_out":212,"duration_ms":2059,"temperature":1.0,"reasoning_tokens":118,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T11:23:08.661256+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run Study 1 with a manipulation check asking each participant to identify the target group of each comment immediately after flagging. If the ingroup bias vanishes in those who fail the check, the effect is tied to attention to target; a real-world test could use platform logs of who flags which content and compare against the group identity of the comment's target.","supporting_citations":[],"review_version":1}