REVIEW 3 major objections 5 minor 16 references
Ingroup bias is prevalent in user reports of hate and abuse online
T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read People flag abuse aimed at their own group 17–63% more often than the same abuse aimed at the other side.
desk verdict The central finding is real and worth publishing, but the abstract's 17–63% range doesn't match the reported numbers, and the missing manipulation check leaves a small but real gap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core instrument is a mock flagging task: each participant reads a neutral news excerpt and then 48 randomized comments (24 abusive, 24 mildly critical) aimed at a named group, and clicks a flag icon to report any they would normally report. The design is pre-registered; equal numbers of participants come from both sides of each issue, and the same comment wording is directed at the ingroup for half of participants and the outgroup for the other half, isolating the effect of target identity. Study 4b presents both targets in a single feed, adding a within-subjects test of whether seeing both sides' abuse reduces bias.
What would settle it
Re-run Study 1 with a manipulation check asking each participant to identify the target group of each comment immediately after flagging. If the ingroup bias vanishes in those who fail the check, the effect is tied to attention to target; a real-world test could use platform logs of who flags which content and compare against the group identity of the comment's target.
Extended reading notes
Core claim
On the paper's own terms: user flagging is a real but biased signal. In a mock social media feed, participants flagged about half of abusive comments and less than 5% of mild criticism, but they consistently flagged more when the comment attacked their own group. The effect held across all four intergroup contexts and ranged from a 17% to a 63% increase in flagging for ingroup-directed abuse. Seeing both groups abused in the same feed reduced but did not eliminate the bias. The authors conclude that flags cannot be treated as a neutral severity measure; they are shaped by the reporter's group membership.
Load-bearing premise
Participants' subjective group membership is self-reported and the target group is not checked after the task, so the whole ingroup-vs-outgroup comparison rests on respondents both encoding who the comment attacks and feeling the declared affiliation as their own.
Editorial extensions
If this is right
- When users are explicitly invited to flag, reporting rates are high: roughly half of abusive comments are flagged, so low engagement may be partly a design problem rather than user indifference.
- Outgroup-directed abuse is systematically under-flagged, meaning moderation queues built from user reports will under-represent attacks on outgroups relative to ingroups.
- The bias is not limited to hate speech: mild criticism of one's own group is more likely to be misclassified as reportable than the same criticism of an outgroup.
- Simultaneously showing abuse against both groups weakens the bias, suggesting that design choices about what users see can partially correct it.
Reading between the lines
- On a real platform, this bias would compound the vulnerability of marginalized groups: outgroup-directed abuse is exactly the category users are least motivated to report, so the people most targeted also get the least crowdsourced protection.
- The mock-feed setting with explicit instructions and attention checks likely yields higher overall flagging than natural browsing; in a noisy real feed, the relative weight of ingroup bias could be even larger or could dilute as users skim.
- A direct design implication the authors do not draw: platforms could experimentally test whether showing users abuse aimed at both sides in a single view—as in Study 4b—reduces bias in live flagging data.
- Because no manipulation check confirms that participants perceived the target group, this bias estimate may be an upper bound for attentive users; skimmers who do not encode the target might show no such bias, which itself would change what the finding means for moderation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports five pre-registered, US-based online experiments (total N ≈ 1,581 after exclusions) testing whether users' flagging of online hate/abuse is biased by group identity. Using mock social media feeds, participants saw abusive and mildly critical comments directed at Republicans/Democrats, pro-/anti-vaccination groups, climate-change believers/non-believers, and pro-/anti-abortion-rights groups. In all studies, participants flagged a majority of abusive comments (roughly 50–65%) and very few critical comments. The key claim is a main effect of target: comments directed at the ingroup are flagged significantly more than comments directed at the outgroup, with the effect replicated in all four social contexts and in both between-subjects (Studies 1–4a) and within-subjects (Study 4b) designs. The authors argue that user flags are therefore not a pure severity signal but are systematically skewed by social identity, with implications for platform moderation and intervention design.
Significance. If the central claim holds, the paper makes an important empirical contribution to the emerging literature on user flagging and online content moderation. Its strengths are substantial: all five studies are pre-registered, adequately powered, use balanced samples recruited on both sides of each issue, and combine a pre-registered ANOVA with a preregistered or complementary GLMM. The paradigm—using a mock feed with matched abusive/critical items—is a useful tool for measuring reporting bias experimentally. The paper also reports open data and materials at an OSF link. However, the quantitative headline in the abstract (a 17%–63% increase in flagging ingroup-directed abuse) is not supported by the reported proportions or odds ratios. This is a load-bearing error because it materially overstates the effect size and robustness that the data actually demonstrate. The qualitative pattern of ingroup bias is nonetheless consistent across all studies, which is the core message of the paper.
major comments (3)
- [Abstract] The abstract states that 'participants were between 17% and 63% more likely to flag abuse directed at the ingroup than at the outgroup.' This range is not derivable from any statistics reported in the manuscript. For abusive comments, the relative increases are: Study 1, (50.5−43.8)/43.8 = 15.3%; Study 2, (55.3−49.4)/49.4 = 11.9%; Study 3, (56.5−48.6)/48.6 = 16.3%; Study 4a, (65.4−53.7)/53.7 = 21.8%; Study 4b, (60.4−56.6)/56.6 = 6.7%. None of these reach 63%, and three fall below 17%. If the intended metric is the GLMM odds ratios (1.9, 2.73, 3.25, 2.66, 1.51), those correspond to 90%–225% increases in odds, not 17–63%. Please either correct the abstract to the actual range (≈7%–22% relative increase) or specify exactly how the 17–63% figure was computed. This is a central issue because the abstract is the most visible statement of the paper's findings and the quoted range overstates the
- [Study 1 Methods, pp. 9–14] No manipulation check is reported for the target-group manipulation. The design relies on participants perceiving whether each comment is directed at the ingroup or outgroup, but the only checks mentioned are two attention checks. If participants did not encode the target, the observed target main effects could be attenuated or driven by item-level differences despite the matching of items. Given that the central claim is that flagging is biased by group identity, the authors should report whether participants correctly identified the target group (e.g., from a recall check in the OSF materials) or acknowledge this as a limitation. Without such evidence, the internal validity of the key independent variable is not fully established.
- [General Discussion, pp. 34–36] The General Discussion states that 'in all four social contexts that we examined, we found a robust ingroup bias effect in flagging.' However, the effect is asymmetrical in the political context: Study 1 shows a significant ingroup bias only among Republicans; Democrats flagged ingroup- and outgroup-directed abuse at similar rates (M_IG = 46.1%, M_OG = 45.6%, p = .901). The same asymmetry appears in Study 4b for abortion stance (pro-abortion participants show bias for abuse but not criticism; anti-abortion participants show bias for criticism but not abuse). The phrase 'robust ingroup bias effect' as a blanket summary is therefore too strong. Please qualify the claim to reflect the group-dependent nature of the effect, or provide a meta-analytic justification for treating the overall effect as robust.
minor comments (5)
- [Title / Metadata] The submission title as provided is 'Ingroup bias is prevalent in user reports of hate and abuse online,' but the full-text manuscript title is 'Social bias is prevalent in user reports of hate and abuse online.' Please ensure the metadata and full text are consistent.
- [Study 4b, p. 29] The phrase 'much online conversation may polarised in this way' contains a grammatical error; should be 'may be polarized.' Check for similar typos throughout (e.g., Table 1: 'supprters' and 'thieving').
- [General Discussion, p. 36] The claim that exposing participants to both ingroup- and outgroup-directed comments 'narrowed the social bias effect' is a post hoc comparison between Studies 4a and 4b. This cross-experiment comparison is not formally tested; the effect sizes (ηp² = .06 vs .054) are not directly comparable without an inferential statistic. Please present this as a tentative observation, not a finding.
- [Abstract] The abstract says 'approximately half of the abusive comments in each study reported.' In Studies 4a and 4b, overall abuse flagging was 59.6% and 58.5%, respectively. 'Approximately half' is acceptable, but 'approximately 50–60%' would be more precise.
- [Methods/Results] For Study 2, the footnote clarifies that the GLMM was the preregistered primary analysis but the ANOVA is reported first. For consistency, consider making the preregistration status of the GLMM explicit in all study Methods sections, not just as a footnote, so readers know which analysis corresponds to the pre-registration.
Circularity Check
No circularity: the central claim is an empirical behavioral finding, not a derivation from its own inputs.
full rationale
This paper reports five pre-registered experiments measuring flagging behaviour. There is no mathematical derivation chain in which a predicted quantity is definitionally identical to a fitted input: the ingroup/outgroup manipulation is operationalized from participants' own self-reported affiliation (e.g., 'the Target was an ingroup or an outgroup depending on participants’ own political affiliation'), and the outcome (flagging rates) is independently measured behaviour. No parameter is fitted to the target effect and then called a prediction; the target main effects are simply observed differences between conditions. Self-citations (e.g., Enock et al., 2020, 2025; Bright et al., 2024) appear only as background context or as prior survey evidence, not as load-bearing arguments that force the results. The abstract's '17% and 63%' range appears not to follow from the reported proportions, but that is a numerical reporting/consistency issue, not a circular reduction of the conclusion to its premises. Accordingly, no specific circular step can be identified by quoting the paper's own equations or construction.
Assumptions & free parameters
assumptions (4)
- domain assumption Self-reported group membership (e.g., Democrat/Republican, pro/anti-vaccination, climate believer/non-believer, pro/anti-abortion) corresponds to a psychologically meaningful ingroup in the experimental context.
- domain assumption Participants perceived the target group of each mock comment as intended.
- domain assumption Flagging behavior in the mock feed with explicit instructions reflects real-world online flagging.
- domain assumption The researcher classification of comments as abusive vs critical is valid; abusive items are genuinely hateful and critical items are not.
Cite this review
Pith. "Pith review of Ingroup bias is prevalent in user reports of hate and abuse online." pith.science (2026). https://pith.science/paper/OMDUSJ6Z
@misc{pith2026251004748,
author = {Pith},
title = {Pith review of: Ingroup bias is prevalent in user reports of hate and abuse online},
year = {2026},
howpublished = {\url{https://pith.science/paper/OMDUSJ6Z}},
note = {Machine review of arXiv:2510.04748}
}
read the original abstract
The prevalence of online hate and abuse is a pressing global problem. While tackling such societal harms is a priority for research across the social sciences, it is a difficult task, in part because of the magnitude of the problem. People's engagement with reporting mechanisms ('flagging') online is an increasingly important part of monitoring and addressing harmful content at scale. However, users may not flag content routinely enough, and when users do engage, they may be biased by group identity and political beliefs. Across five well-powered and pre-registered online experiments, we examine the extent of ingroup bias in people's flagging of hate and abuse in four different intergroup contexts: political affiliation, vaccination opinions, beliefs about climate change, and stance on abortion rights. Overall, participants reported abuse reliably, with approximately half of the abusive comments in each study reported. However, a pervasive ingroup bias was present whereby across studies, participants were between 17% and 63% more likely to flag abuse directed at the ingroup than at the outgroup. Our findings offer new insights into the nature of user flagging online, an understanding of which is crucial for enhancing user intervention against online hate and thus ensuring a safer online environment.
Figures
Reference graph
Works this paper leans on
-
[1]
Enock1, Helen Z
1 Social bias is prevalent in user reports of hate and abuse online Authors: Florence E. Enock1, Helen Z. Margetts1,2,3 and Jonathan Bright1 1 Public Policy Programme, The Alan Turing Institute, The British Library, 96 Euston Road, London. NW1 2DB. 2 Oxford Internet Institute, University of Oxford, Stephen A. Schwarzman Centre for the Humanities, Radcliff...
2016
-
[4]
and online hate can also provoke and justify violent attacks offline (Enock & Over, 2023; Leader Maynard & Benesch, 2016; Ofcom & Kick It Out, 2025; Siegel, 2020). As such, working to tackle online hate and abuse is a priority area for research across the social sciences, but it is a difficult task, in part because of the magnitude of the problem. Many in...
2023
-
[5]
The target of the comments was either the Democrats or the Republicans 10 depending on which between-subjects condition participants were in
and they were designed to include a various types of harmful language, including incitements of violence, dehumanizing slurs, and general abuse (Leader Maynard & Benesch, 2016). The target of the comments was either the Democrats or the Republicans 10 depending on which between-subjects condition participants were in. Examples of abusive comments were, ‘H...
2016
-
[7]
None of the interactions were significant (all ps > .05)
= 1.69, p =.194, ηp²= .005, showing pro-vaccination and anti-vaccination participants flagged to a similar extent overall. None of the interactions were significant (all ps > .05). Overall, comments directed at the ingroup were flagged to a greater extent than equivalent comments directed at the outgroup, both for abusive and critical speech, and this was no...
2022
-
[12]
https://doi.org/10.1038/s41558-022-01527-x Gillespie, T. (2018). Custodians of the Internet: Platforms, content moderation, and the hidden decisions that shape social media. Yale University Press. https://books.google.com/books?hl=en&lr=&id=cOJgDwAAQBAJ&oi=fnd&pg=PA1&dq=Custodians+of+the+Internet:+Platforms,+content+moderation,+and+the+hidden+decisions+th...
-
[20]
41 Iyengar, S., & Westwood, S. J. (2015). Fear and Loathing across Party Lines: New Evidence on Group Polarization. American Journal of Political Science, 59(3), 690–707. https://doi.org/10.1111/ajps.12152 Johansson, P., Enock, F., Hale, S., Vidgen, B., Bereskin, C., Margetts, H., & Bright, J. (2022). How can we combat online misinformation? A systematic ...
arXiv 2015
-
[41]
https://doi.org/10.1140/epjds/s13688-025-00556-8 Chang, R.-C., Rao, A., Zhong, Q., Wojcieszak, M., & Lerman, K. (2023). # RoeOverturned: Twitter Dataset on the Abortion Rights Controversy. Proceedings of the International AAAI Conference on Web and Social Media, 17, 997–1005. https://ojs.aaai.org/index.php/ICWSM/article/view/22207 Cikara, M., & Van Bavel,...
-
[308]
= 16.12, p < .001, ηp²= <.05. Pairwise comparisons showed that pro-abortion rights participants flagged ingroup-directed abuse to a greater extent than outgroup-directed abuse, (MIG = 63.1%, MOG = 56.8%, p < .001), but flagged critical comments to a similar extent for both target groups (MIG = 2.1%, MOG = 1.6%, p = .325). Anti-abortion rights participants fl...
2023
Show all 16 references
-
[312]
= 6.45, p =.012, ηp²=.020, suggesting that the effect of target group was dependent on political affiliation. Pairwise comparisons showed that Republicans flagged abuse directed at the ingroup to a greater extent than abuse directed at the outgroup (MIG = 54.9%, MOG = 41.8%, p =...
2020
-
[315]
None of the other interactions were significant (ps > .05)
= 6.73, p =.010, ηp² = .021, though these effects were not relevant for our 27 key research questions2. None of the other interactions were significant (ps > .05). Overall, comments directed at the ingroup were flagged to a greater extent than equivalent comments directed at the...
2023
-
[316]
None of the other interactions were significant (all ps > .05)
= 17.74, p < .001, ηp²= .053, which was not relevant to our key research questions and showed that while criticism was flagged to a similar extent by both groups, climate change believers flagged abuse to a greater extent than non-believers. None of the other interactions were s...
2015
- [851]
-
[2021]
(Meta, 2025). High levels of online hate and abuse are problematic for many reasons – exposure can cause severe harm to the psychological wellbeing of targets, with experiences linked to depression, anxiety, fear, low self-esteem and escalation of self-harm (Keipi et al., 2016...
2025
-
[2023]
efforts should focus not only on increasing overall engagement with these tools, but also on enhancing the accuracy and fairness of reports. Across all studies, while participants consistently flagged abusive comments to a greater extent than critical ones and comments directed...
2024 arXiv
-
[2025]
and in the US, survey research found that 41% of adults in America had directly experienced online harassment, with a quarter reporting personal experience with stalking, harassment, and physical threats (Vogels, 2021). Platform reports also highlight the scale of the problem ...
2021
-
[7811]
https://doi.org/10.1038/s41586-020-2281-1 Kahan, D. M. (2012). Ideology, motivated reasoning, and cognitive reflection: An experimental study. Judgment and Decision Making, 8, 407–424. Kahan, D. M., Hoffman, D. A., Braman, D., & Evans, D. (2012). They saw a protest: Cognitive i...
2012 doi
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.