REVIEW 3 major objections 6 minor 1 cited by
Prosocial Design in Trust and Safety
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This review argues that prosocial design—deliberately shaping platform interfaces to encourage healthy interaction—can measurably reduce harmful behavior and the spread of misinformation, citing field experiments on norms reminders…
desk verdict Useful synthesis for practitioners, but the effectiveness-conditioned evidence library makes the general claim of efficacy too strong. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that organizes the argument is the temporal classification of interventions into proactive, interactive, and reactive designs, defined by when an intervention touches user engagement. Proactive designs sit upstream of an action (e.g., a sticky-note reminder of community rules, an inoculation video, an accuracy prompt); interactive designs act at the moment of engagement (e.g., an interstitial asking a user to reconsider a hostile comment, a fact-check label, a community note); reactive designs operate after harm occurs (e.g., explaining why a post was removed, notifying users of consequences and appeals). This taxonomy carries the review because it lets the authors map each tested design pattern to a point in the user journey and infer that the same intervention logic can be reused across platforms.
What would settle it
A pre-registered, adequately powered replication that ran the same norms-reminder, interstitial, and accuracy-prompt designs across several unrelated platforms and found no average reduction in rule-breaking or misinformation sharing would directly undercut the paper's central claim. A more targeted version: a multi-platform field experiment testing accuracy prompts on non-English, non-Western social media populations that found zero or reversed effects on sharing false news.
Extended reading notes
Core claim
The paper's core claim is that platforms can reduce harmful behavior and misinformation not only by detecting and punishing bad actors but by designing the environment so that ordinary, good-faith users are nudged, reminded, or given the chance to reflect before they act. Against the backdrop of Trust and Safety's focus on abuse minimization, the authors present a working definition of Prosocial Design—design patterns, features, and processes that foster healthy interactions while ensuring safety, wellbeing, and dignity—and review experimental evidence for interventions placed at three temporal stages: proactive (norms reminders, inoculation, accuracy prompts), interactive (comment interstitials, fact-check labels, community notes), and reactive (explanations for content removal, due-process notifications). The reported effects include fewer rule-breaking posts, more civil comments, reduced recidivism after removal when explanations are given, and large drops in reshares of misleading posts once community notes are displayed. The authors are explicit that this is a selective review drawn from their own evidence library, and that the evidence base is still thin.
Load-bearing premise
The review's conclusion depends on assuming that effects observed in a handful of platform-specific field experiments—conducted at particular times and on particular user populations—will generalize to other platforms, communities, and contexts at scale.
Editorial extensions
If this is right
- Trust and safety teams can treat upstream design changes—norms reminders and accuracy prompts—as evidence-backed complements to reactive moderation.
- Moderation actions that include explanations and due-process information are more likely to reduce repeat rule-breaking than silent removals.
- Community notes, when displayed, can cut reshares of misleading posts by a large margin, though their overall impact is limited by how few posts receive notes and how late notes appear.
- Designers should expect some prosocial interventions to backfire, so internal testing of any specific pattern is warranted before adoption.
Reading between the lines
- The temporal framework implies a testable priority: upstream interventions should generally be cheaper per harmful incident prevented than reactive ones, but the review does not present cost data, so that ordering remains an inference.
- The evidence gap on reactive misinformation interventions suggests an unfilled design niche: platforms could test personalized, fact-based corrections delivered after a user shares false content, using the review's own suggested conditions of timeliness, credibility, and detail.
- If the generalizability assumption holds, the same design patterns could be ported to adjacent spaces such as internal workplace collaboration tools or civic participation platforms, but the paper does not claim this.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This chapter by Grüning and Kamin presents an overview of Prosocial Design, an approach to platform design and governance that explicitly uses design choices to foster healthy interactions, safety, wellbeing, and dignity. The authors articulate several principles (dignity of the user, non-neutral design, proactivity), situate Prosocial Design relative to Trust and Safety and related frameworks, and then—as the main contribution—offer a selective review of empirical research on prosocial design interventions aimed at reducing harmful behavior and mitigating misinformation, with additional examples for agency, wellbeing, discourse, and community. The review is organized by the proactive/interactive/reactive taxonomy introduced in the authors' earlier work (Grüning et al., 2024) and draws on the Prosocial Design Network's library of evidence-based design patterns. The chapter concludes that Prosocial Design is a nascent but promising field, while acknowledging the limited number of thoroughly tested interventions.
Significance. If the claims were fully supported, the chapter would provide a valuable bridge between Trust and Safety practice and behavioral design research, consolidating a dispersed literature of field experiments and quasi-experiments, including several high-quality RCTs (e.g., Matias, 2019; Katsaros et al., 2021; Pennycook & Rand, 2022). The proactive/interactive/reactive taxonomy is a useful organizing device for practitioners. However, because the review is explicitly selective and outcome-conditioned—the library is limited to designs 'for which there is public evidence that they are effective'—the evidence base cannot support the abstract's claim to 'demonstrate' that Prosocial Design is an effective approach. The chapter does include some null results (e.g., Aslett et al., 2022; Hoes et al., 2024; Dörr et al., 2025) and the Discussion appropriately notes that few interventions have been thoroughly tested, but the selection problem is not fully addressed. The paper ships no new data or code; its value is as a framing and review contribution, and its practical recommendations depend on the representativeness of the evidence.
major comments (3)
- [Selective review of Prosocial Design research] The central claim that Prosocial Design 'can be an effective approach' (Abstract) rests on an evidence library that is effectiveness-conditioned by construction. The second paragraph of this section states that the library is 'limited to designs that have been tested and for which there is public evidence that they are effective in producing prosocial outcomes.' Because inclusion depends on the outcome, the review cannot distinguish a robust approach from cherry-picked successes. The Discussion (final pages) concedes that 'there are still few interventions that have been thoroughly tested' and that most tested designs target misinformation, but this caveat does not repair the selection problem. I recommend either (a) softening the central claim to an existence proof—'some tested prosocial design patterns show promise'—with explicit cautions about generalizing to untested platforms and populations, or (b) adding a systematic accounting of the underlying evidence base, including null results, failed replications, and unpublished studies, and a description of how the library is compiled and updated so that the reader can judge representativeness.
- [Introduction and Selective review section] The authors' dual role as leaders of the Prosocial Design Network and as authors of the organizing taxonomy (Grüning et al., 2024) creates a potential conflict that is not disclosed in the review methodology. The selection of patterns from the network's own library, using the authors' own framework, may reflect confirmation bias. A reader cannot tell whether inclusion decisions and interpretations are independent of the authors' advocacy goals. I recommend adding a brief positionality statement and a description of the library's inclusion criteria (e.g., who screens designs, what counts as 'public evidence,' whether decisions are recorded) to strengthen transparency. Without this, the review reads more as a promotional document than as a balanced synthesis.
- [Prosocial Designs to reduce harmful behavior; Prosocial Designs to reduce misinformation] The review treats designs across different evidential tiers as roughly comparable. For example, the Reddit observational study by Jhaver et al. (2019) is presented alongside randomized controlled trials such as Katsaros et al. (2021), and in-house or unpublished sources (Katsaros & Grüning, in preparation; Weijun et al., 2022, a Medium post; Matias et al., 2020, a project report; Lin et al., 2024, a preprint) are cited without noting their non-peer-reviewed status. This matters because the chapter's practical conclusion—that presenting 'evidence that a prosocial design is effective' can persuade platforms to adopt it—depends on the strength and quality of that evidence. I recommend adding a table or explicit provenance markers that distinguish preregistered field experiments, quasi-experiments, survey experiments, preprints, and internal industry reports, and include effect sizes and confidence intervals where available, so that practitioners can calibrate their confidence.
minor comments (6)
- [Discussion] In the paragraph beginning 'If Prosocial Design is gaining traction,' the text reads 'The Council on Technology and Social Cohesion... includes a heavy emphasis on Prosocial Design (cite).' The placeholder '(cite)' should be replaced with an actual reference to the cited policy brief or report.
- [Prosocial Designs to reduce harmful behavior, Interactive examples] In the first paragraph under 'Interactive examples,' the typo 'intersitials' appears; it should be 'interstitials.'
- [References] The reference to Yadav & Garg is missing a publication year; the text cites it as 2023, but the list entry has no year. The year should be added to the reference list entry.
- [References and text throughout] Several cited sources are non-peer-reviewed or in preparation: Katsaros & Grüning (in preparation), Weijun et al. (2022) on Medium, Matias et al. (2020) project report, Lin et al. (2024) PsyArXiv preprint, and Bor et al. (2020) PsyArXiv preprint. These should be flagged as such in the text (e.g., 'preprint' or 'internal industry report') at least at first mention, so readers can weigh the evidence appropriately.
- [References] The reference list contains formatting artifacts such as 'V olume,' 'V iews,' and 'V olunteers' (apparent ligature or space errors). These should be cleaned for consistency.
- [Throughout] The organization name is spelled inconsistently as 'New_ Public' and 'New_Public' in different places in the text and references; please standardize to the organization's preferred spelling.
Circularity Check
No equation-level circularity; central empirical claims rest on external field experiments, but the review's evidence pool is effectiveness-conditioned, so the demonstration of effectiveness is partly guaranteed by inclusion criteria.
-
other
[Selective review of Prosocial Design research (first paragraph)]
"The design patterns we include are drawn from Prosocial Design Network's library of evidence-based design solutions. This library does not contain all existing prosocial design patterns, let alone all imaginable patterns; it is limited to designs that have been tested and for which there is public evidence that they are effective in producing prosocial outcomes."
The chapter's stated main contribution is to 'review relevant research to demonstrate how Prosocial Design can be an effective approach to reducing rule-breaking and other harmful behavior and how it can help to stem the spread of harmful misinformation.' The review sample is explicitly restricted to designs for which public evidence of effectiveness already exists. Therefore the positive conclusion is entailed by the sampling rule rather than by an independent, representative assessment of prosocial design interventions. The existence claim ('some well-tested designs work') survives, but the stronger, general 'effective approach' claim is not tested by the selected cases.
full rationale
This is a narrative review, not a quantitative derivation, so the standard fitted-parameter and equation-level circularity patterns do not apply. The central examples—norms reminders (Matias 2019), comment interstitials (Goldberg 2020; Katsaros et al. 2021), accuracy prompts (Pennycook & Rand 2022; Lin et al. 2024), community notes (Chuai et al. 2023/2024; Renault et al. 2023), and moderation explanations (Jhaver et al. 2019)—are external empirical studies with independent content. The chapter's own definition of Prosocial Design is stipulated, not derived from the reviewed evidence, and the evidence is used to illustrate, not to prove, the definition. The main concern is that the pool of reviewed patterns comes from the Prosocial Design Network's library and is restricted to designs that already have public evidence of effectiveness; consequently the review cannot detect null or backfiring interventions, and the breadth of the practical claim is weaker than the selection procedure suggests. The authors acknowledge the selectivity and explicitly state that 'there are still few interventions that have been thoroughly tested to have a positive impact on desired prosocial outcomes.' Because the review is transparent, makes only an existence-oriented claim, and rests on external benchmark studies, the circularity burden is low: score 2.
Assumptions & free parameters
assumptions (3)
- domain assumption Design is not neutral; design choices influence user behavior.
- domain assumption A subset of rule-breaking users act in good faith and can be reformed.
- domain assumption Cited primary studies are reliable and accurately reported.
Cite this review
Pith. "Pith review of Prosocial Design in Trust and Safety." pith.science (2026). https://pith.science/paper/VTCRPNS2
@misc{pith2026250612792,
author = {Pith},
title = {Pith review of: Prosocial Design in Trust and Safety},
year = {2026},
howpublished = {\url{https://pith.science/paper/VTCRPNS2}},
note = {Machine review of arXiv:2506.12792}
}
read the original abstract
This chapter presents an overview of Prosocial Design, an approach to platform design and governance that recognizes design choices influence behavior and that those choices can or should be made toward supporting healthy interactions and other prosocial outcomes. The authors discuss several core principles of Prosocial Design and its relationship to Trust and Safety and other related fields. As a primary contribution, the chapter reviews relevant research to demonstrate how Prosocial Design can be an effective approach to reducing rule-breaking and other harmful behavior and how it can help to stem the spread of harmful misinformation. Prosocial Design is a nascent and evolving field and research is still limited. The authors hope this chapter will not only inspire more research and the adoption of a prosocial design approach, but that it will also provoke discussion about the principles of Prosocial Design and its potential to support Trust and Safety.
Forward citations
Cited by 1 Pith paper
-
Filtering Offensive Content Changes Its Visibility but Not User Behavior: Two Randomized Controlled Trials with 200,000 Users on Nextdoor
Filtering offensive content on Nextdoor sharply reduced its views but had no meaningful effect on user behavior in two RCTs with 200,000 users.
Reference graph
Works this paper leans on
-
[1]
Argyle, L. P., Bail, C. A., Busby, E. C., Gubler, J. R., Howe, T., Rytting, C., Sorensen, T., & Wingate, D. (2023). Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale. Proceedings of the National Academy of Sciences, 120(41), e2311627120. https://doi.org/10.1073/pnas.2311627120 Argyle, L. P., Bus...
-
[4]
from https://www.everythinginmoderation.co/prosocial-design/ Iyer, R. (2022). Content moderation is a dead end. Designing Tomorrow. Retrieved (May 29,
work page 2022
-
[5]
Adaptable Commitment Interfaces
from https://psychoftech.substack.com/p/content-moderation-is-a-dead-end 23 Lin, H., Garro, H., Wernerfelt, N., Shore, J. C., Hughes, A., Deisenroth, D., Barr, N., Berinsky, A. J., Eckles, D., Pennycook, G., & Rand, D. G. (2024). Reducing misinformation sharing at scale using digital accuracy prompt ads. PsyArXiv. https://doi.org/10.31234/osf.io/u8anb Jha...
arXiv 2024
-
[6]
https://doi.org/10.1038/s44271-023-00052-7 Grüning, D. J., Kamin, J., Saltz, E., Acosta, T., DiFranzo, D., Goldberg, B., Leavitt, A., Menczer, F., Musgrave, T., Wang, Y ., & Wojcieszak, M. (2025). Independently testing prosocial interventions: Methods and recommendations from 31 researchers. Annals of the New York Academy of Sciences. https://doi.org/10.3...
arXiv 2025
-
[7]
https://newpublic.org/purpose/core-beliefs New_Public (2022, March)
from https://kgi.georgetown.edu/wp-content/uploads/2025/02/Better-Feeds_-Algorithms-That- Put-People-First.pdf New_ Public (n.d.) Purpose | New Public. https://newpublic.org/purpose/core-beliefs New_Public (2022, March). The Signals: The qualities of flourishing digital spaces. New_Public. Retrieved (May 28,
work page 2025
-
[8]
from https://docs.google.com/presentation/d/ 1UAsy8ZlCoRwgwOLQNblvuV5gVdPBhS2SmZVtXWT20L4/edit? slide=id.g9c2b1f0ede_1_10#slide=id.g9c2b1f0ede_1_10 Nyhan, B., & Reifler, J. (2010). When corrections fail: The persistence of political misperceptions. Political Behavior, 32(2), 303-330. 25 Offer-Westort, M., Rosenzweig, L. R., & Athey, S. (2024). Battling th...
arXiv 2010
-
[9]
Post Guidance for Online Communities
from https://techandsocialcohesion.org/wp-content/uploads/2025/04/EUI-Policy-Brief- Advancing-Prosocial-Tech-Design-and-shaping-the-EU-platform-design-governance.pdf Ribeiro, M. H., West, R., Lewis, R., & Kairam, S. (2024). Post Guidance for Online Communities. arXiv. https://arxiv.org/abs/2411.16814 Roozenbeek, J., & Van der Linden, S. (2019). Fake news ...
work page Pith review arXiv 2024
-
[10]
Schmidt, A. T., & Engelen, B. (2020). The ethics of nudging: An overview. Philosophy compass, 15(4), e12658. Simon, G. (2020). OpenWeb tests the impact of “nudges” in online discussions. OpenWeb Blog. https://www.openweb.com/blog/openweb-improves-community-health-with-real-time- feedback-powered-by-jigsaws-perspective-api Srinivasan, K. B., Danescu-Nicule...
Show all 10 references
-
[2025]
from https://www.everythinginmoderation.co/trust- safety-values-action/ Hunsberger, A. (2025). Is it prosocial design’s time to shine?. Everything in Moderation. Retrieved (May 29,
2025
-
[2333]
B., Bueno, N
https://doi.org/10.1038/s41467-022-30073-5 Pereira, F. B., Bueno, N. S., Nunes, F., & Pavão, N. (2024). Inoculation reduces misinformation: experimental evidence from multidimensional interventions in brazil. Journal of Experimental Political Science, 11(3), 239-250. https://d...
2024
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.