{"id":"e3fb307c-83f7-4e69-9690-b4e789a90e3d","arxiv_id":"2506.12792","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review chapter synthesizes evidence that proactive design features can reduce harmful behavior and misinformation spread, while acknowledging the field is young.","lead":"This chapter argues that platform design choices can be made to encourage healthy interactions, and reviews field experiments showing that features like rule reminders, writing interstitials, and accuracy prompts reduce harmful behavior and misinformation sharing. It is a state-of-the-field overview for Trust and Safety practitioners and researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Effectiveness-conditioned evidence library makes the review's positive conclusions non-representative; the practical claim needs a systematic accounting of null results.","rationale":"The chapter is a review, not new empirical work, and it is largely careful: the Discussion explicitly warns that few interventions have been thoroughly tested, that most tested designs target misinformation, and that some prosocial designs backfire. Those caveats are to the authors' credit and make the chapter more trustworthy than a purely promotional piece. However, the selection mechanism described in the 'Selective review' section—drawing only from an evidence library containing designs with public evidence of effectiveness—means the reviewed set is not a representative sample of the intervention space. This is a structural issue rather than a matter of individual study accuracy. For the modest claim that 'some prosocial designs can reduce harm and misinformation,' the examples suffice; for the stronger practical claim that 'Prosocial Design is an effective approach,' the absence of a systematic census of null and mixed results leaves the strength of the conclusion unsupported. The reader's weakest assumption was about external generalizability; my concern is closely related but internal to the evidence base: even before asking whether effects generalize to other settings, we cannot tell whether the selected studies are representative of the full evidence base. A preregistered systematic review with inclusion independent of outcome direction would resolve this. This concern reinforces, rather than overturns, the reader's conditional acceptance: the chapter is worth publishing but should either qualify the abstract's claim or add a transparent statement that the review is illustrative and effectiveness-conditioned, not a systematic assessment.","tokens_in":14598,"tokens_out":6522,"duration_ms":74901,"concrete_test":"Run a preregistered systematic search of prosocial design interventions (e.g., norms reminders, comment interstitials, moderation explanations, accuracy prompts, prebunking, community notes) across peer-reviewed and preprint databases, unrestricted by outcome valence. Code every study for direction (positive/null/negative), outcome type (observed behavior vs. self-report), and setting (field vs. lab). Then compare the proportion of positive results and pooled effect sizes against the subset cited in this chapter. If the unrestricted set shows a substantially lower proportion of positive results or high heterogeneity, revise the abstract to 'some tested designs show promise' rather than 'can be an effective approach.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that tested prosocial design patterns are effective at reducing harmful behavior and misinformation. The review's evidence base is deliberately restricted to designs in the Prosocial Design Network's library that 'have been tested and for which there is public evidence that they are effective' (Selective review section). This makes the sample effectiveness-conditioned: null results and failed replications are excluded by construction. An existence claim ('some designs can work') survives that selection; the chapter's practical claim that Prosocial Design is an effective approach does not, because the reader cannot distinguish a robust approach from cherry-picked successes. The review cites backfire/null results only in passing (NewsGuard field experiment, implied truth effect, Hoes et al.; Dörr et al. in Discussion) without a systematic census, and publication bias in this literature is likely given that many included sources are preprints, internal industry reports, or unpublished (e.g., Bor et al. 2020; Chuai et al. 2024; Katsaros & Grüning, in preparation; Lin et al. 2024; Weijun et al. 2022). Consequently the headline conclusion overstates the strength of evidence unless the selected set is shown to be representative.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This chapter by Grüning and Kamin presents an overview of Prosocial Design, an approach to platform design and governance that explicitly uses design choices to foster healthy interactions, safety, wellbeing, and dignity. The authors articulate several principles (dignity of the user, non-neutral design, proactivity), situate Prosocial Design relative to Trust and Safety and related frameworks, and then—as the main contribution—offer a selective review of empirical research on prosocial design interventions aimed at reducing harmful behavior and mitigating misinformation, with additional examples for agency, wellbeing, discourse, and community. The review is organized by the proactive/interactive/reactive taxonomy introduced in the authors' earlier work (Grüning et al., 2024) and draws on the Prosocial Design Network's library of evidence-based design patterns. The chapter concludes that Prosocial Design is a nascent but promising field, while acknowledging the limited number of thoroughly tested interventions.","tokens_in":14790,"tokens_out":4984,"duration_ms":55724,"significance":"If the claims were fully supported, the chapter would provide a valuable bridge between Trust and Safety practice and behavioral design research, consolidating a dispersed literature of field experiments and quasi-experiments, including several high-quality RCTs (e.g., Matias, 2019; Katsaros et al., 2021; Pennycook & Rand, 2022). The proactive/interactive/reactive taxonomy is a useful organizing device for practitioners. However, because the review is explicitly selective and outcome-conditioned—the library is limited to designs 'for which there is public evidence that they are effective'—the evidence base cannot support the abstract's claim to 'demonstrate' that Prosocial Design is an effective approach. The chapter does include some null results (e.g., Aslett et al., 2022; Hoes et al., 2024; Dörr et al., 2025) and the Discussion appropriately notes that few interventions have been thoroughly tested, but the selection problem is not fully addressed. The paper ships no new data or code; its value is as a framing and review contribution, and its practical recommendations depend on the representativeness of the evidence.","major_comments":[{"comment":"The central claim that Prosocial Design 'can be an effective approach' (Abstract) rests on an evidence library that is effectiveness-conditioned by construction. The second paragraph of this section states that the library is 'limited to designs that have been tested and for which there is public evidence that they are effective in producing prosocial outcomes.' Because inclusion depends on the outcome, the review cannot distinguish a robust approach from cherry-picked successes. The Discussion (final pages) concedes that 'there are still few interventions that have been thoroughly tested' and that most tested designs target misinformation, but this caveat does not repair the selection problem. I recommend either (a) softening the central claim to an existence proof—'some tested prosocial design patterns show promise'—with explicit cautions about generalizing to untested platforms and populations, or (b) adding a systematic accounting of the underlying evidence base, including null results, failed replications, and unpublished studies, and a description of how the library is compiled and updated so that the reader can judge representativeness.","section":"Selective review of Prosocial Design research"},{"comment":"The authors' dual role as leaders of the Prosocial Design Network and as authors of the organizing taxonomy (Grüning et al., 2024) creates a potential conflict that is not disclosed in the review methodology. The selection of patterns from the network's own library, using the authors' own framework, may reflect confirmation bias. A reader cannot tell whether inclusion decisions and interpretations are independent of the authors' advocacy goals. I recommend adding a brief positionality statement and a description of the library's inclusion criteria (e.g., who screens designs, what counts as 'public evidence,' whether decisions are recorded) to strengthen transparency. Without this, the review reads more as a promotional document than as a balanced synthesis.","section":"Introduction and Selective review section"},{"comment":"The review treats designs across different evidential tiers as roughly comparable. For example, the Reddit observational study by Jhaver et al. (2019) is presented alongside randomized controlled trials such as Katsaros et al. (2021), and in-house or unpublished sources (Katsaros & Grüning, in preparation; Weijun et al., 2022, a Medium post; Matias et al., 2020, a project report; Lin et al., 2024, a preprint) are cited without noting their non-peer-reviewed status. This matters because the chapter's practical conclusion—that presenting 'evidence that a prosocial design is effective' can persuade platforms to adopt it—depends on the strength and quality of that evidence. I recommend adding a table or explicit provenance markers that distinguish preregistered field experiments, quasi-experiments, survey experiments, preprints, and internal industry reports, and include effect sizes and confidence intervals where available, so that practitioners can calibrate their confidence.","section":"Prosocial Designs to reduce harmful behavior; Prosocial Designs to reduce misinformation"}],"minor_comments":[{"comment":"In the paragraph beginning 'If Prosocial Design is gaining traction,' the text reads 'The Council on Technology and Social Cohesion... includes a heavy emphasis on Prosocial Design (cite).' The placeholder '(cite)' should be replaced with an actual reference to the cited policy brief or report.","section":"Discussion"},{"comment":"In the first paragraph under 'Interactive examples,' the typo 'intersitials' appears; it should be 'interstitials.'","section":"Prosocial Designs to reduce harmful behavior, Interactive examples"},{"comment":"The reference to Yadav & Garg is missing a publication year; the text cites it as 2023, but the list entry has no year. The year should be added to the reference list entry.","section":"References"},{"comment":"Several cited sources are non-peer-reviewed or in preparation: Katsaros & Grüning (in preparation), Weijun et al. (2022) on Medium, Matias et al. (2020) project report, Lin et al. (2024) PsyArXiv preprint, and Bor et al. (2020) PsyArXiv preprint. These should be flagged as such in the text (e.g., 'preprint' or 'internal industry report') at least at first mention, so readers can weigh the evidence appropriately.","section":"References and text throughout"},{"comment":"The reference list contains formatting artifacts such as 'V olume,' 'V iews,' and 'V olunteers' (apparent ligature or space errors). These should be cleaned for consistency.","section":"References"},{"comment":"The organization name is spelled inconsistently as 'New_ Public' and 'New_Public' in different places in the text and references; please standardize to the organization's preferred spelling.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a book chapter, and the review standard for a chapter might reasonably be lower than for a systematic review. However, the abstract's claim is broad, and the evidence selection is not transparent enough to support it without revision. The authors' leadership roles in the Prosocial Design Network and the reliance on their own library and taxonomy are not inherently disqualifying, but the conflict should be explicitly disclosed and the evidence base described more systematically. The issues I raise are local and fixable: reframe the central claim or add a census of null results, add a positionality statement, and clarify the provenance and quality of included studies. I do not see a load-bearing error in the specific study descriptions, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The chapter is a clean, readable overview of Prosocial Design aimed at Trust and Safety practitioners. The main contribution is the synthesis: it pulls together a scattered set of field experiments and organizes them by when they intervene (proactive/interactive/reactive). That taxonomy is not new—it comes from the authors' earlier framework paper—but applying it to a concrete set of designs is useful. The writing is clear and the chapter is honest about tensions (e.g., dignity vs. paternalism) and about the field's immaturity. It also mentions null results and backfire effects (NewsGuard, implied truth effect, Hoes et al.) even though they are not central to the story.\n\nThe soft spot is the evidence base. The review is explicitly drawn from the Prosocial Design Network's library, which includes only designs 'that have been tested and for which there is public evidence that they are effective.' That selection on effectiveness means the review supports an existence claim—some prosocial designs work—but not the stronger claim in the abstract that Prosocial Design can be an effective approach to reducing harm and misinformation. A reader cannot distinguish a robust approach from cherry-picked successes. The authors could fix this by adding a systematic accounting of null results, or at least by softening the abstract's 'demonstrate' to 'illustrate.' The reliance on preprints, industry blog posts, and in-preparation manuscripts (e.g., Katsaros & Grüning, in preparation) also weakens the 'public evidence' standard. Production issues, like an unfinished '(cite)' in the Discussion and some reference inconsistencies (Horta Ribeiro vs. Ribeiro), are easy to fix.\n\nOverall, this is a serviceable practitioner-oriented chapter, not a rigorous systematic review. A serious referee should engage it because the authors are honest and the topic matters; a good reviewer can help them calibrate the claims and clean up the references. I would not cite it in my own work, but I could see handing it to a T&S team as a primer.","headline":"Useful synthesis for practitioners, but the effectiveness-conditioned evidence library makes the general claim of efficacy too strong.","tokens_in":15285,"tokens_out":3297,"would_cite":false,"duration_ms":36265,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that prosocial design—deliberately shaping platform interfaces to encourage healthy interaction—can measurably reduce harmful behavior and the spread of misinformation, citing field experiments on norms reminders…","keywords":["prosocial design","trust and safety","platform governance","behavioral interventions","misinformation","online moderation","social norms","user wellbeing"],"falsifier":"A pre-registered, adequately powered replication that ran the same norms-reminder, interstitial, and accuracy-prompt designs across several unrelated platforms and found no average reduction in rule-breaking or misinformation sharing would directly undercut the paper's central claim. A more targeted version: a multi-platform field experiment testing accuracy prompts on non-English, non-Western social media populations that found zero or reversed effects on sharing false news.","tokens_in":14391,"feed_emoji":"🛡️","tokens_out":6596,"duration_ms":66234,"temperature":0.7,"pith_summary":"This review chapter argues that the design choices platforms make are never neutral: they steer user behavior, and they can be deliberately aimed at prosocial outcomes such as safety, dignity, and healthy interaction. Surveying field experiments and large-scale analyses, it contends that tested design patterns—norms reminders, comment interstitials, explanations for moderation, and accuracy prompts—reduce rule-breaking, harmful speech, and the spread of misinformation. The authors organize these interventions by when they act (before, during, or after an interaction) and draw on a curated library of evidence-based solutions. They also state plainly that the field is young, the review is selective, and only a few interventions have been thoroughly tested, with most evidence concentrated on misinformation. The sympathetic reader takes the paper as making the case that prosocial design is a viable complement to reactive trust-and-safety work, not as claiming a settled evidence base.","feed_headline":"Design nudges can cut harmful posts and misinformation","feed_subtitle":"A review of field experiments shows reminders, interstitials, and accuracy prompts steer online behavior toward healthier interactions.","key_machinery":"The mechanism that organizes the argument is the temporal classification of interventions into proactive, interactive, and reactive designs, defined by when an intervention touches user engagement. Proactive designs sit upstream of an action (e.g., a sticky-note reminder of community rules, an inoculation video, an accuracy prompt); interactive designs act at the moment of engagement (e.g., an interstitial asking a user to reconsider a hostile comment, a fact-check label, a community note); reactive designs operate after harm occurs (e.g., explaining why a post was removed, notifying users of consequences and appeals). This taxonomy carries the review because it lets the authors map each tested design pattern to a point in the user journey and infer that the same intervention logic can be reused across platforms.","core_discovery":"The paper's core claim is that platforms can reduce harmful behavior and misinformation not only by detecting and punishing bad actors but by designing the environment so that ordinary, good-faith users are nudged, reminded, or given the chance to reflect before they act. Against the backdrop of Trust and Safety's focus on abuse minimization, the authors present a working definition of Prosocial Design—design patterns, features, and processes that foster healthy interactions while ensuring safety, wellbeing, and dignity—and review experimental evidence for interventions placed at three temporal stages: proactive (norms reminders, inoculation, accuracy prompts), interactive (comment interstitials, fact-check labels, community notes), and reactive (explanations for content removal, due-process notifications). The reported effects include fewer rule-breaking posts, more civil comments, reduced recidivism after removal when explanations are given, and large drops in reshares of misleading posts once community notes are displayed. The authors are explicit that this is a selective review drawn from their own evidence library, and that the evidence base is still thin.","pith_inferences":["The temporal framework implies a testable priority: upstream interventions should generally be cheaper per harmful incident prevented than reactive ones, but the review does not present cost data, so that ordering remains an inference.","The evidence gap on reactive misinformation interventions suggests an unfilled design niche: platforms could test personalized, fact-based corrections delivered after a user shares false content, using the review's own suggested conditions of timeliness, credibility, and detail.","If the generalizability assumption holds, the same design patterns could be ported to adjacent spaces such as internal workplace collaboration tools or civic participation platforms, but the paper does not claim this."],"forward_implications":["Trust and safety teams can treat upstream design changes—norms reminders and accuracy prompts—as evidence-backed complements to reactive moderation.","Moderation actions that include explanations and due-process information are more likely to reduce repeat rule-breaking than silent removals.","Community notes, when displayed, can cut reshares of misleading posts by a large margin, though their overall impact is limited by how few posts receive notes and how late notes appear.","Designers should expect some prosocial interventions to backfire, so internal testing of any specific pattern is warranted before adoption."],"supporting_citations":[{"why":"Field experiment on the r/Science subreddit showing a sticky-note rules reminder increased rule-following and new-user posting; anchors the proactive norms-reminder pattern.","marker":"Matias (2019)"},{"why":"Field evidence that posting group norms at forum entry reduced comments reported for abuse; supports the proactive norms category.","marker":"Kim et al. (2022)"},{"why":"Facebook field experiment linking rule reminders after suspension to increased later rule adherence.","marker":"Tyler et al. (2021)"},{"why":"Twitter experiment showing a write-time interstitial reduced offensive tweets by 6%; central example of interactive design.","marker":"Katsaros et al. (2021)"},{"why":"OpenWeb interstitial test reporting a 12.5% increase in civil comments; another interactive example.","marker":"Goldberg et al. (2020)"},{"why":"Large-scale analysis of subreddits finding explanations for content removal reduced future removals, stronger for longer explanations.","marker":"Jhaver et al. (2019)"},{"why":"Survey experiments plus a Facebook/Twitter field study establishing accuracy prompts reduce misinformation sharing.","marker":"Pennycook & Rand (2022)"},{"why":"Field and lab studies showing inoculation videos build generalized resistance to misinformation.","marker":"Roozenbeek et al. (2022)"},{"why":"Quasi-experimental studies on X indicating community notes reduce reshares of misleading posts by over 50% once displayed.","marker":"Chuai et al. (2024)"},{"why":"Internal Twitter study reporting a 25-34% reduction in retweets when a community note is attached; corroborates the community-notes effect.","marker":"Wojcik et al. (2022)"}],"fun_headline_variants":["Design nudges reduce harmful posts, review finds","Prosocial design: shaping platforms for safer interactions","Prompts, reminders, and notes curb online harm","How platform design can stem misinformation and abuse","Designing for prosocial outcomes in trust and safety"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's conclusion depends on assuming that effects observed in a handful of platform-specific field experiments—conducted at particular times and on particular user populations—will generalize to other platforms, communities, and contexts at scale.","fun_headline_variants_meta":{"raw":{"variants":["Design nudges reduce harmful posts, review finds","Prosocial design: shaping platforms for safer interactions","Prompts, reminders, and notes curb online harm","How platform design can stem misinformation and abuse","Designing for prosocial outcomes in trust and safety"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2541,"prompt_tokens":879,"completion_tokens":1662,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":1590}},"tokens_in":495,"tokens_out":1662,"duration_ms":15713,"temperature":1.0,"reasoning_tokens":1590,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:41:49.975126+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A pre-registered, adequately powered replication that ran the same norms-reminder, interstitial, and accuracy-prompt designs across several unrelated platforms and found no average reduction in rule-breaking or misinformation sharing would directly undercut the paper's central claim. A more targeted version: a multi-platform field experiment testing accuracy prompts on non-English, non-Western social media populations that found zero or reversed effects on sharing false news.","supporting_citations":[],"review_version":1}