{"id":"b0a6a2e2-0b47-40b4-9ecf-61b9d89c66f1","arxiv_id":"2504.19158","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A restorative-justice-inspired guided sensemaking tool produced significantly higher self-reported sensemaking, guidance, support, and empowerment than an unstructured action-planning task in a 32-participant controlled experiment.","lead":"SnuggleSense is a guided web tool that helps people who have experienced online harm reflect on what happened, consider what they need, and build a step-by-step action plan, with suggestions drawn from similar survivors. A 32-person controlled experiment found that participants rated this structured process significantly higher for sensemaking, guidance, support, and empowerment than writing an action plan on their own.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'unstructured process' control is itself a structured writing task, so the headline claim outruns the experiment.","rationale":"The reader's weakest assumption correctly identifies the validity of the control condition as the load-bearing premise. The paper's own description of the Unstructured condition in Section 4.1 contradicts the label 'unstructured process of making sense of the harm' used in the abstract: participants are explicitly required to write a sequence of action items with stakeholders and actions, which is itself a structured format. This makes the experiment a comparison between two structured elicitation methods, not a comparison between structure and its absence. The within-subject design, the randomized order, and the reported statistics support the literal comparison to the writing task, so I would not reject the paper. But the generalizable claim about enhancing sensemaking relative to natural, unguided sensemaking is not established by the current evidence. The paper has genuine strengths: it reports a null result for agency honestly, discloses positionality, and provides a detailed description of the system and its restorative-justice rationale. These do not resolve the control-validity concern. The proposed additional naturalistic baseline would directly test whether the advantage persists when the comparison group is allowed to make sense of harm in whatever way they normally would. Since the reader's verdict is already CONDITIONAL and this concern is the same one the reader identified, the appropriate status judgment is UNCHANGED: conditional acceptance pending stronger outcome validation and a more representative baseline.","tokens_in":23925,"tokens_out":2514,"duration_ms":27592,"concrete_test":"Run a follow-up controlled experiment with the same population, recall timeframe, and outcome measures, but replace the current Unstructured condition with a naturalistic control: ask participants to 'make sense of the harm and develop whatever plan you would normally make, using whatever format you choose (writing, notes, talking aloud),' without prescribing stakeholders, action items, or a timeline. Keep the time limit, setting, and self-report scales identical. If SnuggleSense still shows a significant and comparable effect size on the sensemaking rating, the claim survives; if the difference shrinks below significance or below d ≈ 0.5, the reported effect is an artifact of the current control's imposed structure.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract—'SnuggleSense significantly enhances sensemaking compared to an unstructured process of making sense of the harm'—rests on the comparison described in Section 4.1. The control condition is not unstructured sensemaking: participants are instructed to 'develop an action plan by directly writing a sequence of action items, each specifying a stakeholder and their corresponding actions.' That prompt imposes the same stakeholder-action decomposition and linear output format that SnuggleSense scaffolds, while stripping away whatever natural process a survivor would actually use (free-form reflection, iterative revision, conversation, journaling, or no written output at all). The reported 1.31-point advantage in self-rated sensemaking (Section 5.1, Figure 3) may therefore reflect the artificial sparseness and rigidity of the writing-only baseline rather than a genuine improvement over natural sensemaking. This is not an internal statistical inconsistency; the data support the literal comparison to the writing task. But the construct validity of the headline claim depends on the control being representative of 'an unstructured process of making sense of the harm,' and the paper provides no evidence for that equivalence. The problem is compounded by the primary outcome being a single unvalidated self-report item, so the measured gap could be partly a format or demand effect. The authors' own limitations section acknowledges the experimental setting, but it does not address the representativeness of the control condition.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents SnuggleSense, a web-based system that supports online-harm survivors through a structured sensemaking process inspired by restorative justice: guided reflection on the harm, emotions, impacts, and needs; creation of stakeholder-and-action items on interactive sticky notes; recommendations from similar survivors; and chronological organization on a timeline. The authors evaluate the system in a within-subject, controlled experiment with 32 university students who had experienced online harm in the previous six months, comparing SnuggleSense to a control condition in which participants were asked to develop an action plan by directly writing a sequence of action items, each specifying a stakeholder and corresponding actions. The paper reports significantly higher self-report ratings for the structured condition on guidance, support, sensemaking, and empowerment, as well as significantly more stakeholders and action items in the plans produced with SnuggleSense. The qualitative data and participant quotes are used to illustrate the perceived benefits, and the authors discuss design implications for survivor-centered social computing and restorative justice.","tokens_in":24092,"tokens_out":6978,"duration_ms":68329,"significance":"If the central claim holds, the paper makes a useful contribution to the HCI/CSCW literature on survivor-centered responses to online harm. It offers a concrete, implemented system, a transparent similarity metric that is not fitted against outcome measures (§3.2.3), randomized order of conditions, a positionality statement, and a candid limitations section. The qualitative quotes substantiate that users found the structured process helpful. However, the headline quantitative claim is weakened by two construct-validity problems described below, and the experiment is best read as an exploratory demonstration of a design concept with a perceived-helpfulness evaluation, not as a rigorous causal test that SnuggleSense enhances sensemaking compared to genuinely unstructured sensemaking. With revised claims or additional control conditions, the work could be a valuable starting point for future restorative-justice-oriented tools.","major_comments":[{"comment":"The abstract's claim that SnuggleSense 'significantly enhances sensemaking compared to an unstructured process of making sense of the harm' outruns the experiment. The control condition is described as participants being asked to 'develop an action plan by directly writing a sequence of action items, each specifying a stakeholder and their corresponding actions' (§4.1). This is itself a structured task that imposes the same stakeholder-action decomposition and linear output format that SnuggleSense scaffolds, while omitting the system's other elements. The data therefore support the more limited claim that SnuggleSense is rated more helpful than a writing-only baseline; they do not compare against the natural, open-ended, iterative, or social processes that 'unstructured sensemaking' would typically involve. The paper should either add a genuinely unstructured control condition (e.g., free-form reflection or journaling without an imposed schema) or revise the abstract, introduction, and discussion (§6.1.1) to describe the comparison to a text-only action-planning task.","section":"Abstract; §4.1"},{"comment":"The central outcome—'assistance in sensemaking'—is a single-item, self-authored rating scale with no demonstrated reliability or validity. There is no multi-item scale, no prior validation, and no convergent or objective measure of sensemaking. Single-item self-reports are particularly weak for an abstract construct, because the rating may reflect the perceived structure of the interface rather than an improved psychological process. Demand characteristics are also salient: the five survey categories (guidance, support, agency, sensemaking, empowerment) explicitly mirror the system's design goals, the researcher was present via Zoom during the study (§4.5), and the within-subject design invites direct comparison between conditions. I recommend either (a) reporting a validated or at least multi-item sensemaking scale, (b) adding a behavioral or objective measure, or (c) explicitly reframing the conclusions to 'perceived improvement in sensemaking' and discussing these measurement threats in the limitations section.","section":"§4.2.1; Figure 3"},{"comment":"The plan-complexity results (more stakeholders and more action items in the Structured condition; §5.4.1) are treated in §6.1.1 as evidence that the system 'enhances the knowledge survivors need to address the harm.' This inference is questionable because 42.22% of the action items in the Structured condition were adopted directly from the system's suggestions, and the reflection questions explicitly prompt users to consider multiple stakeholder types. More numerous actions and stakeholders may therefore reflect the scaffolding and the suggestion feature rather than a cognitive gain in sensemaking. The discussion should either test the incremental effect of the suggestion feature (e.g., by including a structured condition without recommendations) or soften the causal language about the content-level differences.","section":"§5.4; §6.1.1"}],"minor_comments":[{"comment":"The ACM CCS Concepts and Additional Key Words fields contain placeholder text unrelated to the paper (e.g., 'Computer systems organization→ Embedded systems; Redundancy; Robotics; Networks→ Network reliability' and 'datasets, neural networks, gaze detection, text tagging'). These should be corrected before publication.","section":"Page 1 (CCS Concepts and Keywords)"},{"comment":"The notation 'I_{ik,jk}' is unclear; the indicator function I should be defined explicitly as a function of whether survivors i and j both selected (or both did not select) a particular option. Also, the text 'if both survivors either selected or did not select an option, 1/n is added' should specify 'for that option and question' to make the summation in the equation concrete.","section":"§3.2.3, Eq. (1)"},{"comment":"Some entries are formatted inconsistently (e.g., '72%' versus '87.50%'), and the percentages per category exceed 100% because participants can appear in multiple categories; the table captions should state that percentages are per participant and are not mutually exclusive.","section":"§5.4.2, Tables 4 and 5"},{"comment":"The text contains a typo: 'SunggleSense' should read 'SnuggleSense' in the final paragraph of this subsection.","section":"§6.3.4"},{"comment":"Please clarify whether the full 15-minute allotment applied to the entire SnuggleSense process (including the reflection questions and browsing of recommendations) or only to the action-item creation phase; if the former, differing time pressure between conditions should be noted as a limitation, and if the latter, the condition comparison of effort is not balanced.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely topic, and the authors are transparent about many methodological choices. The central issue is that the abstract and discussion make a claim about 'unstructured sensemaking' that the control condition cannot support; this is fixable by either redesigning the control or recalibrating the claims. The single-item outcome measure is a further load-bearing weakness. I would encourage the editors to invite a revision rather than reject, given the strength of the qualitative material and the potential contribution of the design concept. I also noticed that the CCS Concepts and reference [43] appear to be unfinished placeholders, which suggests the manuscript is not yet in a publishable form; these should be fixed in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the system is a real contribution: SnuggleSense integrates guided reflection, peer-similarity recommendations via a defined Sij metric, and sticky-note timeline planning in a way I haven't seen in the cited prior tools. The experiment is carefully run, the stats are transparent, the null result on agency is reported honestly, and the qualitative excerpts give a credible picture of why survivors found the tool helpful. This is serious work, not a toy. Second, the headline claim outruns the experiment. The abstract says SnuggleSense enhances sensemaking \"compared to an unstructured process,\" but the control condition is not unstructured sensemaking. Participants were asked to write a sequence of action items, each specifying a stakeholder and their corresponding actions. That prompt imposes the same stakeholder-action decomposition and linear output format that SnuggleSense scaffolds, minus the guidance. So the comparison is really SnuggleSense versus a stripped-down writing task, not versus how survivors would naturally make sense of harm. The 1.31-point advantage in self-rated sensemaking may partly reflect the artificial sparseness of the baseline. The paper provides no evidence that the writing task matches natural unguided sensemaking. That is a construct validity problem, and the stress-test note lands on it correctly. The soft spots are otherwise proportionate. The primary outcome is a single-item self-report with no validation; the within-subject design and lack of multiple-comparison correction add noise; and the increased action-plan complexity is partly driven by in-system suggestions. None of this is fatal to the literal comparison — the data support what was actually tested — but it means the general claim about survivor empowerment needs stronger measurement and a more naturalistic baseline. I also want to give credit where it's due: the positionality statement is thoughtful, the limitations section acknowledges the experimental setting, and the authors do not oversell the null agency result. The citation pattern looks fine; the self-citations to Xiao et al. are directly relevant prior work. Who should read this? CSCW and HCI researchers working on online harm, restorative justice, or survivor-centered tools. It is also a useful teaching case for evaluation design in social-computing systems: it shows how a well-intentioned control condition can quietly encode the intervention's core assumptions. My recommendation: send it to peer review. The system is novel, the study is honestly reported, and the central flaw is addressable — either by reframing the claims to match the actual comparison, or better, by adding a more naturalistic baseline condition. It deserves serious referee time, with the expectation of revision.","headline":"A genuinely new support system for online harm survivors with a solid but overclaimed evaluation: the 'unstructured' control is itself a structured writing task, so the size of the reported effect should be read with caution.","tokens_in":24708,"tokens_out":1652,"would_cite":true,"duration_ms":17920,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A structured, restorative-justice-inspired process helps online harm survivors make sense of the harm, a controlled experiment finds.","keywords":["online harm","sensemaking","restorative justice","survivor empowerment","content moderation","action plans","peer recommendations","social computing"],"falsifier":"A study that lets survivors make sense of harm with no imposed format, such as free writing, talking with a friend, or browsing support threads, and compares that against SnuggleSense on the same self-report scales would settle whether the structured process itself, rather than the sparse control prompt, drives the gains.","tokens_in":23656,"feed_emoji":"🕊️","tokens_out":6035,"duration_ms":58768,"temperature":0.7,"pith_summary":"The paper introduces SnuggleSense, a web-based system that guides people who have experienced online interpersonal harm through reflective questions, personalized suggestions drawn from similar survivors, and interactive sticky-note timelines for building an action plan. It claims that this structured sensemaking process significantly improves survivors' self-reported ability to make sense of the harm when compared with simply writing out an action plan on their own. The motivation is that content moderation is perpetrator-centered and leaves survivors' needs for understanding, emotional support, validation, and agency largely unmet. If the claim is right, a restorative-justice-inspired tool could complement moderation by helping survivors process what happened and see a wider set of people and actions available for responding to it.","feed_headline":"Guided tool beats unstructured writing for online harm survivors","feed_subtitle":"Survivors rated a guided restorative-justice tool higher on sensemaking, support, and empowerment than free-form writing.","key_machinery":"The central mechanism is SnuggleSense's structured sensemaking pipeline, modeled on restorative justice practices of pre-conference and circles: a guided reflection sequence asks survivors to describe the harm, their feelings, the impacts, and their needs; a similarity score over multiple-choice harm descriptors matches each survivor to three similar prior users whose action items are offered as recommendations; and survivors assemble stakeholder-action items on a chronological sticky-note timeline, with an option to contribute their own plan back to the shared repository. This pipeline is what carries the argument, because it converts the abstract idea of survivor-centered sensemaking into a concrete, repeatable procedure that can be compared against unguided planning.","core_discovery":"In a within-subject experiment with 32 recent survivors, the paper found that participants rated the structured SnuggleSense process significantly higher than an unstructured writing task on sensemaking (M = 6.06 versus 4.75, p < .001), guidance, support, and empowerment, while ratings of agency did not differ significantly between the two conditions. Action plans produced with SnuggleSense contained more distinct stakeholders (M = 4.34 versus 3.16) and more action items (M = 6.25 versus 4.50), and the recommended items from similar survivors were adopted by 81.25% of participants. The paper interprets these results as evidence that a structured, restorative-justice-inspired process expands survivors' awareness of available resources and community support, shifting their plans away from self-focused coping such as ignoring, blocking, or deleting and toward involving family, friends, online community members, and even asking offenders to explain their motivations.","pith_inferences":["Beyond the paper, a natural test is whether the larger action plans produced under SnuggleSense translate into actual steps taken over subsequent weeks; the paper itself flags longitudinal follow-up as future work.","The similarity engine, which currently matches survivors on four multiple-choice harm descriptors, could be extended with needs-based or identity-based similarity measures to serve more diverse survivor populations, a possibility the paper leaves open.","The same guided pipeline could be adapted for adjacent goals, such as evidence documentation, formal reporting, or community education, by changing the reflection prompts and the targets of the recommendations."],"forward_implications":["If the result holds, survivor-facing platforms can treat structured reflection plus peer suggestions as a viable complement to reporting and content moderation.","Users in the structured condition produced action plans with significantly more stakeholders and more action items, suggesting the system widens the solution space survivors consider.","Because agency ratings did not differ between conditions, the design suggests guidance can be added without making survivors feel they have lost control over their own plan.","The observed shift from self-directed coping toward community involvement and requests for explanation aligns survivors' action plans more closely with restorative justice ideals of healing and restoration."],"supporting_citations":[{"why":"Supplies the definition of sensemaking that the experiment's main outcome measure targets.","marker":"[68]"},{"why":"Provides the pre-conference-style action-plan structure of stakeholders, actions, and chronological ordering that SnuggleSense implements.","marker":"[72]"},{"why":"Defines restorative justice and the pre-conference and circle practices from which the system draws inspiration.","marker":"[76]"},{"why":"Frames the contrast between content-moderation questions and restorative-justice questions that shapes the system's reflection prompts.","marker":"[50]"},{"why":"Documents survivors' unmet needs beyond content moderation, the problem the system is designed to address.","marker":"[54]"},{"why":"Grounds the paper's claim that survivor empowerment is central to restorative justice.","marker":"[1]"},{"why":"An earlier survivor-support platform that mobilizes online community members, which SnuggleSense extends toward sensemaking.","marker":"[7]"},{"why":"An earlier friend-sourced moderation tool representing the community-based support approach that SnuggleSense positions itself alongside.","marker":"[39]"}],"fun_headline_variants":["Structured process boosts survivors' sensemaking and support","Guided sensemaking tool empowers online harm survivors","Structured restorative process yields richer survivor action plans","Survivors gain more support from guided sensemaking tool","Guided sensemaking outshines free-form writing for survivors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that the control task, writing an action plan directly as a list of stakeholder-action items, is a fair stand-in for how survivors naturally make sense of harm; if unguided sensemaking is usually more free-form, social, or iterative, the measured advantage may be partly an artifact of how bare the control prompt is.","fun_headline_variants_meta":{"raw":{"variants":["Structured process boosts survivors' sensemaking and support","Guided sensemaking tool empowers online harm survivors","Structured restorative process yields richer survivor action plans","Survivors gain more support from guided sensemaking tool","Guided sensemaking outshines free-form writing for survivors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000789,"raw_usage":{"total_tokens":3444,"prompt_tokens":877,"completion_tokens":2567,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":2486}},"tokens_in":493,"tokens_out":2567,"duration_ms":17812,"temperature":1.0,"reasoning_tokens":2486,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T06:01:21.965471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A study that lets survivors make sense of harm with no imposed format, such as free writing, talking with a friend, or browsing support threads, and compares that against SnuggleSense on the same self-report scales would settle whether the structured process itself, rather than the sparse control prompt, drives the gains.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the definition of sensemaking that the experiment's main outcome measure targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the pre-conference-style action-plan structure of stakeholders, actions, and chronological ordering that SnuggleSense implements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines restorative justice and the pre-conference and circle practices from which the system draws inspiration."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames the contrast between content-moderation questions and restorative-justice questions that shapes the system's reflection prompts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the paper's claim that survivor empowerment is central to restorative justice."}],"review_version":1}