{"id":"392659d4-1776-49b1-9b69-03460d815ac2","arxiv_id":"2506.13739","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Social robots for mental wellbeing can work as coaches or via video, and should be designed with clinicians and studied long term.","lead":"This perspective paper distills six lessons from the authors' studies of social robots for mental wellbeing, including that robots need not be companions, virtual delivery works, and adaptation is optional. It argues robots should support human care rather than replace therapists.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Insights generalize from small same-group samples; the paper's own Discussion admits this, so a conditional verdict is appropriate.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern I see: the six insights are presented as transferable design lessons, yet they rest largely on the authors' own small-sample, non-clinical studies without independent replication. This is the strongest point of vulnerability for the paper's central claim, because the 'grounded in evidence' assertion in the abstract is only as strong as the evidence base, and that base is mostly self-referential. The paper partially inoculates itself by disclaiming comprehensiveness and by acknowledging in the Discussion that current studies limit generalizability and replicability. However, the insights themselves are still stated in a general, prescriptive voice, so the acknowledged limitation does not fully close the gap. I do not think this concern warrants rejection or unverdicting: the paper is a perspective piece, not a systematic review, and the central framing—robots as supportive tools that complement human practitioners—is a plausible and responsible stance. A conditional acceptance, as the reader recommends, is the right level: the recommendations are reasonable heuristics but should be treated as hypotheses pending independent replication. My proposed concrete test targets the most load-bearing empirical pillar by asking whether a core longitudinal effect replicates under more rigorous, controlled, multi-site conditions. If that replication fails, Insights 2, 5, and 6 weaken substantially; if it succeeds, the paper's credibility improves accordingly. Since this is exactly the kind of condition the reader's verdict already encodes, the verdict should remain UNCHANGED.","tokens_in":22261,"tokens_out":3845,"duration_ms":44032,"concrete_test":"Pre-register and run a multi-site replication of the core paradigm behind Insights 2, 5, and 6: the longitudinal robot self-disclosure intervention (Laban et al. [20], §2.2 and §2.6). Use an independent lab, N ≥ 100, diverse non-clinical participants, random assignment to robot, chatbot, or waiting-list control, and validated wellbeing measures at baseline, post-intervention, and follow-up. If the robot condition does not outperform control on self-disclosure depth or wellbeing outcomes, the transferability of these insights is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that robots are best understood as supportive tools 'grounded in evidence'—depends on the six insights in §2.2–2.6 being transferable beyond the specific studies that generated them. The evidentiary base is dominated by the authors' own studies with small, non-clinical samples: e.g., n=17 in Spitale et al. [81] for the VITA longitudinal coaching study, n=41 in Abbasi et al. [34] for child wellbeing assessment, and the long-term self-disclosure studies [20, 17] with modest sample sizes. Many of these studies lack control conditions, rely heavily on self-report, and were run by the same research group, leaving novelty effects, demand characteristics, and experimenter expectations as plausible rival explanations. The paper itself flags this in the Discussion: 'Current studies often rely on short-term measures and non-clinical populations, limiting the generalizability of findings, as well as their replicability.' Because the insights are phrased as general design lessons rather than as context-bound hypotheses, the gap between the cited evidence and the conclusions is load-bearing. This does not invalidate the central position, but it means the 'grounded in evidence' framing is stronger than the current empirical base supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that social robots for mental wellbeing should be understood as supportive tools rather than replacements for human therapists, and it organizes the argument around six 'insights' drawn from the authors' collective research and selected literature: (1) there is no single ground-truth measure of wellbeing; (2) robots need not act as companions to be effective; (3) virtual/mediated interaction is a viable and promising mode; (4) clinicians should be involved in design; (5) one-off interactions are of limited value compared with longitudinal engagement; and (6) adaptation and personalization are not always necessary. The paper also discusses therapeutic alliance, validation requirements, dependency risks, privacy, and fairness, and it concludes with recommendations for collaborative, evidence-based, ethically grounded deployment.","tokens_in":22581,"tokens_out":3924,"duration_ms":41006,"significance":"If the six insights are accepted as transferable design lessons, they provide a useful counterweight to techno-optimistic accounts of robotic mental-health support. The paper's central reframing—away from replacement and toward complementing human practitioners—is timely and well aligned with a growing consensus in HRI and mental-health ethics. The authors are also to be credited for explicitly discussing dependency risks, the need for 'designing for exit,' and a concrete fairness concern (the VLM-based assessment bias reported in [128]). However, the empirical basis for the insights is narrow: most supporting studies are small, same-group, non-clinical, and often without control conditions. The paper is honest about this in the Discussion, but the abstract's 'grounded in evidence' framing is stronger than the current evidence permits. Because the insights are phrased as general conclusions rather than context-bound observations, the strength of the evidentiary base is load-bearing for the paper's contribution.","major_comments":[{"comment":"The six insights are presented as generalizable design lessons, but the supporting empirical base is dominated by the authors' own studies with small non-clinical samples (e.g., n=17 in [81], n=41 in [34], and the caregiver studies [17,20]). The paper itself acknowledges in the Discussion that 'Current studies often rely on short-term measures and non-clinical populations, limiting the generalizability of findings, as well as their replicability.' This acknowledgment is in tension with the abstract's claim that the insights are 'grounded in evidence.' To make the central claim defensible, the authors should either (a) reframe the six insights as context-bound observations or testable hypotheses, or (b) add a systematic or semi-systematic literature search demonstrating that the patterns hold beyond their own studies.","section":"Section 2 (Insights 1–6) and Discussion"},{"comment":"The insight 'robot doesn't need to be a companion to improve wellbeing' is supported mainly by feasibility arguments (cost, market availability) and by the authors' own coach-deployment studies, rather than by direct comparative evidence. Notably, the one comparative study cited, Jeong et al. [39], found that the companion role was more effective than the coach role in building therapeutic alliance and enhancing wellbeing. The claim that companionship is unnecessary does not follow from the cited evidence. Please either present direct comparisons of coach and companion roles on wellbeing outcomes or soften the insight to state that coach roles can be effective in specific deployment contexts.","section":"Section 2.2"},{"comment":"The claim that 'one-off HRI may be not enough for improving mental wellbeing, but it can help!' is not directly supported by the studies cited. The one-off studies [34,58] assess wellbeing or support anxiety reduction in a specific procedure, and the evidence for within-session emotional gains comes from a longitudinal study [78] that is not a one-off interaction. The paper should clarify what 'help' means here (e.g., assessment value, short-term mood effects) and should cite studies with pre-post measures of one-off interactions if that is the claim.","section":"Section 2.5"},{"comment":"The virtual-modality insight rests primarily on two same-group studies [20,17] with self-reported outcomes and no control condition. The additional equivalence studies cited [53–55] concern general perceptions and behavior in HRI, not wellbeing outcomes specifically. To support the strong claim that virtual interactions can effectively support wellbeing, the paper should either provide direct evidence of comparable wellbeing outcomes for physical versus virtual robot delivery or limit the claim to feasibility and accessibility advantages.","section":"Section 2.3"}],"minor_comments":[{"comment":"There is a typo: 'reseach' should be 'research.'","section":"Section 2.1"},{"comment":"The phrase 'relaying on context and settings' should be 'relying on context and settings.'","section":"Section 2.3"},{"comment":"The Introduction states that the paper is 'not intended as a comprehensive review,' but Section 4 opens with 'This paper has provided a comprehensive examination.' Please align these statements to avoid an internal contradiction.","section":"Section 1 and Section 4"},{"comment":"References [14] and [34] appear to be the same conference paper (Abbasi et al., 'Can robots help in the evaluation of mental wellbeing in children?') and should be consolidated.","section":"References"},{"comment":"The phrase 'It is important to be vigilant about this post-study as well' is ambiguous; presumably 'post-screening' is intended.","section":"Section 2.1"},{"comment":"There is a missing space around the comma in 'by using a LLM) , with participants'; insert the space and remove the comma before the relative clause.","section":"Section 2.6"}],"recommendation":"major_revision","confidential_remarks":"This is a heavily self-referential paper: most of the cited empirical support comes from the authors' own prior studies. That is not disqualifying for a position paper, but the 'insights' framing implies a level of generality that the evidence base does not yet support. A major revision that humbles the claims or adds a credible external evidence base would make the contribution publishable. I would not recommend rejection because the paper's central framing--robots as complementary, carefully designed tools--is valuable and the limitations are acknowledged in the Discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, this is a perspective piece from the Cambridge HRI group, and it reads like a mature synthesis of their own decade of work on robots and wellbeing. The punchline: the six insights are sensible, clearly written, and will be useful as a framing reference, but they are not new empirical results, and the evidence base is mostly the authors' own small-sample studies. The paper itself admits this more honestly than most.\n\nWhat is actually new: the explicit articulation of robots as supportive tools rather than therapist replacements, and the sixth insight that adaptation is optional, supported by the mentalization literature. The paper does well in being explicit about limitations: non-clinical populations, short-term measures, need for RCTs and longitudinal work. The discussion of dependency, privacy, and fairness is thoughtful, and I appreciated the concrete example of gender bias in the VLM-based assessment.\n\nSoft spots: the insights in §2.2–2.6 are largely restatements of the authors' prior findings. Self-citation is heavy, though much of it is their own empirical work. Several generalizations rest on small samples (n=17, n=41) without control conditions. The stress-test note worried about this, and I think the concern is real but not fatal: the paper positions these as insights from their experiences, not as systematic review conclusions. The bigger issue is that the 'grounded in evidence' framing in the abstract is stronger than the base supports. If the authors had labeled these as lessons learned or hypotheses to test, the gap would shrink.\n\nBottom line: this is a useful, readable perspective for HRI researchers entering the wellbeing space, and for clinicians collaborating with roboticists. It deserves a serious referee—the synthesis is valuable and the self-awareness is genuine. I'd recommend sending it to review, and I'd expect the reviewers to ask for more careful hedging of the generalization claims.","headline":"A clear, honest synthesis of the authors' own HRI wellbeing work: the six insights are sensible and the paper is self-aware about its limits, but the 'grounded in evidence' framing outruns the small, self-referential sample base.","tokens_in":22995,"tokens_out":1823,"would_cite":true,"duration_ms":20613,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Social robots can support mental wellbeing as evidence-based coaching tools rather than therapist replacements, and six insights from field studies guide their design.","keywords":["human-robot interaction","socially assistive robots","mental wellbeing","robot ethics","virtual human-robot interaction","long-term interaction","wellbeing measurement","adaptation and personalization"],"falsifier":"A preregistered, multi-site randomised controlled trial in a clinical or help-seeking population showing that a non-adaptive coach robot produces no wellbeing improvement while an adaptive companion robot does, or a matched virtual-versus-physical comparison in which users confide significantly less to the virtual robot over repeated sessions, would directly contradict the paper's main insights.","tokens_in":22054,"feed_emoji":"🤖","tokens_out":6991,"duration_ms":69274,"temperature":0.7,"pith_summary":"This paper argues that social robots can genuinely support mental wellbeing, but only when they are designed as supportive tools rather than replacements for human therapists. Drawing on a body of empirical studies and deployments, mostly from the authors' own research group, it distils six transferable insights: there is no single objective ground truth for wellbeing; a robot does not need to be a companion to help; virtual and video-mediated interaction can work as well as physical presence; clinicians and wellbeing professionals should be co-designers; one-off interactions are useful but sustained engagement matters; and personalisation is optional because people often read social intent into even non-adaptive robots. The stated goal is to steer future research and practical deployment toward evidence-based, ethically guarded uses of robots in mental health and wellbeing contexts.","feed_headline":"Robots support wellbeing as coaches, not therapist substitutes","feed_subtitle":"Field-tested design lessons: virtual delivery works, clinicians belong in the loop, adaptation optional.","key_machinery":"The paper's organising device is a set of six insights, each presented as a background challenge followed by a transferable design lesson. The mechanism carrying the argument is cumulative, convergent evidence from the authors' repeated studies—long-term self-disclosure sessions, robotic positive-psychology coaching at work, group mindfulness in a cafe, child wellbeing assessment, and an emotion-regulation intervention—alongside selective external comparisons of companion versus coach roles and virtual versus physical delivery. A supporting psychological mechanism is mentalization, the tendency to attribute intention and meaning to a robot, which allows standardised, low-adaptivity robots to still be experienced as responsive and supportive.","core_discovery":"The central claim, stated directly in the abstract, is that 'rather than positioning robots as replacements for human therapists, we argue that they are best understood as supportive tools that must be designed with care, grounded in evidence, and shaped by ethical and psychological considerations.' The paper's supporting observation is that wellbeing is multifaceted and subjective: no single questionnaire, self-report, or behavioural signal can serve as a unified ground truth, so evaluation and design must work with gold standards and multiple perspectives. From there the authors generalise across their own longitudinal and in-the-wild studies to argue that coach-like roles in workplaces, schools, and public spaces can be effective, that mediated interactions can preserve perceived social presence, and that even simple, non-adaptive robots can produce positive emotional outcomes through users' mentalization and meaning-making.","pith_inferences":["Editorial inference: because the paper is explicitly not a systematic review and the insights come from one group's studies, the insights are best read as transferable hypotheses; a systematic review with preregistered inclusion criteria would test how broadly they hold.","Editorial inference: if mentalization makes low-adaptivity robots effective, then cheap, non-personalised robots could be scaled across public and community settings; the same mechanism, however, makes dependency more likely, so 'designing for exit' should become a standard requirement rather than an afterthought.","Editorial inference: the virtual-modality results suggest a resource-leveraging model in which one physical robot plus a telehealth-style video pipeline reaches many users; a testable extension is comparing adherence and clinical outcomes between online and in-person robot coaching over several months.","Editorial inference: the paper's fairness observation about a robot assessment pipeline misclassifying girls' stories more often than boys' implies that any move toward clinical use of robot-led wellbeing assessment should require stratified validation across gender, age, and cultural groups before deployment."],"forward_implications":["Wellbeing robots can be deployed as coaches, trainers, or facilitators in shared spaces, not only as personal companions, widening feasible deployment contexts.","Video-mediated robot interactions can preserve engagement and perceived social presence, making support accessible to geographically isolated, mobility-limited, or home-bound populations.","Design processes should include clinicians, psychologists, and wellbeing coaches alongside end users, because professional judgment can veto superficially appealing features such as free verbal adaptation.","One-off sessions are appropriate for assessment, early design iteration, and circumscribed situational support, while sustained change and therapeutic alliance require longitudinal deployments.","Researchers should treat adaptation as a cost-benefit decision rather than a default requirement, since standardised interactions can be effective and easier to evaluate, and users may perceive understanding even without personalisation."],"supporting_citations":[{"why":"Supplies the ground-truth versus gold-standard distinction that underlies Insight 1.","marker":"[25]"},{"why":"Long-term self-disclosure study used as evidence for virtual interaction, limited adaptivity, and emotional attachment.","marker":"[20]"},{"why":"Replication with informal caregivers, showing virtual delivery and emotional attachment to the robot.","marker":"[17]"},{"why":"Compares robot companion and coach roles longitudinally, grounding Insight 2 and long-term benefits.","marker":"[39]"},{"why":"In-the-wild workplace study of robotic positive-psychology coaches with clinician involvement.","marker":"[19]"},{"why":"Group mindfulness deployment in a cafe, showing the coach role and variability in role perception.","marker":"[40]"},{"why":"Design and ethical recommendations, including clinician input and limits on verbal adaptation.","marker":"[18]"},{"why":"Child-robot wellbeing assessment study, supporting clinician collaboration and the limits of one-off studies.","marker":"[34]"},{"why":"Robot-led emotion-regulation intervention showing within-session emotional trends and comparison with a more adaptive LLM robot.","marker":"[78]"},{"why":"VITA framework and longitudinal robotic coaching study showing wellbeing gains over four weeks.","marker":"[81]"}],"fun_headline_variants":["Robots as wellbeing coaches, not therapist stand-ins","Six insights for robots in mental wellbeing support","Design lessons: robots aid wellbeing as coaches","Virtual robot sessions can support mental wellbeing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the authors' own studies, mostly small, non-clinical, and produced by one research group, are representative enough to support generalisable insights about robot roles, virtual delivery, and adaptivity.","fun_headline_variants_meta":{"raw":{"variants":["Robots as wellbeing coaches, not therapist stand-ins","Six insights for robots in mental wellbeing support","Design lessons: robots aid wellbeing as coaches","Virtual robot sessions can support mental wellbeing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1167,"prompt_tokens":865,"completion_tokens":302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":481,"tokens_out":302,"duration_ms":3263,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:26:09.329128+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered, multi-site randomised controlled trial in a clinical or help-seeking population showing that a non-adaptive coach robot produces no wellbeing improvement while an adaptive companion robot does, or a matched virtual-versus-physical comparison in which users confide significantly less to the virtual robot over repeated sessions, would directly contradict the paper's main insights.","supporting_citations":[{"cited_title":"International Journal of Social Robotics 16, 1–27 (2024) https: //doi.org/10.1007/s12369-023-01076-z","cited_arxiv_id":null,"evidence_quote":"Long-term self-disclosure study used as evidence for virtual interaction, limited adaptivity, and emotional attachment."},{"cited_title":"International Journal of Social Robotics (2025) https://doi.org/10.1007/s12369-024-01207-0","cited_arxiv_id":null,"evidence_quote":"Replication with informal caregivers, showing virtual delivery and emotional attachment to the robot."}],"review_version":1}