{"id":"fa115d0f-964c-4cf1-8e2a-aaea3c1c7bb4","arxiv_id":"2506.17831","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-week factorial study with 26 students found that engagement and evaluative intimacy increased with a physical robot but decreased with a chatbot.","lead":"In a two-week study, 26 university students did daily CBT exercises at home with either a physical robot or a chatbot. Engagement and emotionally intimate disclosure increased for the robot group over time but decreased for the chatbot group, suggesting physical presence may matter for sustained digital therapy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Embodiment claim is confounded with output modality: the robot spoke via AWS Polly while the chatbot appears to have used text, so the time-by-embodiment interaction may be a speech effect rather than an effect of physical embodiment.","rationale":"The paper is a small, clearly reported empirical study, and the reader's CONDITIONAL verdict is appropriate. My stress-test pass did not find a reason to move away from that verdict; instead, it sharpens the same load-bearing concern the reader identified. The central claim in the abstract and conclusions is causal: physical embodiment drives increased engagement and intimacy over time. The methods, however, operationalize the independent variable as a bundle that includes physical presence, spoken output via AWS Polly, and robot head movements (Section III-A3), while the chatbot condition is text-based. Because the design does not hold output modality constant, the time-by-embodiment interactions for evaluative intimacy and engagement cannot be uniquely attributed to embodiment. That is not a disagreement with the statistical findings; it is a limitation on the construct the findings can support. The proposed three-arm follow-up is the direct test: it controls the spoken transcript and voice across embodied and non-embodied conditions, isolating physical presence. Until such a test is run or the claims are reframed as 'multimodal embodied agent vs text chatbot,' the conclusions overreach. I kept the reader's CONDITIONAL verdict because the concern is addressable and does not invalidate the descriptive results, but the paper should not be accepted without either the additional comparison or explicitly narrowed language.","tokens_in":11656,"tokens_out":3456,"duration_ms":37935,"concrete_test":"Run a three-arm follow-up sharing identical GPT-3.5 transcripts and identical Joanna TTS audio: (1) physical Blossom robot with speech, (2) a laptop avatar or plain screen with the same speech, and (3) text-only chatbot. If the time-by-condition interaction for evaluative intimacy and active engagement is absent or attenuated in arm 2 relative to arm 1, the effect is attributable to speech or multimodal delivery, not to physical embodiment; if arm 2 matches arm 1, embodiment is not needed. Power the study for the reported interaction effect sizes (partial eta-squared .09 and .03).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim attributes the divergent time trends to 'physical embodiment.' That inference requires the robot and chatbot conditions to differ only in embodiment. They do not. Section III-A3 specifies that the Blossom robot verbalized LLM-generated chat messages with the AWS Polly 'Joanna' voice and made synchronized head movements, while the chatbot condition is described only as a web application and presumably presented text. Speech is a well-known social cue that can increase engagement and perceived presence, so the observed interactions (evaluative intimacy F(1,24)=7.30, p=.01; engagement F(1,24)=5.14, p=.03) could be caused by the voice or by multimodal presentation rather than by physical presence. The Discussion in Section V repeatedly concludes that 'embodiment is critical' and credits 'the SAR's physical embodiment,' but the design cannot separate embodiment from speech. This is a construct-validity problem, not a statistical one: even if the ANOVA results are exactly correct, they do not support the causal claim about embodiment. A secondary concern is baseline imbalance (robot Day 1 evaluative intimacy .27 vs chatbot .20) with no equivalence test, so part of the interaction could reflect regression to the mean, but the main issue remains the modality confound.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a two-week deployment study with 26 university students who completed daily CBT exercises with either an LLM-powered Blossom robot (with AWS Polly speech and synchronized head movements) or a disembodied chatbot accessed via a web application. Transcripts were manually annotated for descriptive intimacy, evaluative intimacy, and engagement. Two-way mixed ANOVAs on first-day versus last-day sessions found significant time-by-condition interactions for evaluative intimacy (F(1,24)=7.30, p=.01) and engagement (F(1,24)=5.14, p=.03), with sample means suggesting increases in the robot condition and decreases in the chatbot condition. The authors conclude that physical embodiment drives greater engagement and intimate disclosure over time.","tokens_in":11879,"tokens_out":3858,"duration_ms":40842,"significance":"If the causal claim about physical embodiment were supportable, this would be a useful longitudinal contribution to HRI and digital mental health: it is one of the few comparisons of a physically present SAR versus a disembodied chatbot for daily CBT exercises, conducted in participants' residences over two weeks. Strengths include manual annotation with a published intimacy scheme, reported inter-coder reliability, and transparent reporting of ANOVA statistics. The core interaction findings are plausible and worth reporting, but the interpretation requires substantial revision because of the embodiment/speech confound and the absence of within-condition simple-effects tests.","major_comments":[{"comment":"The robot condition confounds physical embodiment with speech output: Section III-A3 states that the Blossom robot verbalized the LLM-generated messages with the AWS Polly 'Joanna' voice and made synchronized head movements, while the chatbot condition is described only as a web application and the paper does not state whether it presented text, speech, or an animated character. The observed interactions for evaluative intimacy and engagement therefore cannot be attributed to physical embodiment per se; they may reflect speech, multimodal presentation, or other social cues. The causal language in the Abstract, Section V, and Conclusion ('embodiment is critical', 'the SAR's physical embodiment') goes beyond what this design can establish. Please reframe the manipulation as a physically present speaking robot versus a text-based chatbot, or provide evidence that speech was controlled across conditions.","section":"Section III-A3 and Section IV-A"},{"comment":"The directional claims that engagement and intimacy 'increased over time in the physical robot condition, while both measures decreased in the chatbot condition' are not directly supported by the reported analyses because no simple-effects tests within each condition are reported. A significant interaction (F(1,24)=7.30, p=.01 for evaluative intimacy; F(1,24)=5.14, p=.03 for engagement) establishes only that the time trends differ between conditions, not that either trend is significantly different from zero. Please report within-condition paired comparisons between first and last day, with means, standard deviations, test statistics, and effect sizes, and adjust the abstract and conclusions to match those results.","section":"Abstract and Sections IV-A and IV-C"},{"comment":"H2b ('CBT exercises will affect evaluative intimacy outcomes more than descriptive intimacy outcomes') is declared supported, but the paper never performs a statistical comparison between the two dependent variables. Separate ANOVAs on evaluative and descriptive intimacy cannot establish that the two outcomes differ in magnitude or trajectory; a proper test would require a within-subjects comparison of the two intimacy measures in a single model or an explicit test of the difference between their effect sizes. If such a test is not available, H2b should be described as an informal observation rather than a supported hypothesis.","section":"Section V, H2b"},{"comment":"The two conditions show day-1 baseline differences in evaluative intimacy (robot M=0.27 vs chatbot M=0.20) and engagement (robot M=0.84 vs chatbot M=0.79), but no baseline equivalence test is reported. With group sizes of 14 and 12, chance imbalance or regression to the mean could contribute to the observed interaction. Please report a test of day-1 differences between conditions and, if possible, a sensitivity analysis based on change scores or baseline-adjusted models.","section":"Section IV-A and IV-C"}],"minor_comments":[{"comment":"The inter-coder reliability section reports average percent agreement and an average Cohen's kappa of 0.603, described as 'substantial'; please clarify whether this is Cohen's kappa or another variant, how the average is computed across annotator pairs, and whether kappa was computed on a per-item or per-transcript basis.","section":"Section III-B2"},{"comment":"The 'effect size' values (0.21, 0.09, 0.03) are reported without stating that they are partial eta-squared, which is the standard measure associated with these F-tests; please label them explicitly.","section":"Section IV-A and IV-C"},{"comment":"The study is described as a 'two-week' study in the Abstract and as a '15-day' study in the Conclusion; please harmonize these descriptions and clarify whether 'last day' refers to a fixed study day or to each participant's final completed session, especially since exercises were optional after day 8.","section":"Section III-A1 and Section VI"},{"comment":"Three separate ANOVAs are presented without any correction for multiple comparisons; given the small sample and secondary nature of descriptive intimacy and engagement analyses, please acknowledge this or report adjusted p-values.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The central interaction findings are likely valid and interesting, but the paper's strongest claims about physical embodiment are undercut by the speech confound. I believe the issues are fixable within the manuscript's scope if the authors reframe the manipulation, add simple-effects tests, and temper the causal language. I would not recommend rejection because the empirical pattern is worth reporting, but the revision needs to be substantive."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper: it is an honest, small longitudinal field study, and the headline result is real in the narrow sense – engagement and evaluative intimacy went up over two weeks with the robot and down with the chatbot. But the paper's central attribution to 'physical embodiment' does not survive contact with its own methods. The robot spoke (AWS Polly), and the chatbot presumably presented text. So the differential time trend could be a voice effect or a multimodal effect, not embodiment per se.\n\nWhat is genuinely new is the factorial time-by-embodiment comparison in LLM-delivered CBT homework; I haven't seen that specific design elsewhere. The annotation is careful: they used a published intimacy coding scheme, trained four annotators, and report kappas around 0.60. The ANOVA structure is standard, with type III sums of squares and effect sizes. The interaction for evaluative intimacy (F(1,24)=7.30, p=.01, eta²=.09) is plausible given the means.\n\nThe soft spots are real. The modality confound is the strongest one. It's not a statistical quibble; it's a construct-validity problem. The discussion repeatedly says 'embodiment is critical,' but the design cannot tell embodiment from speech. A secondary issue is baseline imbalance: Day 1 evaluative intimacy is 0.27 in the robot group and 0.20 in the chatbot group, with no equivalence test, so part of the interaction could be regression to the mean. Also, the abstract's 'increased/decreased' language is not backed by within-condition simple effects tests; they only report the interaction. And the engagement interaction (p=.03) would likely not survive a correction for three ANOVAs. These are all fixable in revision: report simple effects, add a speaking chatbot or text-only robot control, and discuss the confound explicitly.\n\nOverall, a solid empirical contribution to a narrow question, with a load-bearing caveat. I'd bring it to a reading group as a good example of how easy it is to conflate embodiment with other cues. A serious referee should see it, but the authors need to reframe the claim before it's publishable.\n\nRecommendation: accept the paper for review with a request for major revision.","headline":"Real longitudinal data, but the 'embodiment' claim is confounded by speech; the study deserves review, not blind acceptance.","tokens_in":12452,"tokens_out":3755,"would_cite":true,"duration_ms":33196,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A physical robot, but not a chatbot, increased engagement and intimate disclosure over two weeks of daily CBT exercises.","keywords":["LLM","socially assistive robots","chatbot","cognitive behavioral therapy","self-disclosure","intimacy","engagement","embodiment"],"falsifier":"Run the same two-week protocol with four conditions: robot with voice, robot with text, chatbot with voice, and chatbot with text; if voice rather than physical presence drives the temporal effect, the robot-with-text condition should show the chatbot-style decline and the chatbot-with-voice should show the robot-style rise.","tokens_in":11469,"feed_emoji":"🤖","tokens_out":4886,"duration_ms":44427,"temperature":0.7,"pith_summary":"The paper claims that the physical form of an AI therapy agent changes how people open up over time. In a two-week study of daily cognitive behavioral therapy exercises, participants speaking with an embodied robot increased their active engagement and their disclosure of opinions, judgments, and emotions from the first day to the last, while participants using a text chatbot decreased on both measures. The statistically significant interaction between embodiment and time is the paper's central evidence that physical presence matters for sustaining therapeutic self-disclosure, not just for initial reactions. If the finding holds, it points toward socially assistive robots, rather than chatbots, as the more promising LLM-based format for at-home mental health support.","feed_headline":"Robot keeps therapy users engaged; chatbot loses them in 2 weeks","feed_subtitle":"Daily CBT transcripts show disclosure and engagement rising with a robot, falling with a chatbot.","key_machinery":"The argument runs on a two-by-two factorial mixed ANOVA with embodiment (between subjects) and time (first vs last day, within subjects) as factors, applied to transcripts hand-annotated by four coders. The annotation scheme, based on Morton's intimacy framework, separates descriptive intimacy (facts about oneself) from evaluative intimacy (opinions, judgments, emotions), and codes engagement as active or passive; intercoder reliability was substantial (Cohen's kappa = 0.603). The embodied condition used the handcrafted Blossom robot with AWS Polly speech and synchronized head motion, while the chatbot condition presented the same LLM-generated CBT exercises in a web application without a physical body or spoken voice. This setup is what lets the interaction term in the ANOVA carry the claim that embodiment changes how disclosure evolves over time.","core_discovery":"In the paper's own terms, the discovery is an interaction: over a two-week daily CBT program, the embodied robot condition improved while the disembodied chatbot condition deteriorated. For evaluative intimacy, the robot group's mean high-intimacy disclosure percentage rose from 0.27 to 0.42, while the chatbot group fell from 0.20 to 0.12, yielding F(1,24)=7.30, p=0.01, partial eta-squared 0.09. For active engagement, the robot group rose from 0.84 to 0.90 and the chatbot group fell from 0.79 to 0.69, yielding F(1,24)=5.14, p=0.03, partial eta-squared 0.03. The main effect of embodiment was also significant for evaluative intimacy (robot 0.35 vs chatbot 0.16, p=0.006). The authors interpret these results as support for the claim that embodiment of an LLM-powered agent is critical for encouraging engagement and intimate disclosure in longitudinal therapeutic exercises.","pith_inferences":["The design confounds physical embodiment with voice output: only the robot spoke aloud. A follow-up with robot-plus-text and chatbot-plus-voice arms is needed to isolate whether the temporal effect comes from the body or the voice.","Because only the first and last days were analyzed, the shape of the trajectory is unknown; daily or session-level modeling could show whether the divergence is steady or driven by a single point.","With partial eta-squared values of 0.09 and 0.03, the effects are small in variance-explained terms, and the clinical significance for depression or anxiety outcomes is not established.","The sample is 26 university students with PHQ-9 below the depression threshold, so generalizing to clinical populations or older adults remains an open question."],"forward_implications":["Developers of LLM-based therapeutic chatbots should consider adding physical embodiment if the goal is sustaining engagement over repeated sessions.","Longitudinal chatbot studies may show declining engagement and disclosure, matching the paper's chatbot arm, so embodiment may counteract that decline.","Evaluative intimacy, not descriptive intimacy, is the measure that responded to embodiment, suggesting emotional disclosure is the sensitive outcome for embodied agents.","At-home CBT homework supported by a robot could be a viable complement to human therapy, consistent with the paper's proposal.","Eight CBT sessions have been shown sufficient for outcomes; the two-week divergence observed here suggests embodiment may help maintain homework participation through that window."],"supporting_citations":[{"why":"Supplies the two-dimensional intimacy framework (descriptive vs evaluative) that structures the annotation scheme.","marker":"[36]"},{"why":"Provides the baseline expectation that a physical robot outperforms other embodiments in engagement.","marker":"[34]"},{"why":"Prior study by the same group using the LLM-powered Blossom robot, on which the current protocol builds.","marker":"[7]"},{"why":"Supports the social-presence advantage of embodied agents that the paper uses to explain its results.","marker":"[58]"},{"why":"Documents loss of interest in chatbots over time, which motivates the interpretation of the chatbot arm's decline.","marker":"[61]"},{"why":"Provides the content-analysis methodology used to guide transcript annotation and coding.","marker":"[52]"}],"fun_headline_variants":["Robot therapy boosts intimacy; chatbot fades in 2 weeks","Embodiment matters: robot lifts CBT disclosure, chatbot drops","Daily CBT: robot users open up, chatbot users clam up","Two weeks: robot gains engagement, chatbot loses it","LLM robot outshines chatbot in longitudinal CBT study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The robot condition differs from the chatbot condition in both physical presence and spoken voice, so the paper's credited embodiment effect may actually be an effect of speech output, and the groups' day-one intimacy scores are not shown to be statistically equivalent.","fun_headline_variants_meta":{"raw":{"variants":["Robot therapy boosts intimacy; chatbot fades in 2 weeks","Embodiment matters: robot lifts CBT disclosure, chatbot drops","Daily CBT: robot users open up, chatbot users clam up","Two weeks: robot gains engagement, chatbot loses it","LLM robot outshines chatbot in longitudinal CBT study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000522,"raw_usage":{"total_tokens":2524,"prompt_tokens":944,"completion_tokens":1580,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":1498}},"tokens_in":560,"tokens_out":1580,"duration_ms":10608,"temperature":1.0,"reasoning_tokens":1498,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:59:56.509612+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two-week protocol with four conditions: robot with voice, robot with text, chatbot with voice, and chatbot with text; if voice rather than physical presence drives the temporal effect, the robot-with-text condition should show the chatbot-style decline and the chatbot-with-voice should show the robot-style rise.","supporting_citations":[{"cited_title":"Intimacy and reciprocity of exchange: A comparison of spouses and strangers","cited_arxiv_id":null,"evidence_quote":"Supplies the two-dimensional intimacy framework (descriptive vs evaluative) that structures the annotation scheme."},{"cited_title":"Effect of a robot on user perceptions,","cited_arxiv_id":null,"evidence_quote":"Provides the baseline expectation that a physical robot outperforms other embodiments in engagement."},{"cited_title":"More than just a pretty face,","cited_arxiv_id":null,"evidence_quote":"Supports the social-presence advantage of embodied agents that the paper uses to explain its results."},{"cited_title":"Users’ experiences with chatbots: findings from a questionnaire study,","cited_arxiv_id":null,"evidence_quote":"Documents loss of interest in chatbots over time, which motivates the interpretation of the chatbot arm's decline."},{"cited_title":"Krippendorff,Content Analysis: An Introduction to Its Methodology (second edition)","cited_arxiv_id":null,"evidence_quote":"Provides the content-analysis methodology used to guide transcript annotation and coding."}],"review_version":1}