{"id":"acb093f5-cfd4-4390-8f70-39eab5d5dbd7","arxiv_id":"2504.21702","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A 600-person study of a well-being chatbot found that intrinsic motivation factors predicted self-reported awareness and behavioral intention, while conversational style (formal vs. informal, with or without emojis) had no significant effect.","lead":"The authors tested a chatbot named Allegra that coaches people toward healthier habits, in a study with 600 participants on Prolific. They found that users who found the bot interesting, useful, or trustworthy also reported more awareness and stronger intention to change, while the bot's formal versus informal style had no measurable effect.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The strongest path estimates in §5.4 may be inflated by common-method variance and cannot, on their own, support the causal wording used for H1/H2.","rationale":"The Reader's weakest assumption is also the one I would stress: the model's load-bearing inputs are one-shot self-reports, and the paper's causal verbs ('confirm', 'causal influence') are not supported by the design. The proposed method-factor re-analysis is feasible with the existing data and would directly test whether the shared method component drives the paths. I considered other candidate concerns, such as the poor RMSEA and the internal reporting inconsistencies; these are real but secondary. I also considered the absence of equivalence testing for the conversational-style null effect; that weakens the negative claim but does not threaten the positive causal claim as directly. Because the reader already marked the paper CONDITIONAL and my concern matches the reader's central worry, no verdict change is needed. The paper should be revised to replace causal language with associational language, report a common-method sensitivity analysis or explicitly justify its absence, and correct the statistical reporting errors in §5.3 and §5.4.","tokens_in":14582,"tokens_out":6121,"duration_ms":68748,"concrete_test":"Re-estimate the §5.4 lavaan model with an added common-method latent factor on which every questionnaire item loads, constraining the method loadings to equality across constructs to keep the model identified. If the IM→AC and IM→BI standardized paths drop below 0.3 or lose significance, common-method variance is a plausible explanation for the claimed causal influence; if the paths remain large and significant, the common-method objection is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that intrinsic motivation 'influences' awareness creation and behavioural intention. This is based on §5.4, where a cross-sectional SEM on same-questionnaire responses yields IM→AC = 0.93 and IM→BI = 0.45. All constructs (intrinsic motivation items, awareness creation, behavioural intention) were measured once, immediately after the chat, in a single self-report instrument. A general positive evaluation of the Allegra experience could therefore inflate all item groups. The EFA finding of three factors argues against a single method factor entirely explaining the data, but it does not rule out a shared method component that still biases the structural paths. In addition, RQ1 is answered from absolute mean levels (AC 4.58, BI 5.79) without any no-chat or reading-only baseline, so the wording 'confirm' and 'causal influence' in Sections 5.2 and 5.4 goes beyond what this design can establish. Secondary reporting inconsistencies (the swapped ANOVA p-values in §5.3 and calling RMSEA 0.10 'good or very good') do not change the core concern but further support a conditional rather than accepting verdict.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a user study of a scripted conversational agent (Allegra) designed to promote well-being awareness and behavioural intention. Six hundred Prolific participants interacted with one of three versions differing in conversational style (informal, informal with multimedia elements, formal) and then completed a questionnaire measuring intrinsic motivation factors (interest, value, trust, relatedness), awareness creation (AC), and behavioural intention (BI). The authors use exploratory and confirmatory factor analyses (EFA/CFA) within a structural equation modelling framework to test H1 (intrinsic motivation influences BI) and H2 (intrinsic motivation influences AC), and they complement this with a sentiment analysis of free-text comments and a content analysis of conversation-topic choices. They report that intrinsic motivation has a strong positive path to AC (0.93) and a moderate path to BI (0.45), and that conversational style has no significant effect on behavioural intention. The abstract and conclusions present these results as evidence of causal influence.","tokens_in":14755,"tokens_out":5072,"duration_ms":50485,"significance":"If interpreted as correlational associations, this is a reasonably large (N=600) empirical contribution to the design of conversational well-being tools. Its strengths include a clearly described instrument based on established IMI constructs, a three-arm comparison of conversational styles, and an explicit null result for style effects, which is useful for the HCI literature. The paper does not ship code or data, but the methodological description is sufficiently detailed to permit replication. The main value is provisional and hypothesis-generating; the causal framing and the fit-statistic presentation currently exceed what the cross-sectional, single-source design can support.","major_comments":[{"comment":"The statement that the data 'confirm' H1 and H2 and that intrinsic motivation has a 'causal influence' is not supported by the design. All three construct families (intrinsic motivation, awareness creation, behavioural intention) were measured in one self-report questionnaire immediately after the chat, with no manipulation or temporal separation of the predictor; the SEM paths therefore estimate associations that can be inflated by common-method variance. The EFA's three-factor solution argues against a single method factor entirely explaining the data, but it does not rule out a shared method component that biases the structural paths. Please reword the claims as correlational, add an explicit common-method-bias limitation, and consider longitudinal or multi-source designs in future work.","section":"Section 5.4, Figure 5, H1/H2"},{"comment":"The answer to RQ1 is based solely on absolute means (AC 4.58, BI 5.79) on a 1–7 scale, with no no-intervention or reading-only control condition. Mean levels above the scale midpoint do not demonstrate that the conversational approach 'influence[s]' awareness creation or behavioural intention; they only describe the participants' reported levels after exposure. Please reframe this as a descriptive finding and explicitly acknowledge the absence of a baseline.","section":"Section 5.2, RQ1"},{"comment":"The fit indices are misreported. RMSEA = 0.10 is conventionally considered poor or at best borderline, not 'good or very good'; the text also contains 'RMSEA > 0.8', which appears to be a typo for 0.08 or 0.10. Because the authors use the fit statement to support the confirmatory conclusion, the fit indices should be reported accurately (e.g., CFI/TLI around 0.90 and RMSEA = 0.10 suggest marginal fit) and the implications for H1/H2 should be discussed rather than glossed over.","section":"Section 5.4, fit statistics"}],"minor_comments":[{"comment":"The composite IM–BI correlation of 0.42 is inconsistent with the item-level correlations of 0.54–0.58 from which it is presumably derived; please clarify how the composite was computed, since a simple average of the four reported item correlations would be about 0.57.","section":"Table 4"},{"comment":"The ANOVA p-values appear swapped: F(1, 596) = 65.16 would correspond to an extremely small p-value, while F(1, 596) = 9.86 would correspond to a p around 0.002; please verify the reported values.","section":"Section 5.3"},{"comment":"Several items are phrased negatively (e.g., 'Allegra did not hold my attention at all') and it is not stated whether they were reverse-scored before computing Cronbach's alpha; the third behavioural-intention item is phrased as an open question rather than a 1–7 statement and should be aligned with the other items.","section":"Table 2"},{"comment":"The phrase 'may be a reason for RMSEA > 0.8' is unclear; also, 'Explanatory Factor Analysis' should be 'Exploratory Factor Analysis'.","section":"Section 5.4"},{"comment":"The manuscript would benefit from a dedicated limitations subsection that explicitly lists the same-source design, lack of a baseline, no long-term follow-up, and possible demand effects from the chat setting.","section":"Conclusions"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an HCI venue and the data are, for the most part, transparently reported. The central issue is that the causal language in the abstract, Section 5.2, and Section 5.4 goes beyond what a cross-sectional, single-questionnaire design can establish, and the fit-statistic reporting needs correction. These are fixable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on arXiv:2504.21702. The paper is a decent field study of a scripted well-being chatbot (Allegra) with 600 Prolific participants, three style conditions. The genuinely useful finding is the null: conversational style (informal, informal+multimedia, formal) had no measurable effect on behavioural intention. That's a real contribution to a subfield that often assumes style matters. The dataset is original, the protocol is transparent, and the authors are honest that the interaction is fully scripted and short, which they rightly flag as a possible reason for the null.\n\nWhat I'd push back on is the framing around H1/H2. The abstract says the results 'confirm the positive effect of intrinsic motivation factors on both awareness creation and behavioural intention.' The SEM gives paths of 0.93 and 0.45 from a single cross-sectional questionnaire, with all items collected at the same time after the chat. That design cannot support 'causal influence,' and the authors use that phrase in Section 5.4. A general positive evaluation of Allegra could inflate all item groups. The EFA showing three factors argues against one pure method factor, but not against shared method variance biasing the structural paths. And RQ1 is answered from absolute means (AC 4.58, BI 5.79) with no no-chat or reading-only baseline, so 'confirm' is too strong there too.\n\nThere are also minor but real reporting issues. RMSEA = 0.10 is conventionally poor, not 'good or very good.' The ANOVA p-values in Section 5.3 appear swapped (F=65.16 with p=0.0017, F=9.86 with p<2.2e-16 is backwards). Table 4 shows IM-BI correlation as 0.42 while the text says 0.54–0.58 for individual items; the EFA supports three factors but the model collapses all IM sub-scales into one, so the relative roles of trust vs. interest are not tested. These don't sink the paper, but they need fixing.\n\nThe core issue is proportion: the empirical contribution (a null on style, a clean protocol, a useful dataset) is smaller than the causal claims attached to it. The paper is revisable. Temper the abstract, add a discussion of common method bias or ideally a baseline condition, report effect sizes and equivalence bounds for the style null, correct the statistical errors.\n\nWho is this for? Practitioners designing health chatbots and researchers working on chatbot evaluation. It deserves a serious referee; it's not a desk reject. I'd send it out with a request for major revision, and the revision should be judged on whether the claims shrink to fit the design.\n\nRecommendation: engage with it.","headline":"A useful null result on conversational style is buried under causal language that the cross-sectional design cannot support.","tokens_in":15317,"tokens_out":3818,"would_cite":true,"duration_ms":32134,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that intrinsic motivation, not conversational style, drives both well-being awareness and intention to change in a scripted chatbot counselling chat.","keywords":["conversational agent","well-being","healthy lifestyle","intrinsic motivation","awareness creation","behavioural intention","chatbot","structural equation modelling"],"falsifier":"Run the same chat experiment, then check a month later whether participants actually did the healthy action they promised, and also ask how much they simply liked the chatbot; if the apparent effect of motivation shrinks once general liking is accounted for, or if motivated users do not follow through, the model's causal claim fails.","tokens_in":14338,"feed_emoji":"💬","tokens_out":8862,"duration_ms":87853,"temperature":0.7,"pith_summary":"This paper tests whether a scripted conversational counsellor can create awareness of healthy lifestyles and strengthen a user's intention to act, and which factors drive those outcomes. Analysing questionnaire responses from 600 participants who chatted with 'Allegra' in one of three conversational styles, it uses structural equation modelling to confirm that the intrinsic motivation factors of interest, value, trust and relatedness have a strong positive effect on awareness creation (path coefficient 0.93) and a substantial positive effect on behavioural intention (0.45), both significant at p<0.001. It reports no statistically significant effect of formal, informal, or multimedia-enriched informal conversational style on behavioural intention. If the claims hold, designers of well-being chatbot campaigns should concentrate on strengthening users' motivation rather than on fine-tuning the tone of the dialogue.","feed_headline":"Motivation, not chatbot style, drives well-being intent","feed_subtitle":"Interest, value, trust and relatedness predict awareness and intention; formal vs informal tone made no difference.","key_machinery":"The machinery is a three-part measurement and modelling chain. First, a scripted chat with a counsellor named Allegra follows a four-step coaching structure (Goal, Reality, Options, Will) and ends by asking the user whether they will follow a well-being suggestion in the coming month. Second, the experience is assessed with a 1-to-7 questionnaire whose items operationalise four intrinsic-motivation constructs from the Intrinsic Motivation Inventory — interest, value, trust, relatedness — together with awareness creation and behavioural intention; internal consistency is reported as high, with Cronbach's alpha 0.95 for the motivation items and 0.81 for each of the two outcome scales. Third, structural equation modelling, including exploratory and confirmatory factor analysis, turns those ratings into path coefficients connecting the latent factors, and the same model is re-fit across the three stylistic variants to test the style hypothesis.","core_discovery":"The paper's central claim is that a fully scripted, survey-like conversational agent presented as a well-being counsellor can create awareness and solicit behavioural intention, and that those outcomes are driven by intrinsic motivation rather than by the agent's linguistic style. After an exploratory factor analysis collapsed interest, value, trust and relatedness into a single intrinsic-motivation latent factor, a confirmatory factor analysis found strong paths from intrinsic motivation to awareness creation (0.93) and to behavioural intention (0.45), both at p<0.001, with about 63 percent of the variance in the two outcomes explained. Pairwise t-tests and a grouped confirmatory analysis found no significant differences among the three conversational-style conditions. The paper interprets this as support for its hypotheses H1 and H2, while noting that actual behaviour change was not measured and that longer or AI-driven interactions might behave differently.","pith_inferences":["A testable extension the paper leaves implicit: style effects may emerge in longer or repeated use, because interest and trust can decay or accumulate over multiple sessions; a longitudinal version of the same three-style comparison would give style a fairer test.","Because awareness creation carries almost all of the motivational influence (0.93), an inference beyond the paper is that manipulations aimed at behaviour should target awareness first; behavioural intention may then follow indirectly rather than being directly purchasable by content nudges.","The path coefficients rest on self-reports collected in the same session; a replication that separates the chatbot evaluation from a later behavioural follow-up, for example checking whether the promised walking actually happened, would tell whether the causal reading survives objective measurement."],"forward_implications":["If intrinsic motivation is the active ingredient, well-being chatbot designs should focus on features that raise interest, perceived value, trust, and relatedness, and evaluation checklists should measure those constructs rather than relying on style choices.","The null style result implies that, for a single scripted encounter, formal wording and informal wording with emojis and GIFs are roughly interchangeable in their effect on behavioural intention; decisions between them can be driven by audience preference rather than by expected conversion.","The 0.93 path from intrinsic motivation to awareness creation suggests that awareness gains are largely mediated by motivation, so simply adding more health information to a chatbot may not create awareness if the interaction does not engage the user.","High average behavioural intention scores, around 5.8 out of 7, after a roughly three-minute chat indicate that a single coaching conversation can shift stated intentions, but the paper's own scope stops at intention and does not establish lasting behaviour change."],"supporting_citations":[{"why":"Supplies the Self-Determination Theory basis for treating intrinsic motivation as the driver of engagement and well-being.","marker":"[10]"},{"why":"Provides the Intrinsic Motivation Inventory from which the four motivation constructs and their questionnaire items are taken.","marker":"[11]"},{"why":"Defines the four-step coaching structure (Goal, Reality, Options, Will) that organises the Allegra conversation.","marker":"[29]"},{"why":"Supplies the definition of behavioural intention as the perceived probability of future use, the main dependent variable.","marker":"[31]"},{"why":"Describes the conversational-survey toolkit used to build the fully scripted Allegra experiences.","marker":"[32]"},{"why":"Documents the crowdsourcing platform used to recruit the 600 participants across the three experimental campaigns.","marker":"[33]"},{"why":"Gives the exploratory factor analysis method that identifies the three latent factors in the data.","marker":"[35]"},{"why":"Gives the confirmatory factor analysis method used to estimate and test the path coefficients for H1 and H2.","marker":"[36]"}],"fun_headline_variants":["Motivation, not chatbot tone, drives well-being intent","Intrinsic motivation beats chatbot style for well-being","Well-being intent hinges on motivation, not talk style","Chatbot style irrelevant: motivation fuels well-being intent","For well-being chatbots, motivation outranks conversational flair"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that what people tick on the questionnaire genuinely measures their motivation, awareness, and intention, and that the correlations among those answers can be read as causes; if a vague liking for the chatbot pushed up every answer, the path coefficients would not prove a distinct motivational effect.","fun_headline_variants_meta":{"raw":{"variants":["Motivation, not chatbot tone, drives well-being intent","Intrinsic motivation beats chatbot style for well-being","Well-being intent hinges on motivation, not talk style","Chatbot style irrelevant: motivation fuels well-being intent","For well-being chatbots, motivation outranks conversational flair"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000742,"raw_usage":{"total_tokens":3334,"prompt_tokens":990,"completion_tokens":2344,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":606,"completion_tokens_details":{"reasoning_tokens":2267}},"tokens_in":606,"tokens_out":2344,"duration_ms":17551,"temperature":1.0,"reasoning_tokens":2267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:55:46.180846+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same chat experiment, then check a month later whether participants actually did the healthy action they promised, and also ask how much they simply liked the chatbot; if the apparent effect of motivation shrinks once general liking is accounted for, or if motivated users do not follow through, the model's causal claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Self-Determination Theory basis for treating intrinsic motivation as the driver of engagement and well-being."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Intrinsic Motivation Inventory from which the four motivation constructs and their questionnaire items are taken."},{"cited_title":"Whitmore, Coaching for Performance: The Principles and Practice of Coaching and Leadership fully revised 25th-anniversary edition, Hachette UK, 2010","cited_arxiv_id":null,"evidence_quote":"Defines the four-step coaching structure (Goal, Reality, Options, Will) that organises the Allegra conversation."},{"cited_title":"Palan, C","cited_arxiv_id":null,"evidence_quote":"Documents the crowdsourcing platform used to recruit the 600 participants across the three experimental campaigns."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the exploratory factor analysis method that identifies the three latent factors in the data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the confirmatory factor analysis method used to estimate and test the path coefficients for H1 and H2."}],"review_version":1}