REVIEW 3 major objections 5 minor 35 references
A Conversational Approach to Well-being Awareness Creation and Behavioural Intention
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims that intrinsic motivation, not conversational style, drives both well-being awareness and intention to change in a scripted chatbot counselling chat.
desk verdict A useful null result on conversational style is buried under causal language that the cross-sectional design cannot support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a three-part measurement and modelling chain. First, a scripted chat with a counsellor named Allegra follows a four-step coaching structure (Goal, Reality, Options, Will) and ends by asking the user whether they will follow a well-being suggestion in the coming month. Second, the experience is assessed with a 1-to-7 questionnaire whose items operationalise four intrinsic-motivation constructs from the Intrinsic Motivation Inventory — interest, value, trust, relatedness — together with awareness creation and behavioural intention; internal consistency is reported as high, with Cronbach's alpha 0.95 for the motivation items and 0.81 for each of the two outcome scales. Third, structural equation modelling, including exploratory and confirmatory factor analysis, turns those ratings into path coefficients connecting the latent factors, and the same model is re-fit across the three stylistic variants to test the style hypothesis.
What would settle it
Run the same chat experiment, then check a month later whether participants actually did the healthy action they promised, and also ask how much they simply liked the chatbot; if the apparent effect of motivation shrinks once general liking is accounted for, or if motivated users do not follow through, the model's causal claim fails.
Extended reading notes
Core claim
The paper's central claim is that a fully scripted, survey-like conversational agent presented as a well-being counsellor can create awareness and solicit behavioural intention, and that those outcomes are driven by intrinsic motivation rather than by the agent's linguistic style. After an exploratory factor analysis collapsed interest, value, trust and relatedness into a single intrinsic-motivation latent factor, a confirmatory factor analysis found strong paths from intrinsic motivation to awareness creation (0.93) and to behavioural intention (0.45), both at p<0.001, with about 63 percent of the variance in the two outcomes explained. Pairwise t-tests and a grouped confirmatory analysis found no significant differences among the three conversational-style conditions. The paper interprets this as support for its hypotheses H1 and H2, while noting that actual behaviour change was not measured and that longer or AI-driven interactions might behave differently.
Load-bearing premise
The study assumes that what people tick on the questionnaire genuinely measures their motivation, awareness, and intention, and that the correlations among those answers can be read as causes; if a vague liking for the chatbot pushed up every answer, the path coefficients would not prove a distinct motivational effect.
Editorial extensions
If this is right
- If intrinsic motivation is the active ingredient, well-being chatbot designs should focus on features that raise interest, perceived value, trust, and relatedness, and evaluation checklists should measure those constructs rather than relying on style choices.
- The null style result implies that, for a single scripted encounter, formal wording and informal wording with emojis and GIFs are roughly interchangeable in their effect on behavioural intention; decisions between them can be driven by audience preference rather than by expected conversion.
- The 0.93 path from intrinsic motivation to awareness creation suggests that awareness gains are largely mediated by motivation, so simply adding more health information to a chatbot may not create awareness if the interaction does not engage the user.
- High average behavioural intention scores, around 5.8 out of 7, after a roughly three-minute chat indicate that a single coaching conversation can shift stated intentions, but the paper's own scope stops at intention and does not establish lasting behaviour change.
Reading between the lines
- A testable extension the paper leaves implicit: style effects may emerge in longer or repeated use, because interest and trust can decay or accumulate over multiple sessions; a longitudinal version of the same three-style comparison would give style a fairer test.
- Because awareness creation carries almost all of the motivational influence (0.93), an inference beyond the paper is that manipulations aimed at behaviour should target awareness first; behavioural intention may then follow indirectly rather than being directly purchasable by content nudges.
- The path coefficients rest on self-reports collected in the same session; a replication that separates the chatbot evaluation from a later behavioural follow-up, for example checking whether the promised walking actually happened, would tell whether the causal reading survives objective measurement.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a user study of a scripted conversational agent (Allegra) designed to promote well-being awareness and behavioural intention. Six hundred Prolific participants interacted with one of three versions differing in conversational style (informal, informal with multimedia elements, formal) and then completed a questionnaire measuring intrinsic motivation factors (interest, value, trust, relatedness), awareness creation (AC), and behavioural intention (BI). The authors use exploratory and confirmatory factor analyses (EFA/CFA) within a structural equation modelling framework to test H1 (intrinsic motivation influences BI) and H2 (intrinsic motivation influences AC), and they complement this with a sentiment analysis of free-text comments and a content analysis of conversation-topic choices. They report that intrinsic motivation has a strong positive path to AC (0.93) and a moderate path to BI (0.45), and that conversational style has no significant effect on behavioural intention. The abstract and conclusions present these results as evidence of causal influence.
Significance. If interpreted as correlational associations, this is a reasonably large (N=600) empirical contribution to the design of conversational well-being tools. Its strengths include a clearly described instrument based on established IMI constructs, a three-arm comparison of conversational styles, and an explicit null result for style effects, which is useful for the HCI literature. The paper does not ship code or data, but the methodological description is sufficiently detailed to permit replication. The main value is provisional and hypothesis-generating; the causal framing and the fit-statistic presentation currently exceed what the cross-sectional, single-source design can support.
major comments (3)
- [Section 5.4, Figure 5, H1/H2] The statement that the data 'confirm' H1 and H2 and that intrinsic motivation has a 'causal influence' is not supported by the design. All three construct families (intrinsic motivation, awareness creation, behavioural intention) were measured in one self-report questionnaire immediately after the chat, with no manipulation or temporal separation of the predictor; the SEM paths therefore estimate associations that can be inflated by common-method variance. The EFA's three-factor solution argues against a single method factor entirely explaining the data, but it does not rule out a shared method component that biases the structural paths. Please reword the claims as correlational, add an explicit common-method-bias limitation, and consider longitudinal or multi-source designs in future work.
- [Section 5.2, RQ1] The answer to RQ1 is based solely on absolute means (AC 4.58, BI 5.79) on a 1–7 scale, with no no-intervention or reading-only control condition. Mean levels above the scale midpoint do not demonstrate that the conversational approach 'influence[s]' awareness creation or behavioural intention; they only describe the participants' reported levels after exposure. Please reframe this as a descriptive finding and explicitly acknowledge the absence of a baseline.
- [Section 5.4, fit statistics] The fit indices are misreported. RMSEA = 0.10 is conventionally considered poor or at best borderline, not 'good or very good'; the text also contains 'RMSEA > 0.8', which appears to be a typo for 0.08 or 0.10. Because the authors use the fit statement to support the confirmatory conclusion, the fit indices should be reported accurately (e.g., CFI/TLI around 0.90 and RMSEA = 0.10 suggest marginal fit) and the implications for H1/H2 should be discussed rather than glossed over.
minor comments (5)
- [Table 4] The composite IM–BI correlation of 0.42 is inconsistent with the item-level correlations of 0.54–0.58 from which it is presumably derived; please clarify how the composite was computed, since a simple average of the four reported item correlations would be about 0.57.
- [Section 5.3] The ANOVA p-values appear swapped: F(1, 596) = 65.16 would correspond to an extremely small p-value, while F(1, 596) = 9.86 would correspond to a p around 0.002; please verify the reported values.
- [Table 2] Several items are phrased negatively (e.g., 'Allegra did not hold my attention at all') and it is not stated whether they were reverse-scored before computing Cronbach's alpha; the third behavioural-intention item is phrased as an open question rather than a 1–7 statement and should be aligned with the other items.
- [Section 5.4] The phrase 'may be a reason for RMSEA > 0.8' is unclear; also, 'Explanatory Factor Analysis' should be 'Exploratory Factor Analysis'.
- [Conclusions] The manuscript would benefit from a dedicated limitations subsection that explicitly lists the same-source design, lack of a baseline, no long-term follow-up, and possible demand effects from the chat setting.
Circularity Check
H1 is partially circular: two intrinsic-motivation items ask about future interaction with Allegra, which is the same construct the paper defines as behavioural intention.
-
self definitional
[Section 3 (RQ1 definition), Table 2 (questionnaire items), Section 5.4 (CFA/H1 confirmation)]
"behavioural intention, which represents the individual perceived probability that he/she will use the conversational tool in the future ... Table 2: Relatedness item 'I'd like a chance to interact with Allegra again'; Value item 'I would be willing to chat with Allegra again because it has some value to me' ... Section 5.4: 'the data shows that intrinsic motivation has a statistically significant and quite strong positive effect on both awareness creation (0.93) and behavioural intention (0.45).'"
The intrinsic-motivation latent variable is operationalized, in part, by items that directly ask about future interaction with and reuse of the conversational tool. That is the same construct the paper defines as behavioural intention ('perceived probability that he/she will use the conversational tool in the future'). When this IM latent is then used as the predictor of BI in the SEM, the estimated path is partly the outcome correlating with its own indicators. The EFA's three-factor grouping does not remove the wording overlap. The awareness-creation path is less affected, so the circularity is partial rather than total.
full rationale
The paper's central derivation is an empirical SEM fitted to questionnaire responses, not a mathematical derivation from first principles, so most of the analysis is not circular in the restricted sense used here. The same-instrument, cross-sectional design creates common-method-variance and causal-inference concerns, but those are validity threats rather than definitional circularity. The one concrete circular element is the item-content overlap: the IM factor includes items about wanting to interact with Allegra again, while BI is defined as the perceived probability of using the tool in the future; the IM-to-BI path therefore has a self-definitional component. The self-citation to the Coney toolkit [32] is a normal tool citation and is not load-bearing for any theoretical claim, and no uniqueness theorem or ansatz is imported from the authors' prior work. The RQ1 'confirmation' from absolute mean scores without a baseline is an unsupported causal inference, but not a circularity. Overall, the H1 result is partially forced by item wording, while H2 and the conversational-style findings retain independent empirical content.
Assumptions & free parameters
free parameters (2)
- Path coefficient IM to awareness creation =
0.93
- Path coefficient IM to behavioural intention =
0.45
assumptions (4)
- domain assumption Self-report questionnaire items validly measure the latent constructs without common-method bias.
- domain assumption The three conversational conditions implement meaningfully different styles.
- domain assumption Prolific participants are representative enough for the study's claims.
- domain assumption Behavioural intention self-reports relate to actual future behaviour.
Cite this review
Pith. "Pith review of A Conversational Approach to Well-being Awareness Creation and Behavioural Intention." pith.science (2026). https://pith.science/paper/CQYW4KII
@misc{pith2026250421702,
author = {Pith},
title = {Pith review of: A Conversational Approach to Well-being Awareness Creation and Behavioural Intention},
year = {2026},
howpublished = {\url{https://pith.science/paper/CQYW4KII}},
note = {Machine review of arXiv:2504.21702}
}
read the original abstract
The promotion of a healthy lifestyle is one of the main drivers of an individual's overall physical and psycho-emotional well-being. Digital technologies are more and more adopted as ''facilitators'' for this goal, to raise awareness and solicit healthy lifestyle habits. This study aims to experiment the effects of the adoption of a digital conversational tool to influence awareness creation and behavioural change in the context of a well-being lifestyle. Our aim is to collect evidence of the aspects that must be taken into account when designing and implementing such tools in well-being promotion campaigns. To this end, we created a conversational application for promoting well-being and healthy lifestyles, which presents relevant information and asks specific questions to its intended users within an interaction happening through a chat interface; the conversational tool presents itself as a well-being counsellor named Allegra and follows a coaching approach to structure the interaction with the user. In our user study, participants were asked to first interact with Allegra in one of three experimental conditions, corresponding to different conversational styles; then, they answered a questionnaire about their experience. The questionnaire items were related to intrinsic motivation factors as well as awareness creation and behavioural change. The collected data allowed us to assess the hypotheses of our model that put in connection those variables. Our results confirm the positive effect of intrinsic motivation factors on both awareness creation and behavioural intention in the context of well-being and healthy lifestyle; on the other hand, we did not record any statistically significant effect of different language and communication styles on the outcomes.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
J. van Olmen, The promise of digital self-management: A reflection about the effects of patient-targeted e-health tools on self-management and well- being, International Journal of Environmental Research and Public Health 19 (3) (2022)
work page 2022
- [2]
-
[3]
G. Laban, T. Araujo, Working together with conversational agents: The relationship of perceived cooperation with service performance evaluations, in: A. Følstad, T. Araujo, S. Papadopoulos, E. L.-C. Law, O.-C. Granmo, E. Luger, P. B. Brandtzaeg (Eds.), Chatbot Research and Design, Springer International Publishing, 2020, pp. 215–228
work page 2020
- [4]
-
[5]
A. B. Kocaballi, L. Laranjo, L. Clark, R. Kocielnik, R. J. Moore, Q. V. Liao, T. Bickmore, Special issue on conversational agents for healthcare and wellbeing, ACM Transactions on Interactive Intelligent Systems (TiiS) 12 (2) (2022) 1–3
work page 2022
-
[6]
Z. Callejas, D. Griol, Conversational Agents for Mental Health and Well- being, Springer International Publishing, 2021, Ch. 11, pp. 219–244
work page 2021
-
[7]
L. Tudor Car, D. A. Dhinagaran, B. M. Kyaw, T. Kowatsch, S. Joty, Y.-L. Theng, R. Atun, Conversational agents in health care: Scoping review and conceptual analysis, Journal of Medical Internet Research 22 (8) (2020) e17158–1:21
work page 2020
-
[8]
C. Liebrecht, L. Sander, C. van Hooijdonk, Too informal? how a chatbot’s communication style affects brand attitude and quality of interaction, in: A. Følstad, T. Araujo, S. Papadopoulos, E. L.-C. Law, E. Luger, M. Good- win, P. B. Brandtzaeg (Eds.), Chatbot Research and Design, Springer In- ternational Publishing, Cham, 2021, pp. 16–31. 21
work page 2021
Show all 35 references
-
[9]
Peters, R
D. Peters, R. A. Calvo, R. M. Ryan, Designing for motivation, engagement and wellbeing in digital experience, Frontiers in Psychology 9 (2018)
2018
-
[10]
E. L. Deci, R. M. Ryan, Self-determination theory, Handbook of theories of social psychology 1 (20) (2012) 416–436
2012
-
[11]
R. M. Ryan, V. Mims, R. Koestner, Relation of reward contingency and interpersonal context to intrinsic motivation: A review and test using cogni- tive evaluation theory., Journal of personality and Social Psychology 45 (4) (1983) 736
1983
-
[12]
K. S. Ostrow, N. T. Heffernan, Testing the validity and reliability of intrin- sic motivation inventory subscales within assistments, in: Artificial Intel- ligence in Education: 19th International Conference, AIED 2018, London, UK, June 27–30, 2018, Proceedings, Part I 19, Spr...
2018
-
[13]
Ehsan, Q
U. Ehsan, Q. V. Liao, M. Muller, M. O. Riedl, J. D. Weisz, Expanding explainability: Towards social transparency in ai systems, in: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, CHI ’21, Association for Computing Machinery, New York, NY, USA, 20...
2021
-
[14]
Q. V. Liao, D. Gruen, S. Miller, Questioning the ai: informing design practices for explainable ai user experiences, in: Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–15
2020
-
[16]
Hsiao, H
J. Hsiao, H. H. T. Ngai, L. Qiu, Y. Yang, C. C. Cao, Roadmap of de- signing cognitive metrics for explainable artificial intelligence (xai), ArXiv abs/2108.01737 (2021)
2021 arXiv
-
[17]
W.-P. Brinkman, Virtual health agents for behavior change: Research per- spectives and directions, in: Proceedings of the Workshop on Graphical and Robotic Embodied Agents for Therapeutic Systems, 2016, pp. 1–17
2016
-
[18]
Dingler, D
T. Dingler, D. Kwasnicka, J. Wei, E. Gong, B. Oldenburg, The use and promise of conversational agents in digital health, Yearbook of Medical Informatics 30 (01) (2021) 191–199
2021
-
[19]
Følstad, M
A. Følstad, M. Skjuve, Chatbots for customer service: User experience and motivation, in: Proceedings of the 1st International Conference on Conver- sational User Interfaces, CUI ’19, Association for Computing Machinery, New York, NY, USA, 2019, pp. 1–9. 22
2019
-
[20]
Fadhil, G
A. Fadhil, G. Schiavo, Y. Wang, B. A. Yilma, The effect of emojis when interacting with conversational interface assisted health coaching system, in: Proceedings of the 12th EAI International Conference on Pervasive Computing Technologies for Healthcare, PervasiveHealth ’18, A...
2018
-
[21]
Peltola, K
J. Peltola, K. Kaipainen, K. Keinonen, N. Kiuru, M. Turunen, Developing a conversational interface for an act-based online program: Understanding adolescents’ expectations of conversational style, in: Proceedings of the 5th International Conference on Conversational User Inter...
2023
-
[22]
Silva, E
G. Silva, E. Canedo, Towards user-centric guidelines for chatbot conversa- tional design, International Journal of Human-Computer Interaction (2022) 1–23
2022
-
[23]
Luger, A
E. Luger, A. Sellen, ”like having a really bad pa” the gulf between user expectation and experience of conversational agents, in: Proceedings of the 2016 CHI conference on human factors in computing systems, Association for Computing Machinery, New York, NY, USA, 2016, pp. 5286–5297
2016
-
[24]
V. Lai, C. Chen, Q. V. Liao, A. Smith-Renner, C. Tan, Towards a science of human-ai decision making: a survey of empirical studies, arXiv preprint arXiv:2112.11471 (2021)
2021 arXiv
-
[25]
Miller, Explanation in artificial intelligence: Insights from the social sciences, Artificial Intelligence 267 (06 2017)
T. Miller, Explanation in artificial intelligence: Insights from the social sciences, Artificial Intelligence 267 (06 2017). doi:10.1016/j.artint. 2018.07.007
2017 doi
-
[26]
T. C. Leonard, Richard H. Thaler, Cass R. Sunstein, Nudge: Improving decisions about health, wealth, and happiness: Yale University Press, New Haven, CT, 2008, 293 pp, $26.00, Constitutional Political Economy 19 (4) (2008) 356–360. doi:10.1007/s10602-008-9056-2 . URL http://li...
2008 doi
-
[27]
Caraban, E
A. Caraban, E. Karapanos, D. Gon¸ calves, P. Campos, 23 ways to nudge: A review of technology-mediated nudging in human-computer interaction, in: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, CHI ’19, Association for Computing Machinery, New Yor...
2019
-
[28]
Leeuwis, L
L. Leeuwis, L. He, Hi, i’m cecil(y) the smoking cessation chatbot: The effec- tiveness of motivational interviewing and confrontational counseling chat- bots and the moderating role of the need for autonomy and self-efficacy, in: A. Følstad, T. Araujo, S. Papadopoulos, E. L.-C...
2022
-
[29]
Whitmore, Coaching for Performance: The Principles and Practice of Coaching and Leadership fully revised 25th-anniversary edition, Hachette UK, 2010
J. Whitmore, Coaching for Performance: The Principles and Practice of Coaching and Leadership fully revised 25th-anniversary edition, Hachette UK, 2010
2010
-
[30]
Panchal, P
S. Panchal, P. Riddell, The grows model: extending the grow coaching model to support behavioural change, International Journal of The Coach- ing Psychologist 16 (2) (2020)
2020
-
[31]
F. D. Davis, Perceived usefulness, perceived ease of use, and user acceptance of information technology, MIS quarterly (1989) 319–340
1989
-
[32]
Celino, G
I. Celino, G. Re Calegari, Submitting surveys via a conversational interface: an evaluation of user acceptance and approach effectiveness, International Journal of Human-Computer Studies 139 (2020) 102410. doi:10.1016/j. ijhcs.2020.102410
2020
-
[33]
Palan, C
S. Palan, C. Schitter, Prolific.ac — a subject pool for online experiments, Journal of Behavioral and Experimental Finance 17 (2018) 22–27
2018
-
[34]
J. D. Brown, The cronbach alpha reliability estimate, JALT Testing & Evaluation SIG Newsletter 6 (1) (2002)
2002
-
[35]
C. D. Stapleton, Basic Concepts in Exploratory Factor Analysis (EFA) as a Tool To Evaluate Score Validity: A Right-Brained Approach., ERIC, 1997
1997
-
[36]
T. A. Brown, Confirmatory factor analysis for applied research, Guilford publications, 2015. 24
2015
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.