REVIEW 4 major objections 6 minor 64 references
Welzijn.AI: Developing Responsible Conversational AI for Elderly Care through Stakeholder Involvement
T0 review · 4 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Through three stakeholder evaluations, this paper argues that early, iterative involvement of experts and elderly users yields concrete design requirements for a conversational AI well-being monitor, and that comprehension and privacy…
desk verdict Honest, modest early-stage stakeholder study whose user-perception findings are weakened by the static mock-up, but the paper's framing mostly saves it; worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the staged, early-phase stakeholder evaluation loop: semi-structured expert interviews summarised as a SWOT table; a co-creation session in which stakeholders write use cases, allocate 100 dollars across them with the Hundred Dollar Method, extract core values, and rank value requirements; and a proof-of-concept static interface (a mock-up of the chatbot and a dashboard) evaluated by elderly residents through Likert and semantic-differential items plus a ranking of seven social characteristics. Together these methods convert stakeholder opinions into named value requirements and design characteristics that can feed the next development iteration.
What would settle it
A field test in which a working Welzijn.AI voice prototype is used daily for several weeks by a larger, more diverse elderly sample would settle the claim: if comprehension and satisfaction remain low after practice sessions, or if empathy no longer outranks transparency, the paper's design requirements would not transfer to real use.
Extended reading notes
Core claim
On its own terms, the paper establishes that three complementary stakeholder evaluations of an early Welzijn.AI concept disclose distinct and partly divergent perspectives on the system. Expert interviews identify strengths (combating loneliness, extracting behavioural patterns), weaknesses (unclear utility and signalling), opportunities (activating social networks, daily-task guidance), and threats (privacy, dependence, wrong conclusions from data). The co-creation session, using the Hundred Dollar Method, ranks value requirements and finds consensus on a gradual conversation flow, safe data storage, a help desk, and demo/test/practice sessions, while developer and caregiver priorities diverge. The proof-of-concept evaluation with 20 elderly residents shows majority-positive perceptions of accessibility, trust, and human-likeness, but not of comprehensibility or satisfaction, and ranks 'responding empathetically' first and 'using natural cues' last among desired social characteristics. The paper concludes that incorporating all stakeholder perspectives in system development remains challenging.
Load-bearing premise
The study assumes that 20 elderly residents' reactions to a static proof-of-concept screen, in one region of the Netherlands, stand in for how the target population will perceive and use a working conversational system in daily life.
Editorial extensions
If this is right
- If the findings hold, the first design priorities for such a system are empathetic and varied interaction, because elderly users ranked these above transparent behaviour and natural cues.
- Deployment must include non-software support: education for caregivers and users, demo and test sessions, and a help desk, since comprehension is the weakest perceived characteristic.
- Privacy and data-access agreements must be settled before rollout, as non-elderly experts consistently flagged them even while elderly users expressed trust.
- The implementation context matters as much as the app: stakeholders expect the system to activate an individual's existing social network rather than replace human caregivers.
- Iterative multi-stakeholder evaluation is a viable early-phase method for responsible AI in health, even though perspectives diverge and no single evaluation settles the design.
Reading between the lines
- A testable extension would be to compare the static mock-up ratings with a working voice prototype, since the paper's elderly participants never spoke to the system and may have under- or overestimated its comprehensibility.
- The finding that elderly rank empathy and personality over transparency suggests designers might satisfy privacy concerns through back-end data governance rather than through what the chatbot says, though the paper does not draw this conclusion.
- The sharpest open question implied by the results is what a deviation signal should trigger; the paper records the disagreement but leaves the escalation protocol unspecified.
- If the 20-person regional sample is representative, the absence of large majorities on most items hints at a heterogeneous elderly population split by tech experience, which would push development toward adaptable rather than one-size-fits-all interfaces.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Welzijn.AI, a voice-based conversational AI system intended to monitor the well-being of elderly people through EQ-5D-5L-aligned dialogue and language biomarkers, and reports three stakeholder evaluations: semi-structured interviews with six experts summarized in a SWOT table, a co-creation session with four stakeholders using the Hundred Dollar Method to rank value requirements, and a proof-of-concept evaluation in which 20 elderly nursing-home residents were shown a static interface and answered Likert and ranking items. The authors claim that these evaluations disclose new perspectives on the system's strengths, weaknesses, design characteristics, and value requirements, notably that elderly users value empathetic and varied interactions, while comprehension and privacy remain open issues. The paper is framed as an illustration of responsible AI development following the CEHRES roadmap.
Significance. If the results are interpreted within their stated scope, the paper is a useful early-phase stakeholder-engagement case study for responsible conversational AI in elderly care. Its strengths include transparency about methods, a concrete link to the CEHRES roadmap, full item and use-case materials in the supplement, and a clear qualitative account of divergent stakeholder values. The paper makes no fitted-parameter or predictive claims, so circularity is not a concern. However, the headline user-perception results in Sections 2.3 and 3.3 are based on a static mock-up rather than the actual conversational system, and the reported percentages pool item responses across only 20 participants without uncertainty measures. These limitations materially affect the strength of the design-requirement conclusions and need to be addressed before the paper can support its current abstract and discussion claims.
major comments (4)
- [§2.3, §3.3, Table 3] The proof-of-concept evaluation used a static interface (Figure 3), but the system under study is a voice-controlled chatbot with speech transcription, LLM dialogue, and text-to-speech modules (Section 1, Figure 1). No participant spoke to or heard Welzijn.AI. Yet Section 3.3.1 states that "Elderly experienced the system as natural, and more human than machine-like," and Table 4 ranks "responding empathetic" as the most important social characteristic. These findings concern properties of conversational behavior that cannot be validly assessed from static screenshots. Please reframe all user-perception and social-characteristics results as reactions to a visual mock-up, explicitly state that conversational properties were not evaluated, and soften the abstract and discussion conclusions that attribute empathy, naturalness, and human-likeness to the conversational system.
- [§2.3, Table 3] The percentages in Table 3 are computed over item responses pooled across items and participants (e.g., accessibility has 160 item-responses from 20 participants, i.e., 8 items per participant), not over participants. Reporting "64% of the elderly responses" conflates items and participants and makes the apparent consensus fragile, since a single participant contributes multiple observations. Please reanalyze the data at the participant level, report per-participant proportions with confidence intervals or a comparable uncertainty measure, and present item-level results so readers can see which items drive each characteristic. Without this, the ordering of characteristics in Table 3 and the claim that "elderly were divided among most topics" are not quantitatively supported.
- [§2, §7] The manuscript states in Section 7 that no separate ethical approval was obtained, and Section 2 indicates only verbal consent from participants, who were nursing-home residents with a mean age of 83.2 years. Given that this is a vulnerable population and the study collected opinions on a health-related AI system, the paper should state the institutional or national policy under which this was exempt (e.g., exclusion from the Dutch WMO), or provide evidence of approval or waiver. As written, the ethical-governance disclosure is insufficient for a health-context study involving elderly participants, and this is a publication-blocking issue for many journals.
- [§3.2.1] The co-creation session involved only four stakeholders, and "consensus" is defined as all stakeholders allocating more than zero dollars to a value requirement. With such a small group and small dollar pools, this is a very weak criterion; the footnote admits as much for U2 but the same logic applies to T7, T8, and E2. Please either pre-specify a consensus threshold, report the full per-stakeholder dollar allocations (Figure 4 must show these), or clearly label these as "requirements that all stakeholders valued non-zero" rather than consensus. The current wording overstates the level of agreement among stakeholder types.
minor comments (6)
- [§2.3] Calling the 20 elderly participants an "expert panel" is confusing; "elderly participants" or "user representatives" would be clearer, since they are not experts in the same sense as the professional stakeholders.
- [Table 3] Please define "Num. statements" explicitly as the number of items times the number of participants, and add a column showing the number of participants who gave a positive response, not just the pooled item-response count.
- [§3.3.1] The phrase "Elderly experienced the system as natural" should be changed to "Elderly rated the static interface as natural" or similar, to avoid implying that the conversational system was experienced.
- [Abstract] The phrase "non-elderly and elderly experts" is awkward; consider "expert interviewees and elderly participants" or "professional stakeholders and elderly users."
- [§2.1] The interviews were "manually examined" but the paper does not state whether this was a single-researcher thematic analysis or an independent coding process; please specify the analysis procedure and whether any inter-rater reliability check was performed.
- [Figure 4] Figure 4 is referenced in the text but not visible in the manuscript body; ensure it is included in the final publication and that individual stakeholder dollar allocations are legible.
Circularity Check
No circularity: stakeholder evaluations are direct elicitations; the paper's conclusions restate the data rather than deriving them from assumptions or self-cited results.
full rationale
This paper reports three qualitative stakeholder evaluations (expert interviews, a co-creation session, and a proof-of-concept evaluation with elderly participants) and summarizes the elicited opinions. There are no fitted parameters, predictive equations, or derived quantities that reduce to the inputs by construction. The design requirements (e.g., empathetic and varied interaction, help desk, practice sessions, safe data storage, agreements on data access) are presented as direct findings from stakeholder responses, not as outputs of a model or as consequences of prior work. Citations to the authors' own prior work appear only as background context for language biomarkers and challenges in speech-based conversational AI, and they are not load-bearing for the paper's conclusions. The limitation that the proof-of-concept was a static interface rather than a working conversational system is a validity concern about proxy measures, not a circularity concern: the claims are explicitly about reactions to the presented proof-of-concept, and the paper does not disguise fitted inputs as predictions. No circular step can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption EQ-5D-5L dimensions capture relevant elderly well-being for monitoring
- domain assumption Language biomarkers such as pitch and topical coherence indicate mental well-being
- domain assumption CEHRES roadmap is an appropriate framework for responsible AI health development
- domain assumption Small expert panels disclose valid and sufficient stakeholder perspectives
- domain assumption Survey and ranking instruments adapted from prior literature measure the intended constructs
Cite this review
Pith. "Pith review of Welzijn.AI: Developing Responsible Conversational AI for Elderly Care through Stakeholder Involvement." pith.science (2026). https://pith.science/paper/4SZ2OO3X
@misc{pith2026250207983,
author = {Pith},
title = {Pith review of: Welzijn.AI: Developing Responsible Conversational AI for Elderly Care through Stakeholder Involvement},
year = {2026},
howpublished = {\url{https://pith.science/paper/4SZ2OO3X}},
note = {Machine review of arXiv:2502.07983}
}
read the original abstract
We present Welzijn.AI as new digital solution for monitoring (mental) well-being in elderly populations, and illustrate how development of systems like Welzijn.AI can align with guidelines on responsible AI development. Three evaluations with different stakeholders were designed to disclose new perspectives on the strengths, weaknesses, design characteristics, and value requirements of Welzijn.AI. Evaluations concerned expert panels and involved patient federations, general practitioners, researchers, and the elderly themselves. Panels concerned interviews, a co-creation session, and feedback on a proof-of-concept implementation. Interview results were summarized in terms of Welzijn.AI's strengths, weaknesses, opportunities and threats. The co-creation session ranked a variety of value requirements of Welzijn.AI with the Hundred Dollar Method. User evaluation comprised analysing proportions of (dis)agreement on statements targeting Welzijn.AI's design characteristics, and ranking desired social characteristics. Experts in the panel interviews acknowledged Welzijn.AI's potential to combat loneliness and extract patterns from elderly behaviour. The proof-of-concept evaluation complemented the design characteristics most appealing to the elderly to potentially achieve this: empathetic and varying interactions. Stakeholders also link the technology to the implementation context: it could help activate an individual's social network, but support should also be available to empower users. Yet, non-elderly and elderly experts also disclose challenges in properly understanding the application; non-elderly experts also highlight issues concerning privacy. In sum, incorporating all stakeholder perspectives in system development remains challenging. Still, our results benefit researchers, policy makers, and health professionals that aim to improve elderly care with technology.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[7]
B. H. Wolfe, Y. J. Oh, H. Choung, X. Cui, J. Weinzapfel, R. A. Cooper, H.-N. Lee, R. Lehto, Caregiving artificial intelligence chatbot for older adults and their preferences, well-being, and social connectivity: Mixed-method study, Journal of Medical Internet Research 27 (2025) e65776
work page 2025
-
[1]
Mental health of older adults, https://www.who
World Health Organisation. Mental health of older adults, https://www.who. int/news-room/fact-sheets/detail/mental-health-of-older-adults , (accessed 29 April 2024) (2023)
work page 2023
-
[2]
World Health Organisation. Social isolation and loneliness among older peo- ple, https://www.who.int/publications/i/item/9789240030749, (accessed 29 April 2024) (2021). 15
arXiv 2021
- [3]
-
[4]
European Union. Measures to tackle labour shortages: Lessons for future policy, https://www.eurofound.europa.eu/en/publications/2023/ measures-tackle-labour-shortages-lessons-future-policy , (accessed 26 March 2025) (2023)
work page 2023
-
[5]
A. H. Sapci, H. A. Sapci, Innovative assisted living tools, remote monitoring technologies, artificial intelligence-driven solutions, and robotic systems for ag- ing societies: systematic review, JMIR aging 2 (2) (2019) e15429
work page 2019
-
[6]
A Review of Challenges in Speech-based Conversational AI for Elderly Care
W. Klaassen, B. van Dijk, M. Spruit, A review of challenges in speech-based conversational ai for elderly care, arXiv preprint arXiv:2412.07388 (2024)
work page Pith review arXiv 2024
- [8]
Show all 64 references
-
[9]
Drougkas, E
G. Drougkas, E. M. Bakker, M. Spruit, Multimodal machine learning for lan- guage and speech markers identification in mental health, BMC Medical Infor- matics and Decision Making 24 (1) (2024) 354
2024
-
[10]
Figueroa-Barra, D
A. Figueroa-Barra, D. Del Aguila, M. Cerda, P. A. Gaspar, L. D. Terissi, M. Dur´ an, C. Valderrama, Automatic language analysis identifies and predicts schizophrenia in first-episode of psychosis, Schizophrenia 8 (1) (2022) 53
2022
-
[11]
Spruit, S
M. Spruit, S. Verkleij, K. de Schepper, F. Scheepers, Exploring language markers of mental health in psychiatric stories, Applied Sciences 12 (4) (2022) 2179
2022
-
[12]
Radford, J
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: International Conference on Machine Learning, PMLR, 2023, pp. 28492–28518. 16
2023
-
[13]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Let- man, A. Mathur, A. Schelten, A. Vaughan, et al., The llama 3 herd of models, arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[14]
J. Kim, J. Kong, J. Son, Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech, in: International Conference on Machine Learning, PMLR, 2021, pp. 5530–5540
2021
-
[15]
J. E. van Gemert-Pijnen, N. Nijland, M. van Limburg, H. C. Ossebaard, S. M. Kelders, G. Eysenbach, E. R. Seydel, A holistic framework to improve the uptake and impact of ehealth technologies, Journal of medical Internet research 13 (4) (2011) e1672
2011
-
[16]
Hagendorff, The ethics of ai ethics: An evaluation of guidelines, Minds and machines 30 (1) (2020) 99–120
T. Hagendorff, The ethics of ai ethics: An evaluation of guidelines, Minds and machines 30 (1) (2020) 99–120
2020
-
[17]
M. Nair, P. Svedberg, I. Larsson, J. M. Nygren, A comprehensive overview of barriers and strategies for ai implementation in healthcare: mixed-method design, Plos one 19 (8) (2024) e0305949
2024
-
[18]
Beecham, T
S. Beecham, T. Hall, C. Britton, M. Cottee, A. Rainer, Using an expert panel to validate a requirements process improvement model, Journal of Systems and Software 76 (3) (2005) 251–275
2005
-
[19]
Leffingwell, D
D. Leffingwell, D. Widrig, Managing software requirements: a unified approach, Addison-Wesley Professional, 2000
2000
-
[20]
D ´ ıaz-Bossini, L
J.-M. D ´ ıaz-Bossini, L. Moreno, Accessibility to mobile interfaces for older peo- ple, Procedia computer science 27 (2014) 57–66
2014
-
[21]
F. D. Davis, Perceived usefulness, perceived ease of use, and user acceptance of information technology, MIS quarterly (1989) 319–340
1989
-
[22]
Heerink, B
M. Heerink, B. Kr¨ ose, V. Evers, B. Wielinga, The influence of social presence on acceptance of a companion robot by older people, Journal of Physical Agents 2 (2) (2008) 33–40
2008
-
[23]
F. Meng, X. Guo, Z. Peng, Q. Ye, K.-H. Lai, Trust and elderly users’ continuance intention regarding mobile health services: the contingent role of health and technology anxieties, Information Technology & People 35 (1) (2022) 259–280. 17
2022
-
[24]
Gelderman, The relation between user satisfaction, usage of information systems and performance, Information & management 34 (1) (1998) 11–18
M. Gelderman, The relation between user satisfaction, usage of information systems and performance, Information & management 34 (1) (1998) 11–18
1998
-
[25]
M. Li, A. Suh, Machinelike or humanlike? a literature review of anthropomor- phism in ai-enabled technology, in: 54th Hawaii International Conference on System Sciences (HICSS 2021), 2021, pp. 4053–4062
2021
-
[26]
T. Fong, I. Nourbakhsh, K. Dautenhahn, A survey of socially interactive robots, Robotics and autonomous systems 42 (3-4) (2003) 143–166
2003
-
[27]
S. C. Mathews, M. J. McShea, C. L. Hanley, A. Ravitz, A. B. Labrique, A. B. Cohen, Digital health: a path to validation, NPJ digital medicine 2 (1) (2019) 38
2019
-
[28]
Sedhom, M
R. Sedhom, M. J. McShea, A. B. Cohen, J. A. Webster, S. C. Mathews, Mobile app validation: a digital health scorecard approach, NPJ Digital Medicine 4 (1) (2021) 111
2021
-
[29]
Zainal, N
A. Zainal, N. F. A. Aziz, N. A. Ahmad, F. H. A. Razak, F. Razali, N. H. Azmi, H. L. Koyou, Usability measures used to enhance user experience in using digital health technology among elderly: A systematic review, Bulletin of Electrical Engineering and Informatics 12 (3) (2023)...
2023
-
[30]
Text and icons are large enough
-
[31]
The ‘settings’ button is easy to spot
-
[32]
The functions of the icons on the screen are clear
-
[33]
The language used is easy to grasp
-
[34]
The contrasts between the background and text boxes is clear
-
[35]
I understand how I can respond with a message after the chatbot’s introduction
-
[36]
I understand where I have to press when I want to send a message
-
[37]
I understand where I need to press when I want messages to be read out loud Comprehensibility (5-point Likert scale):
-
[38]
The system looks needlessly complex
-
[39]
The system looks as if it is easy to use
-
[40]
I need the help of someone with technical knowledge if I want to use this system
-
[41]
I think I can master this system reasonably fast
-
[42]
I can see the point of this system
-
[43]
I think the system is hard to use
-
[44]
I think I can reliably use this system
-
[45]
I think that I will use this system frequently Intention to use (5-point Likert scale):
-
[46]
I intend to use this system
-
[47]
I am not planning to use this system
-
[48]
I would only use this system if the GP recommends it
-
[49]
I would recommend this system to others Perceived trust (5-point Likert scale):
-
[50]
I think that this system is reliable
-
[51]
I think that the technology is reliable in general
-
[52]
If this system is reliable, it will yield good results
-
[53]
I think this system is able to do what it says it will do
-
[54]
I do not trust this system
-
[55]
I think this system can address my needs when it comes to filling out surveys
-
[56]
I think that this system can work on its own without the assistance of humans
-
[57]
Overall I have trust in this system Satisfaction (5-point semantic differential scale):
-
[58]
terrible – wonderful
-
[59]
frustrating – satisfying
-
[60]
satisfactory – not satisfactory
-
[61]
boring – fun Human-likeness (5-point semantic differential scale):
-
[62]
machinelike – humanlike
-
[63]
not conscious – conscious
-
[64]
artificial – natural
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.