{"id":"e2642ba6-d74a-4ef4-b764-780fabc02bfe","arxiv_id":"2502.07983","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Stakeholder evaluations of Welzijn.AI show experts see promise in reducing loneliness and detecting behavior changes, while elderly users value empathy and personality over transparency, and comprehension remains a challenge.","lead":"This paper describes a conversational AI system for monitoring elderly well-being and reports three small stakeholder studies: interviews with experts, a co-creation session, and feedback from 20 elderly users on a static interface. It is useful as a case study of how early user involvement shapes responsible AI design in elderly care.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Static proof-of-concept user evaluation cannot support design requirements for a conversational system: participants never interacted with Welzijn.AI's voice dialogue, so the empathetic/comprehensibility findings rest on an untested proxy.","rationale":"The paper's central claim is modest: three stakeholder evaluations disclose perspectives and reveal that integrating them is challenging. For that claim to carry weight, the proof-of-concept evaluation must tell us something about Welzijn.AI as a conversational system. The least secure condition is that a static visual mock-up preserves the properties being evaluated. This is not an internal inconsistency: the authors are explicit that the evaluation used a proof-of-concept, and their conclusions are appropriately hedged. But the specific quantitative findings used to motivate design requirements are only as valid as that proxy. The dynamic interaction is not an incidental detail; the system's intended mechanisms are spoken-language biomarkers, TTS voice, and LLM conversation, none of which is present in Figure 3. A within-subject comparison against the actual prototype would settle whether the proxy matters. Other limitations noted by the reader—small regional sample, no inter-rater protocol, no uncertainty intervals, no shared data, no ethics review—are real but do not attack the central inferential step as directly. Because the concern is a validity threat rather than a demonstrated failure, conditional acceptance remains the right verdict, and no change to the reader's judgment is needed.","tokens_in":10954,"tokens_out":7616,"duration_ms":78444,"concrete_test":"Run a within-subject replication with the same or matched 20 elderly participants: after the static condition, let each participant interact with the actual Welzijn.AI prototype (or a Wizard-of-Oz equivalent) in a scripted 10-minute conversation covering two EQ-5D-5L topics, then administer the identical accessibility, comprehensibility, trust, satisfaction, and human-likeness items plus the social-characteristic ranking. Pre-specify equivalence bounds, e.g., each characteristic's positive-perception proportion within ±10 percentage points and rank-sum differences for the top three social characteristics within ±10 points, with analysis accounting for within-participant clustering. If interactive estimates fall outside these bounds, the static mock-up is not a valid proxy and the design requirements are not established by the current data; if they fall inside, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's user-level findings in Sections 2.3 and 3.3, which supply the headline design requirements (empathetic and varying interactions; comprehension as an open issue), were obtained by showing 20 nursing-home residents a static proof-of-concept interface (Figure 3). Participants never spoke to or heard Welzijn.AI, although the intended system is a voice-controlled chatbot with speech transcription, LLM dialogue, and text-to-speech modules (Section 1, Figure 1). Tables 3 and 4 therefore record reactions to static screenshots and item wording, not to the conversational behavior being judged. Statements such as \"elderly experienced the system as natural, and more human than machine-like\" and the top ranking of \"responding empathetic\" presuppose that a picture preserves exactly the dynamic properties—empathy, human-likeness, trust, naturalness—that depend on interaction, timing, and voice. In addition, the percentages in Table 3 pool Likert responses across items and participants (e.g., 160 item-responses for accessibility from 20 participants) without participant-level clustering or uncertainty, so the apparent consensus is fragile. The central claim that stakeholder evaluation discloses stable design needs for a conversational AI therefore rests on an untested proxy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Welzijn.AI, a voice-based conversational AI system intended to monitor the well-being of elderly people through EQ-5D-5L-aligned dialogue and language biomarkers, and reports three stakeholder evaluations: semi-structured interviews with six experts summarized in a SWOT table, a co-creation session with four stakeholders using the Hundred Dollar Method to rank value requirements, and a proof-of-concept evaluation in which 20 elderly nursing-home residents were shown a static interface and answered Likert and ranking items. The authors claim that these evaluations disclose new perspectives on the system's strengths, weaknesses, design characteristics, and value requirements, notably that elderly users value empathetic and varied interactions, while comprehension and privacy remain open issues. The paper is framed as an illustration of responsible AI development following the CEHRES roadmap.","tokens_in":11161,"tokens_out":4254,"duration_ms":40187,"significance":"If the results are interpreted within their stated scope, the paper is a useful early-phase stakeholder-engagement case study for responsible conversational AI in elderly care. Its strengths include transparency about methods, a concrete link to the CEHRES roadmap, full item and use-case materials in the supplement, and a clear qualitative account of divergent stakeholder values. The paper makes no fitted-parameter or predictive claims, so circularity is not a concern. However, the headline user-perception results in Sections 2.3 and 3.3 are based on a static mock-up rather than the actual conversational system, and the reported percentages pool item responses across only 20 participants without uncertainty measures. These limitations materially affect the strength of the design-requirement conclusions and need to be addressed before the paper can support its current abstract and discussion claims.","major_comments":[{"comment":"The proof-of-concept evaluation used a static interface (Figure 3), but the system under study is a voice-controlled chatbot with speech transcription, LLM dialogue, and text-to-speech modules (Section 1, Figure 1). No participant spoke to or heard Welzijn.AI. Yet Section 3.3.1 states that \"Elderly experienced the system as natural, and more human than machine-like,\" and Table 4 ranks \"responding empathetic\" as the most important social characteristic. These findings concern properties of conversational behavior that cannot be validly assessed from static screenshots. Please reframe all user-perception and social-characteristics results as reactions to a visual mock-up, explicitly state that conversational properties were not evaluated, and soften the abstract and discussion conclusions that attribute empathy, naturalness, and human-likeness to the conversational system.","section":"§2.3, §3.3, Table 3"},{"comment":"The percentages in Table 3 are computed over item responses pooled across items and participants (e.g., accessibility has 160 item-responses from 20 participants, i.e., 8 items per participant), not over participants. Reporting \"64% of the elderly responses\" conflates items and participants and makes the apparent consensus fragile, since a single participant contributes multiple observations. Please reanalyze the data at the participant level, report per-participant proportions with confidence intervals or a comparable uncertainty measure, and present item-level results so readers can see which items drive each characteristic. Without this, the ordering of characteristics in Table 3 and the claim that \"elderly were divided among most topics\" are not quantitatively supported.","section":"§2.3, Table 3"},{"comment":"The manuscript states in Section 7 that no separate ethical approval was obtained, and Section 2 indicates only verbal consent from participants, who were nursing-home residents with a mean age of 83.2 years. Given that this is a vulnerable population and the study collected opinions on a health-related AI system, the paper should state the institutional or national policy under which this was exempt (e.g., exclusion from the Dutch WMO), or provide evidence of approval or waiver. As written, the ethical-governance disclosure is insufficient for a health-context study involving elderly participants, and this is a publication-blocking issue for many journals.","section":"§2, §7"},{"comment":"The co-creation session involved only four stakeholders, and \"consensus\" is defined as all stakeholders allocating more than zero dollars to a value requirement. With such a small group and small dollar pools, this is a very weak criterion; the footnote admits as much for U2 but the same logic applies to T7, T8, and E2. Please either pre-specify a consensus threshold, report the full per-stakeholder dollar allocations (Figure 4 must show these), or clearly label these as \"requirements that all stakeholders valued non-zero\" rather than consensus. The current wording overstates the level of agreement among stakeholder types.","section":"§3.2.1"}],"minor_comments":[{"comment":"Calling the 20 elderly participants an \"expert panel\" is confusing; \"elderly participants\" or \"user representatives\" would be clearer, since they are not experts in the same sense as the professional stakeholders.","section":"§2.3"},{"comment":"Please define \"Num. statements\" explicitly as the number of items times the number of participants, and add a column showing the number of participants who gave a positive response, not just the pooled item-response count.","section":"Table 3"},{"comment":"The phrase \"Elderly experienced the system as natural\" should be changed to \"Elderly rated the static interface as natural\" or similar, to avoid implying that the conversational system was experienced.","section":"§3.3.1"},{"comment":"The phrase \"non-elderly and elderly experts\" is awkward; consider \"expert interviewees and elderly participants\" or \"professional stakeholders and elderly users.\"","section":"Abstract"},{"comment":"The interviews were \"manually examined\" but the paper does not state whether this was a single-researcher thematic analysis or an independent coding process; please specify the analysis procedure and whether any inter-rater reliability check was performed.","section":"§2.1"},{"comment":"Figure 4 is referenced in the text but not visible in the manuscript body; ensure it is included in the final publication and that individual stakeholder dollar allocations are legible.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The ethical-approval statement in Section 7 is likely to draw scrutiny from the editor and may need to be resolved with institutional documentation before publication. In addition, the user-perception findings should be presented as mock-up evaluations, not as evaluations of a conversational system; this is a framing problem that can be fixed with careful revision, but it is central to the paper's headline claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a straightforward, honest case study of involving stakeholders in early development of a conversational AI for elderly well-being. What's new: the Welzijn.AI system itself and the concrete stakeholder findings—SWOT themes from six Dutch experts, HDM value rankings, Likert percentages and social-characteristic rankings from 20 nursing-home residents. None of that appears in prior literature. The authors use the CEHRES roadmap sensibly and are transparent about the design: they call these expert panels, not experiments, and they say no separate ethics review took place. That level of candor is rarer than it should be.\n\nThe paper does what it claims: it shows that different stakeholders raise complementary and conflicting requirements—empathy and personality from elderly users, privacy and clarity concerns from professionals, and context-level fixes (help desk, practice sessions) from the co-creation group. The qualitative and descriptive results are consistent with the data as reported. There is no fitted model, no circularity, no overclaimed prediction.\n\nThe soft spots are real but mostly proportional. The strongest concern is the static proof-of-concept: participants saw screenshots (Figure 3), never spoke to the system. So Table 4's top ranking of \"responding empathetic\" and Table 3's human-likeness percentages record reactions to a picture, not to conversational behavior. Empathy, naturalness, and trust are exactly the properties that depend on voice, timing, and interaction. The paper's conclusions are modest enough that this does not sink it, but those specific user-perception findings should be labeled as mock-up reactions, not prototype evaluations, and ideally re-tested with the working demo that now exists. Second, the percentages pool item responses across 20 participants (e.g., 160 accessibility responses) with no participant-level clustering or uncertainty, so the apparent consensus is fragile. The authors actually note that majorities were not large, which helps, but confidence intervals or raw counts by participant would be better. Third, the interview analysis is manual with no inter-rater protocol, and the \"little work embeds stakeholders\" claim in the introduction is a bit strong given [7] and related work. Fourth, no data or code are shared, though datasets are available on request; for a study selling reproducibility of process, a fuller supplement would serve the field.\n\nNet: this is a useful early-stage case study for researchers and designers working on responsible AI for elderly care. It deserves a serious referee; with revision focused on the mock-up limitation and the pooled percentages, it would be a solid contribution. I would cite it as a process case study.\n\nRecommendation: send to peer review, conditional on addressing the static-interface issue.","headline":"Honest, modest early-stage stakeholder study whose user-perception findings are weakened by the static mock-up, but the paper's framing mostly saves it; worth refereeing.","tokens_in":11658,"tokens_out":1954,"would_cite":true,"duration_ms":18700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Through three stakeholder evaluations, this paper argues that early, iterative involvement of experts and elderly users yields concrete design requirements for a conversational AI well-being monitor, and that comprehension and privacy…","keywords":["well-being","elderly care","artificial intelligence","remote monitoring","language biomarkers","responsible AI","large language models","conversational AI"],"falsifier":"A field test in which a working Welzijn.AI voice prototype is used daily for several weeks by a larger, more diverse elderly sample would settle the claim: if comprehension and satisfaction remain low after practice sessions, or if empathy no longer outranks transparency, the paper's design requirements would not transfer to real use.","tokens_in":10766,"feed_emoji":"💬","tokens_out":5962,"duration_ms":48466,"temperature":0.7,"pith_summary":"This paper is trying to show that a conversational AI system for monitoring elderly well-being can be developed responsibly by running small, early stakeholder evaluations instead of waiting for a finished product. Three panels—expert interviews, a co-creation session, and a proof-of-concept review by 20 elderly residents—produce a coherent but incomplete picture: stakeholders see the system as a way to reduce loneliness and detect shifts in behaviour, and they converge on concrete requirements such as safe data storage, a help desk, and practice sessions. The same evaluations expose persistent gaps: many elderly do not find the interface comprehensible or satisfying, and non-elderly experts worry about privacy and about what a warning signal should trigger. The paper's contribution is the demonstration that integrating all stakeholder perspectives is difficult but yields design and implementation guidance that a lab-only development process would miss.","feed_headline":"Elderly users want empathy, not transparency, from care chatbots","feed_subtitle":"Stakeholder panels set design needs: empathetic chats, practice sessions, help desk.","key_machinery":"The load-bearing mechanism is the staged, early-phase stakeholder evaluation loop: semi-structured expert interviews summarised as a SWOT table; a co-creation session in which stakeholders write use cases, allocate 100 dollars across them with the Hundred Dollar Method, extract core values, and rank value requirements; and a proof-of-concept static interface (a mock-up of the chatbot and a dashboard) evaluated by elderly residents through Likert and semantic-differential items plus a ranking of seven social characteristics. Together these methods convert stakeholder opinions into named value requirements and design characteristics that can feed the next development iteration.","core_discovery":"On its own terms, the paper establishes that three complementary stakeholder evaluations of an early Welzijn.AI concept disclose distinct and partly divergent perspectives on the system. Expert interviews identify strengths (combating loneliness, extracting behavioural patterns), weaknesses (unclear utility and signalling), opportunities (activating social networks, daily-task guidance), and threats (privacy, dependence, wrong conclusions from data). The co-creation session, using the Hundred Dollar Method, ranks value requirements and finds consensus on a gradual conversation flow, safe data storage, a help desk, and demo/test/practice sessions, while developer and caregiver priorities diverge. The proof-of-concept evaluation with 20 elderly residents shows majority-positive perceptions of accessibility, trust, and human-likeness, but not of comprehensibility or satisfaction, and ranks 'responding empathetically' first and 'using natural cues' last among desired social characteristics. The paper concludes that incorporating all stakeholder perspectives in system development remains challenging.","pith_inferences":["A testable extension would be to compare the static mock-up ratings with a working voice prototype, since the paper's elderly participants never spoke to the system and may have under- or overestimated its comprehensibility.","The finding that elderly rank empathy and personality over transparency suggests designers might satisfy privacy concerns through back-end data governance rather than through what the chatbot says, though the paper does not draw this conclusion.","The sharpest open question implied by the results is what a deviation signal should trigger; the paper records the disagreement but leaves the escalation protocol unspecified.","If the 20-person regional sample is representative, the absence of large majorities on most items hints at a heterogeneous elderly population split by tech experience, which would push development toward adaptable rather than one-size-fits-all interfaces."],"forward_implications":["If the findings hold, the first design priorities for such a system are empathetic and varied interaction, because elderly users ranked these above transparent behaviour and natural cues.","Deployment must include non-software support: education for caregivers and users, demo and test sessions, and a help desk, since comprehension is the weakest perceived characteristic.","Privacy and data-access agreements must be settled before rollout, as non-elderly experts consistently flagged them even while elderly users expressed trust.","The implementation context matters as much as the app: stakeholders expect the system to activate an individual's existing social network rather than replace human caregivers.","Iterative multi-stakeholder evaluation is a viable early-phase method for responsible AI in health, even though perspectives diverge and no single evaluation settles the design."],"supporting_citations":[{"why":"Supplies the five-phase roadmap for digital health development whose contextual-inquiry and value-specification phases structure the three stakeholder evaluations.","marker":"[15]"},{"why":"Supplies the Hundred Dollar Method used to rank value requirements in the co-creation session.","marker":"[19]"},{"why":"Defines the EQ-5D-5L domains (mobility, self-care, daily activities, discomfort/pain, mental health) around which Welzijn.AI's conversations are structured.","marker":"[8]"},{"why":"Provides the social-presence and acceptance items adapted into the proof-of-concept evaluation and the desired social characteristics.","marker":"[22]"},{"why":"Underlies the comprehensibility items that produced the paper's lowest-scoring user-perception characteristic.","marker":"[21]"}],"fun_headline_variants":["Elderly care AI needs empathy over transparency, panels say","Stakeholder input reveals empathy key for elderly AI chatbots","Divergent views on elderly AI: empathy vs privacy","Co-creation with elderly prioritizes empathy in AI design","Merging stakeholder views on elderly AI remains a challenge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that 20 elderly residents' reactions to a static proof-of-concept screen, in one region of the Netherlands, stand in for how the target population will perceive and use a working conversational system in daily life.","fun_headline_variants_meta":{"raw":{"variants":["Elderly care AI needs empathy over transparency, panels say","Stakeholder input reveals empathy key for elderly AI chatbots","Divergent views on elderly AI: empathy vs privacy","Co-creation with elderly prioritizes empathy in AI design","Merging stakeholder views on elderly AI remains a challenge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000655,"raw_usage":{"total_tokens":3043,"prompt_tokens":1033,"completion_tokens":2010,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1943}},"tokens_in":649,"tokens_out":2010,"duration_ms":13703,"temperature":1.0,"reasoning_tokens":1943,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:13:37.998410+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A field test in which a working Welzijn.AI voice prototype is used daily for several weeks by a larger, more diverse elderly sample would settle the claim: if comprehension and satisfaction remain low after practice sessions, or if empathy no longer outranks transparency, the paper's design requirements would not transfer to real use.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the five-phase roadmap for digital health development whose contextual-inquiry and value-specification phases structure the three stakeholder evaluations."},{"cited_title":"Leffingwell, D","cited_arxiv_id":null,"evidence_quote":"Supplies the Hundred Dollar Method used to rank value requirements in the co-creation session."},{"cited_title":"Brooks, E","cited_arxiv_id":null,"evidence_quote":"Defines the EQ-5D-5L domains (mobility, self-care, daily activities, discomfort/pain, mental health) around which Welzijn.AI's conversations are structured."},{"cited_title":"Heerink, B","cited_arxiv_id":null,"evidence_quote":"Provides the social-presence and acceptance items adapted into the proof-of-concept evaluation and the desired social characteristics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Underlies the comprehensibility items that produced the paper's lowest-scoring user-perception characteristic."}],"review_version":1}