{"id":"57db7cad-7a5d-43de-b879-8762fb7ddd41","arxiv_id":"2509.02537","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A usability study with 30 children with congenital heart disease guided the redesign of a metaphor-based game app intended to teach young patients about their condition, though learning gains were not measured.","lead":"This paper describes how 30 children with congenital heart disease helped test and refine a cartoon app called Octo's Heartland that teaches heart anatomy and healthy habits through play. It shows a practical path for involving young patients in designing health education tools, though it does not yet prove the app improves learning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Comprehension outcomes are asserted without any knowledge measure; observed engagement and short verbal answers cannot support the claimed health-literacy and comprehension findings.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing concern: observed play behavior and short verbal answers in single 15-minute sessions are treated as evidence of comprehension, despite the absence of any knowledge measure. The paper's own Section 4.2 undermines this inference by reporting that children struggled to connect metaphors to medical content and that one child explicitly disclaimed knowledge. The abstract's 'comprehension outcomes' is therefore an overstatement that is internally inconsistent with the reported evidence. This is not a mere stylistic issue; the design implications about metaphor-based serious games supporting health literacy depend on understanding, not just engagement. If the concern lands, the paper's central contribution narrows to a usability and engagement study, which is still valuable for a CHD-specific serious game but should not be described as demonstrating comprehension. Given that the design process itself is well-documented, IRB-approved, and grounded in child-engaged methods, the appropriate response is to temper the comprehension claims and add a knowledge measure in the next phase, which is exactly the reader's conditional recommendation. No new concern changes that verdict, so UNCHANGED is appropriate.","tokens_in":15419,"tokens_out":2727,"duration_ms":27478,"concrete_test":"Re-score the existing video and transcript data from the February 2025 usability sessions with a minimal comprehension rubric: for each child, code whether they (a) name the CHD condition associated with an activity, (b) explain in their own words what the central metaphor represents (e.g., the river is the aorta, the narrowing is coarctation), and (c) answer one factual question about the condition immediately after play. Report the proportion of children meeting each criterion, stratified by target age (4–10) vs. older participants. If fewer than half of target-age children can map metaphor to medical concept, the 'comprehension outcomes' claim in the abstract fails and the design implications must be reframed as engagement findings only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central design claim is that child-engaged usability testing produced design improvements that support comprehension and health literacy. RQ2 explicitly asks how insights can be translated into design improvements that 'enhance health literacy.' Yet the evidence for comprehension is thin. Section 4.1 says the team 'observed which activities they chose, usability, moments of excitement and frustration, and CHD content comprehension,' but no comprehension instrument, pre/post test, or task-performance metric is described anywhere in the protocol. The only data are 15-minute sessions, activity-choice observations, affective reactions, and short answers to open-ended questions. Section 4.2 then supplies direct counter-evidence: a participant said 'I don't know much about anything' during the matching game, and 'some children struggled to connect them to medical content' for the nature metaphors. These admissions show that observed engagement does not entail understanding. The abstract's phrase 'comprehension outcomes' and the discussion's claim that reactions were 'used to evaluate usability, emotional engagement, and conceptual understanding' overstate what the method can support. If engagement does not track learning, the design implications—especially the claim that metaphor-based activities make medical content 'more digestible'—may optimize enjoyment while leaving the health-literacy goal unsubstantiated. The paper itself flags this gap in Section 4.2, making the abstract's assertion internally inconsistent with the reported findings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an iterative design and usability testing study for \"Octo's Heartland,\" a metaphor-based digital application intended to support health literacy, engagement, and comprehension for children with congenital heart disease (CHD) aged 4–10. The work is part of a multi-phase participatory design project that also includes a plush toy. In a preliminary study at Camp Odayin, 12 children interacted with an earlier prototype and provided open-ended feedback; in a subsequent usability test, 18 children explored four newly designed activities with different narrative metaphors (Heart Island, Train, Cardia Kingdom, Matching Game). Thematic analysis of observations and transcripts informed design criteria for usability, educational, and emotional outcomes, and led to a final design organized as a train ride through a heart-shaped island with five stops representing four CHD conditions. The paper claims findings about usability, engagement, and comprehension, and proposes serious game (SG) design principles for supporting health literacy in young children with CHD.","tokens_in":15722,"tokens_out":3121,"duration_ms":29540,"significance":"If the central claims hold, the paper makes a useful contribution to participatory design and serious game research for a population rarely included in design and evaluation: young children with CHD. The strengths are concrete and should be credited: the study was conducted with IRB approval in a medically supported camp, engaged children as young as 4 in open-ended exploration rather than task-based protocols, involved 30 children across two phases, and translated qualitative insights into a detailed, coherent final design that is clearly grounded in the four SG principles cited from Kucher. The paper also addresses a genuine gap in child-focused CHD education. However, the significance is weakened by the overclaimed \"comprehension outcomes\" and by a usability-testing sample that includes many children outside the stated target age range. The design recommendations may be plausible, but the evidence as presented does not support knowledge-acquisition or health-literacy claims without further measurement.","major_comments":[{"comment":"The abstract states that findings highlight \"usability, engagement, and comprehension outcomes,\" and §4.1 says the team observed \"CHD content comprehension,\" but no comprehension instrument, knowledge measure, pre/post test, or task-performance metric is described anywhere in the protocol. The only data are 15-minute sessions, activity-choice observations, affective reactions, and short verbal answers. In §4.2 the paper itself supplies counter-evidence: a participant said \"I don't know much about anything\" during the matching game, and \"some children struggled to connect\" nature metaphors to medical content. Observed engagement and brief verbal responses cannot support the claimed comprehension or health-literacy outcomes. Please either add a real knowledge measure (e.g., a short pre/post quiz, parent or clinician rating, or task-based comprehension check) or reframe the claims to \"usability, engagement, and exploratory indications of understanding\" and explicitly report comprehension as anecdotal rather than measured.","section":"Abstract; §4.1"},{"comment":"The usability-testing sample is described as \"Eighteen children with CHD participated (ages 6–16)\", yet the paper's target population is ages 4–10 and §3.1 states the study focus is ages 4–10. Only 10 of the 18 participants are within the target range, and eight are aged 11–16, including a 16-year-old. Design implications drawn from the full sample are then applied to the final design for 4–10-year-olds without any age-based breakdown of which feedback came from which age group. Please report results separately for the target-age subgroup versus older participants, and justify how feedback from adolescents informs design decisions for younger children, or restrict the design implications to the ages actually represented.","section":"§4.1; Table 1"},{"comment":"Thematic coding is described as handwritten notes and transcripts \"coded sentence by sentence\" and organized in two stages, but no coding reliability or trustworthiness information is provided: there is no second coder, no inter-rater reliability or consensus process, no codebook, and no audit trail beyond the list of themes. For a qualitative design study, this is acceptable if the analysis is framed as the research team's interpretive synthesis to guide redesign, but the paper presents the coded themes as findings (e.g., \"Usability findings:\", \"Educational and Emotional findings:\"). Please state the coding procedure transparently (who coded, how disagreements were resolved, whether the codebook evolved) or reframe the results as designer-led analysis rather than independent qualitative findings.","section":"§3.2; §4.2"},{"comment":"The link between the collected data and the final design criteria is asserted but not demonstrated. Section 5.1 lists \"Design Criteria Based on Types of CHD\" and §5.2 describes the final features, but there is no traceable mapping from specific participant observations or quotes to specific criteria or design decisions. For example, the claim that \"problem-solving based activities like the train puzzle were the most appealing\" is stated in §4.2, but no frequency counts, illustrative quotes, or comparison across the four activities are provided to show why it was most appealing or how that led to the train-based navigation in the final design. Given that RQ2 asks how insights are translated into design improvements, the authors should provide a small traceability table or at least explicitly cite representative data points that support each final design criterion.","section":"§5.1; §5.2"}],"minor_comments":[{"comment":"The phrase \"comprehension outcomes\" also appears in the abstract's last sentence; even if the major measurement concern is addressed, the wording should be softened to avoid implying a formal learning assessment was performed.","section":"Abstract; §4.2"},{"comment":"The sentence \"The testing protocol followed principles from Andersen et al. [4], accommodating short attention spans (<30 min)\" is slightly under-specified because the sessions themselves are described as \"around 15 min\" later in the same paragraph; please state the planned session duration explicitly and consistently.","section":"§3.1"},{"comment":"The final design section would benefit from at least one figure per metaphor world (Pulmonary Pathway, Aorta River, Ventricle Village, Atria Leak) to help readers assess the claimed metaphor–medical content alignment; currently Figures 3–6 show only the home screen, learning cards, and two of the worlds in partial views.","section":"§4.2; Figure 3–6"},{"comment":"The paper uses \"Kucher [37]\" for the four SG principles, but the reference list entry does not provide a full first name or publisher information; please complete the bibliographic details to match ACM style.","section":"§2.2; §4.2"},{"comment":"The final design description mixes future tense (\"children begin,\" \"players select\") with present tense (\"The Aorta River mirrors,\" \"Users earn stars\"); please unify the tense for clarity.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a legitimate design-study contribution with a real user population, and the iterative design work appears genuinely valuable. The main issue is overclaiming: the term \"comprehension outcomes\" in the abstract and the discussion of \"conceptual understanding\" in §6 go beyond what the qualitative protocol can support. The authors could fix this either by adding a lightweight knowledge check (which would strengthen the paper substantially) or by carefully reframing all claims as engagement and usability findings with exploratory observations about understanding. The age-range mismatch in the usability test sample is also load-bearing and should be addressed explicitly rather than glossed over. For a CSCW Companion paper, I would not recommend rejection, but the revision must address the comprehension claim and the sample-age issue before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, carefully reported usability study of a CHD-specific serious game for young children. The actual value is the iterative design process and the design criteria that come out of the child sessions, not any demonstrated learning outcome. The abstract says 'comprehension outcomes' but no knowledge measure was used; the paper's own Section 4.2 shows children struggling to connect metaphors to the medical content. So read it as a design paper, not an efficacy paper.\n\nWhat's genuinely new: an app specifically for children with CHD aged 4-10, tested with 30 kids at a medically supported camp, and a clear set of design implications mapped onto Kucher's game-based learning principles. That fills a gap the authors' scoping review identifies. The qualitative methods are appropriate for this age group: observation over think-aloud, video coding, thematic analysis. The redesign decisions follow from the data; the four activities are systematically varied, and the final design is described in enough detail to reproduce.\n\nThe soft spots: the comprehension claim is not supported. The protocol records activity choice, moments of excitement/frustration, and short verbal answers, but there is no pre/post test, no task performance metric, no measure of health literacy. The discussion says reactions were 'used to evaluate ... conceptual understanding', which is a stretch. And the paper's own findings undercut the idea that engagement tracks understanding: one child says 'I don't know much about anything' during the matching game, and some children could not connect the island metaphors to the medical content. That is useful negative evidence, but it means the abstract overstates what the study shows. Also, the coding reliability is not reported (single coder, hand-written notes), and there is no baseline or comparison condition, so any claim about effectiveness is only directional.\n\nProportionally, this is a moderate problem, not a fatal one. The design contribution stands on its own; the 'comprehension outcomes' phrasing should be changed to something like 'observed understanding' or 'comprehension-related behaviors' in the next version.\n\nI'd bring this to a reading group for the method alone, and I'd cite it for the design criteria. A serious referee could ask for the comprehension language to be toned down, but the paper deserves review.","headline":"A useful, honest usability study of a CHD education app; the design refinement is credible, but the 'comprehension outcomes' claim in the abstract outruns the data.","tokens_in":16148,"tokens_out":2349,"would_cite":true,"duration_ms":18767,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that observation-first usability testing with 30 children with congenital heart disease produced the concrete design criteria behind Octo's Heartland, a metaphor-based serious game that teaches heart anatomy and healthy…","keywords":["congenital heart disease","serious games","pediatric health education","usability testing with children","participatory design","metaphor-based learning","health literacy","iterative design"],"falsifier":"Give two comparable groups of children with CHD the redesigned Octo's Heartland and a non-game picture booklet covering the same four conditions, then measure heart knowledge before and after; if the game group shows no larger gain, the claim that metaphor-based gameplay improves comprehension is not supported.","tokens_in":15196,"feed_emoji":"🩺","tokens_out":5409,"duration_ms":48335,"temperature":0.7,"pith_summary":"Children with congenital heart disease (CHD) rarely get educational materials designed for them, and younger children are usually left out of usability research. This paper argues that they can be included: two rounds of testing with 30 children at a medically supported camp, using observation and open-ended play rather than think-aloud tasks, produced a clear set of design changes. Those changes shaped Octo's Heartland, a serious game in which a heart-shaped island, a train, and metaphor-based activities teach cardiac anatomy, four CHD conditions, and healthy habits. The paper's contribution is a worked example of translating child reactions, verbal and non-verbal, into usability and educational design criteria, plus a description of the redesigned app ready for clinical and home testing.","feed_headline":"Child-led testing reshapes a heart-disease learning app for ages 4-10","feed_subtitle":"The Octo's Heartland app turns heart anatomy into train, river, and house puzzles guided by 30 children's reactions.","key_machinery":"The central mechanism is metaphor-based gameplay organized around a heart-shaped island traversed by train, with each CHD condition mapped to a physical metaphor children can manipulate: pulmonary stenosis as a narrowed train track to widen, coarctation of the aorta as a blocked river to clear, and septal defects as openings in a house wall to patch. The evaluation and redesign are carried by four game-based learning principles from prior work—meaningful interactivity, immersiveness, effective feedback, and freedom of exploration—which the paper uses both as a lens for coding observations and as the source of the final design criteria.","core_discovery":"The paper's central claim is that child-engaged usability testing, not expert opinion, should drive the design of pediatric health-education technology, and that this is feasible even for children aged 4-10 with complex medical conditions. In the reported sessions, children's navigation struggles, pacing preferences, excitement, and confusion were coded into themes and then turned into four design criteria: intuitive navigation, flexible pacing, accessible design, and storytelling or problem-solving that keeps an uplifting tone. The resulting application presents four CHD conditions—ventricular septal defect, pulmonary stenosis, coarctation of the aorta, and atrial septal defect—each as a three-level metaphor-based journey, such as widening a train track, clearing a river, or patching a house wall. The paper reports that problem-solving activities were the most appealing, nature metaphors were the best understood, and children wanted more guidance and rewards, all of which are now reflected in the final activity structure.","pith_inferences":["Because some children struggled to connect the metaphors to their medical meaning, a natural next test is to add explicit verbal narration that states what each metaphor means for the child's own heart, and compare comprehension against the current version.","The same design criteria—intuitive navigation, flexible pacing, storytelling, and an uplifting tone—could be lifted directly into educational apps for other pediatric chronic conditions such as asthma or type 1 diabetes.","Observed engagement may be a proxy for interest rather than understanding; a paired pre- and post-knowledge measure would clarify whether the stars-and-levels system actually produces health literacy gains.","The hybrid Octo plush toy and app could serve as a shared communication object during clinic visits, letting children show clinicians what they understand about their condition."],"forward_implications":["Children as young as four can produce usable design feedback through observation-based testing when think-aloud is replaced by watching choices, expressions, and repeated actions.","The design criteria now specify a five-stop heart-shaped island, one stop per CHD condition, with three metaphor-based levels per stop.","The four tested conditions—ventricular septal defect, pulmonary stenosis, coarctation of the aorta, and atrial septal defect—form the curriculum of the finalized prototype.","The app's flexible navigation, star rewards, and progressive difficulty are intended to balance structured learning with free exploration.","The next phase will deploy the redesigned app in home and clinical settings with children and caregivers to test its real-world effectiveness."],"supporting_citations":[{"why":"It supplies the observation-first usability testing protocol for children with short attention spans, including reliance on non-verbal cues.","marker":"[4]"},{"why":"It anchors the multi-phase project in prior stakeholder assessment of children, parents, and providers that defined the educational needs.","marker":"[8]"},{"why":"It establishes through a scoping review that child-focused CHD health education is rare, the gap this app addresses.","marker":"[9]"},{"why":"It provides cooperative inquiry as the participatory design approach that positions children as active design informants.","marker":"[22]"},{"why":"It supports the three-level-per-activity structure as an effective serious game format for primary school children.","marker":"[26]"},{"why":"It supplies the four game-based learning principles used to evaluate activities and derive the final design criteria.","marker":"[37]"},{"why":"It defines serious games as having an explicit educational purpose, the framing that justifies the app's design.","marker":"[38]"},{"why":"It provides participatory research methods for young children that justify including ages 4-10 in the design process.","marker":"[53]"}],"fun_headline_variants":["Child testers reshape heart-disease app with train and river puzzles","30 kids' play tests guide redesign of Octo's Heartland app","Heart app turns anatomy into house and river puzzles for children","Children's input drives design of heart health learning app","Octo's Heartland app redesigned after 30 children test it"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The single load-bearing premise is that what children do and say during a 15-minute play session—facial expressions, repeated actions, and short answers—is valid evidence of what they understand about the medical content, even though no knowledge test was given.","fun_headline_variants_meta":{"raw":{"variants":["Child testers reshape heart-disease app with train and river puzzles","30 kids' play tests guide redesign of Octo's Heartland app","Heart app turns anatomy into house and river puzzles for children","Children's input drives design of heart health learning app","Octo's Heartland app redesigned after 30 children test it"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000249,"raw_usage":{"total_tokens":1523,"prompt_tokens":893,"completion_tokens":630,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":509,"completion_tokens_details":{"reasoning_tokens":542}},"tokens_in":509,"tokens_out":630,"duration_ms":6422,"temperature":1.0,"reasoning_tokens":542,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:36:04.274722+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give two comparable groups of children with CHD the redesigned Octo's Heartland and a non-game picture booklet covering the same four conditions, then measure heart knowledge before and after; if the game group shows no larger gain, the claim that metaphor-based gameplay improves comprehension is not supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the observation-first usability testing protocol for children with short attention spans, including reliance on non-verbal cues."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supports the three-level-per-activity structure as an effective serious game format for primary school children."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the four game-based learning principles used to evaluate activities and derive the final design criteria."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides participatory research methods for young children that justify including ages 4-10 in the design process."}],"review_version":2}