{"id":"6c0388a2-e44f-431a-a86e-2eae996b62bd","arxiv_id":"2501.13258","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Embedding reflective prompts in AR instructions improved objective task understanding and increased voluntary information-seeking in a 16-person within-subject study.","lead":"This paper tested whether adding short reflective questions to step-by-step augmented reality instructions helps people understand a task rather than just follow along. In two hands-on tasks, participants who saw such prompts scored higher on understanding quizzes and looked up more optional information while working.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The objective-understanding quiz may be contaminated by reflective prompt content; since quiz items are not reported, the central claim that reflective prompts improve understanding is not independently verifiable.","rationale":"The central claim of the paper is that reflective prompts improve objective task understanding and increase voluntary information seeking. The objective-understanding result rests entirely on a bespoke quiz that is not reported. The reflective prompts were designed to trigger thoughts about necessity, purpose, and alternatives, which are exactly the kinds of constructs a quiz on 'rationale behind the actions' would likely test. If the quiz items paraphrase the prompts, the observed effect could reflect short-term priming of prompt content rather than a genuine, transferable improvement in understanding. The authors themselves acknowledge in Section 6.5.1 that quizzes have limited ability to capture deeper cognitive engagement and real-world applicability. This is not a case of internal inconsistency, but it is a correctness risk in the interpretation of the main dependent variable. The reader identified this same concern as the weakest assumption. The stated p-value discrepancy for the keyword-interaction result is a secondary issue; it does not change the significance of that finding, but it does indicate a need for careful reporting. Other potential concerns, such as the lack of a control for extra text in the reflective condition, would further test the specificity of the intervention, but the quiz-overlap issue is the most damaging because it directly undermines the primary quantitative claim. A conditional verdict is appropriate: the paper's conclusion is plausible, but it cannot be fully trusted until the quiz materials are made available and the analysis is rerun on items that are independent of prompt content. The recommended concrete test would settle whether the effect survives this threat.","tokens_in":24528,"tokens_out":8159,"duration_ms":86342,"concrete_test":"Contact the authors for the complete quiz item sets (Section 4.3) and the exact reflective prompts used in the evaluation (Section 4.2, Fig. 5). Have two independent coders, blind to the study hypotheses, classify each quiz item as 'overlapping' if it tests the same construct or paraphrases a prompt (e.g., asking why a step is necessary when the prompt asked 'Is this step necessary?'). Then recompute the paired t-test on quiz scores using only non-overlapping items. If the effect is no longer statistically significant or drops below a minimal effect size (e.g., d<0.3), the reflective-prompt advantage is better explained by priming than by improved understanding. As a secondary check, reconcile the reported p-value for keyword clicks between Figure 1 (p<.01) and Section 4.5 (p<.05).","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 describes quizzes that 'assess how well participants remember the task procedure and understand the rationale behind the actions,' but the actual items are not disclosed. The intervention itself consists of three prompt types (Section 3.3.2, Fig. 5) that ask exactly about necessity ('Is this step necessary?'), purpose ('What are we trying to achieve with this step?'), and alternatives ('What if we pour all the water at once?'). The quiz is therefore highly susceptible to content overlap: participants in the reflective condition have been cued, moments before, with the very concepts the quiz later rewards. Because the quiz is the sole quantitative basis for the headline claim of improved 'objective task understanding' (Section 4.5, t(15)=2.33, d=0.582), this is a load-bearing threat to construct validity. The paper's own Limitation 6.5.1 concedes that quizzes mainly capture recall and procedural knowledge rather than deeper transferable understanding, which makes interpreting the quiz gain as 'understanding' even more fragile. The p-value inconsistency for the keyword-interaction result (Fig. 1 says p<.01; Section 4.5 says p<.05) is an additional reporting issue, but the quiz-overlap problem is what most directly undercuts the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether embedding reflective prompts into AR task instructions improves users' understanding of the task being performed. The authors first conduct a formative survey and co-design sessions (N=9) to select three prompt types: Challenging Assumptions, Connection to Outcomes, and Hypothetical Scenarios. They then run a within-subject evaluation (N=16) comparing AR instructions with and without these prompts on coffee-making and circuit-assembly tasks. The main quantitative results are a significant increase in quiz scores (interpreted as objective understanding) and in voluntary clicks on optional explanation keywords, with no measured increase in cognitive load or loss of usability. Qualitative interviews and pre/post questionnaires are used to support design guidelines for reflective AR instructional systems.","tokens_in":24712,"tokens_out":3558,"duration_ms":36758,"significance":"If the central claim survives scrutiny, the contribution is valuable for AR instruction design: it provides a concrete, co-designed prompt taxonomy, a within-subject evaluation with effect sizes, and design guidelines grounded in both quantitative and qualitative data. The paper's strengths include counterbalanced task/condition assignment, standardized clickable keyword explanations, normalization of quiz scores within tasks, and an explicit treatment of limitations. The main risk is that the objective-understanding measure is the sole quantitative basis for the headline claim and its items are not disclosed, so the claim is not currently independently verifiable.","major_comments":[{"comment":"The objective-understanding measure is not verifiable and may be aligned with the intervention. The quizzes are described as assessing 'basic conceptual understanding, factual memory recall, and knowledge transfer' and as testing whether participants 'remember the task procedure and understand the rationale behind the actions,' but no quiz items are reported. The reflective prompts themselves ask questions such as 'Is this step necessary?', 'What are we trying to achieve with this step?', and 'What if we pour all the water at once?' (Fig. 5), which map directly onto the quiz's stated constructs. Participants in the reflective condition were therefore cued with the very concepts the quiz later rewards, so the significant quiz gain (t(15)=2.33, p<.05, d=0.582) could reflect priming or wording overlap rather than improved task understanding. This is load-bearing for the central claim. Please include the full quiz instruments, a mapping of each item to the prompt types, and a reanalysis that excludes or separately reports items overlapping with prompt content. At minimum, the interpretation of the quiz as 'objective understanding' should be softened, especially because Limitation 6.5.1 concedes that quizzes mainly capture recall and procedural knowledge rather than deeper transferable understanding.","section":"§4.3, §4.5, Fig. 5"},{"comment":"The chi-square test of independence is applied to paired within-subject data, violating the independence assumption; each participant contributed one success/failure observation per condition, so the two conditions are not independent samples. A paired test such as McNemar's test should be used. In addition, the reported success rates (94.8% vs. 68.8%) are inconsistent with 16 tasks per condition, which would be 93.75% (15/16) and 68.75% (11/16). Please correct the rates and re-run the appropriate test.","section":"§4.5, Task Success"},{"comment":"The Task B comparison is described as a paired sample t-test (t(7)=3.035, p<.05), but no participant performed Task B in both conditions: half of the participants did Task B with reflective prompts and half without. This is an independent-samples comparison of n=8 per group, not a paired comparison. The same issue applies to the Task A subjective-understanding comparison. Please revise the statistical description and re-report these results with the correct test.","section":"§4.5, Subjective Understanding"}],"minor_comments":[{"comment":"The summary of findings in Fig. 1 reports that reflective prompts increase information-seeking behaviors 'from 0.32 to 0.54, p<.01', while §4.5 reports M=0.538, SD=0.324 for the reflective condition and p<.05. Please align the reported mean and p-value across the figure and the text.","section":"Fig. 1 vs. §4.5"},{"comment":"The procedure for excluding quiz items that participants marked as known before the study is described only in aggregate ('an average of 1.1 questions per participant per quiz'). Please report the number of excluded items per condition and confirm that exclusion rates did not differ systematically between conditions.","section":"§4.3"},{"comment":"The caption contains a typo ('connect the cathode ... to to the negative rail'); please correct it.","section":"Fig. 6 caption"},{"comment":"The click-rate metric is defined as the number of clicks in steps with reflective prompts divided by the number of clickable keywords in those steps, but the non-reflective condition has no reflective prompts. Please clarify how the denominator was defined for the non-reflective condition so the comparison is interpretable.","section":"§4.5, Keyword Interaction"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a CHI-style venue and the formative/design work is solid. The decisive issue is the undisclosed quiz instrument: the main quantitative claim is only as strong as the evidence that the quiz measures transferable understanding rather than cued recall of prompt-like content. This is fixable with an appendix of quiz items and a reanalysis, so I am not recommending rejection, but the revision must address it head-on. The statistical issues in §4.5 (paired vs. independent tests, chi-square on paired data) also need correction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine empirical paper, not a rebranding exercise, but the headline claim — reflective prompts improve objective task understanding — rests on a quiz that is neither disclosed nor clearly independent of the prompts themselves. The behavioral result is the cleaner part of the evidence.\n\nWhat's new: transferring established reflective-prompt strategies from education into AR instruction-following, with a co-design step (N=9) that narrowed a literature-derived taxonomy to three prompt types, and a counterbalanced within-subject evaluation (N=16) across two deliberately contrastive tasks. That's a reasonable contribution to the AR-guidance subfield. The study is competently run and honestly reported: standard paired tests with effect sizes, exclusion of quiz items participants knew beforehand, no significant cognitive-load or usability differences, and null and negative results reported straight — including the counterintuitive finding that subjective understanding dropped for the circuit task. The qualitative material on how prompts interrupt automatic step-following is genuinely useful, and Section 6.5.1 concedes the quizzes capture recall and procedural knowledge more than deep transfer. I'd call the thinking clear and the writing unusually candid.\n\nSoft spots, in order of weight. First, the objective-understanding measure. The reflective prompts ask \"Is this step necessary?\", \"What are we trying to achieve?\", \"What if we pour all the water at once?\" — and the quiz tests rationale and transfer, exactly the concepts the prompts rehearse. Quiz items are not in the paper, so the gain (0.66 SD) could be partly cueing or recall of prompt-like phrasing. The stress-test note is right that this is the main threat to the central claim, though I'd call it a fixable construct-validity issue, not a disqualifying one; publishing the quiz, the prompt schedule, and per-step placement as supplementary material would largely settle it. Second, there is an internal contradiction: Section 4.5 reports Task B subjective understanding was significantly lower with prompts (t(7)=3.035, p<.05), while Section 6.2 says a paired t-test confirmed prompts did not significantly lower subjective understanding. Both can't stand without a scope qualifier. Third, the keyword-interaction p-value is p<.01 in Figure 1 and p<.05 in Section 4.5. Minor. Fourth, the click-rate metric is described as clicks \"in steps with reflective prompts,\" which doesn't parse for the non-reflective condition — presumably they matched step subsets across conditions, but they need to say so. No artifacts or data are posted.\n\nWho this is for: designers of AR instructional systems and anyone studying how to measure learning when the intervention and the test share surface content. It deserves a serious referee; the design is sound enough that the outcome should turn on the supplemental materials, not on desk rejection. My recommendation: send it to review, and ask for the quiz and prompt materials, a clean statement of the click-rate definition, and reconciliation of the contradictory subjective-understanding sentences. With those, the defensible claim is already good: prompts change information-seeking behavior and raise quiz scores, with how much of that counts as \"understanding\" still open.","headline":"A competently run, honestly reported AR study whose headline 'understanding' gain rests on an undisclosed quiz that likely overlaps with the prompt content; the behavioral info-seeking result is the cleaner finding.","tokens_in":25242,"tokens_out":8663,"would_cite":false,"duration_ms":81009,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Adding short reflective questions to AR instructions improves measured task understanding and increases voluntary information seeking.","keywords":["augmented reality","reflective prompts","task understanding","instruction following","epistemic curiosity","information seeking","task guidance","user study"],"falsifier":"Compare the same two conditions using quiz items whose wording provably does not echo the prompt questions, or move the quiz to a new context days later; if the gain disappears or concentrates only on items that reuse prompt vocabulary, the claimed understanding effect is largely priming. A reader could also inspect the unreported quiz items for direct overlap with 'Is this step necessary?', 'What are we trying to achieve?', and 'What if...?' wording.","tokens_in":24297,"feed_emoji":"🥽","tokens_out":7206,"duration_ms":69885,"temperature":0.7,"pith_summary":"This paper asks whether prompting people to reflect while they follow augmented-reality instructions can move them from passive step-following to genuine task understanding. In a within-subject study, 16 novices performed coffee-making and circuit-assembly tasks with standard AR instructions in one condition and the same instructions plus short reflective questions in the other. Tasks with reflective prompts produced significantly higher quiz scores measuring objective understanding, a 0.66 standard deviation increase, and a 68.75% rise in clicks on optional explanation keywords, with no significant change in cognitive load or usability. The authors take this as evidence that small, well-timed questions can deepen comprehension inside an otherwise instruction-following medium.","feed_headline":"Asking 'why?' in AR instructions boosts task understanding","feed_subtitle":"Short reflective questions in AR task guidance raised quiz scores and lifted info-seeking clicks by 69 percent.","key_machinery":"The operative mechanism is the reflective prompt: a short, non-interactive question overlaid in AR text below an instruction step, appearing three seconds after the step. Three prompt types survived the formative co-design: challenging assumptions ('Is this step necessary?'), connecting actions to outcomes ('What are we trying to achieve with this step?'), and hypothetical scenarios ('What if we pour all the water at once?'). The prompts were attached to seven of thirteen coffee steps and five of eight circuit steps, and the system measured information seeking through clicks on predefined clickable keywords that revealed extra explanation. The mechanism works by interrupting automatic execution just long enough for users to notice or articulate questions they would otherwise skip.","core_discovery":"On the paper's own terms, the central claim is that embedding reflective prompts into AR task instructions improves how well users understand the task, not just how well they execute it. The evaluative evidence is a paired comparison in which each of 16 participants did one task with and one without prompts; quiz scores, normalized per task and excluding questions participants already knew, rose by 0.66 standard deviations in the reflective condition (t(15)=2.33, p<.05, d=0.582). Participants also clicked on optional explanation keywords 68.75% more often in the reflective condition (reflective M=0.538 vs non-reflective M=0.32, t(15)=2.943, p<.05, d=0.736). The authors report no significant differences in cognitive load or system usability, and most participants (15/16) described the prompts as non-intrusive; perceived understanding was significantly lower with prompts only for the circuit-assembly task.","pith_inferences":["Extension not settled by the paper: because the quiz items are not reported and the prompts themselves model words like 'necessary' and 'achieve', some of the quiz gain may be wording priming; a replication with novel quiz items and delayed transfer questions would test this.","Extension: the success-rate gap (94.8% vs 68.8%) was not statistically significant in this sample, so a larger replication is needed before treating improved task success as an established outcome of reflective prompts.","Extension: participants said errors and deviations naturally triggered reflection, so an adaptive version that fires prompts at error-prone moments is a testable next step that the current fixed-step system does not evaluate.","Extension: the authors' guidelines imply a system with user-selectable modes for efficient versus reflective guidance; a direct next experiment would ask whether voluntary mode choice preserves the quiz gain or requires forced prompts."],"forward_implications":["AR instruction systems can adopt short, ignorable reflective questions as a lightweight way to raise measured understanding without raising cognitive load or lowering usability.","Click-through on optional explanation keywords can serve as an observable proxy for epistemic curiosity in AR guidance studies.","Prompt placement matters: the design guidelines advise tying prompts to the current step, avoiding high-load moments, and keeping the tone conversational and brief.","If the effect generalizes, reflective prompts belong in tasks where the paper's reflection values--safety, quality, efficiency, customizability, and skill--make understanding valuable."],"supporting_citations":[{"why":"Grounds the core premise that reflection turns experience into learning, which the prompt design depends on.","marker":"[11]"},{"why":"Provides the reflective-prompt scaffolding technique that the AR prompts adapt from educational contexts.","marker":"[24]"},{"why":"Supplies experiential learning theory, the framework for why reflection after action deepens understanding.","marker":"[39]"},{"why":"Documents the problem the paper targets: AR guidance can leave users passively following steps without deep reasoning.","marker":"[59]"},{"why":"Motivates reflection-in-action during task execution, the real-time timing used by the prompts.","marker":"[62]"},{"why":"Closest prior empirical anchor: reflective mechanisms in AR-based learning improved inquiry outcomes.","marker":"[48]"},{"why":"Supports the use of curiosity and information-seeking as learning-relevant behavior measured by keyword clicks.","marker":"[35]"},{"why":"Supports the behavioral measure of curiosity through information-seeking behavior.","marker":"[72]"}],"fun_headline_variants":["Reflective AR prompts boost task understanding by 0.66 SD","Asking 'why' in AR instructions boosts understanding and curiosity","AR prompts that ask 'why' improve learning, not just task completion","Reflective AR questions boost understanding; info-seeking clicks jump 69%","Why-asking AR instructions deepen understanding and spark info-seeking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the post-task quizzes measure transferable task understanding rather than recall of the prompt-like phrasing, and because the quiz items are not reported, the reader cannot check how much the prompts and quiz overlap.","fun_headline_variants_meta":{"raw":{"variants":["Reflective AR prompts boost task understanding by 0.66 SD","Asking 'why' in AR instructions boosts understanding and curiosity","AR prompts that ask 'why' improve learning, not just task completion","Reflective AR questions boost understanding; info-seeking clicks jump 69%","Why-asking AR instructions deepen understanding and spark info-seeking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3244,"prompt_tokens":907,"completion_tokens":2337,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":2247}},"tokens_in":523,"tokens_out":2337,"duration_ms":18300,"temperature":1.0,"reasoning_tokens":2247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:19:00.732383+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the same two conditions using quiz items whose wording provably does not echo the prompt questions, or move the quiz to a new context days later; if the gain disappears or concentrates only on items that reuse prompt vocabulary, the claimed understanding effect is largely priming. A reader could also inspect the unreported quiz items for direct overlap with 'Is this step necessary?', 'What are we trying to achieve?', and 'What if...?' wording.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the core premise that reflection turns experience into learning, which the prompt design depends on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the reflective-prompt scaffolding technique that the AR prompts adapt from educational contexts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior empirical anchor: reflective mechanisms in AR-based learning improved inquiry outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the use of curiosity and information-seeking as learning-relevant behavior measured by keyword clicks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the behavioral measure of curiosity through information-seeking behavior."}],"review_version":1}