{"id":"9dcab3e8-2a7f-4e09-b767-7d47d3dbafe0","arxiv_id":"2608.00876","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Externalizing a system's interpretation of ambiguous spoken queries in XR sports viewing improves inspectability and repair language, but correction occurs in only 38% of misaligned trials, showing transparency and correction affordance are separate design axes.","lead":"This paper tests whether a soccer-viewing headset that shows its guesses about your spoken questions helps you notice and correct misunderstandings. It finds users do inspect the system's assumptions more, yet still fix only 38% of deliberately wrong answers, exposing a gap between seeing and acting.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 38% repair rate and 62% visibility-action gap rest on induced errors engineered to be 'plausible but wrong' and 'subtle, not obvious' (App. A.5); no evidence shows these match the pipeline's natural interpretation errors, so the orthogonality claim lacks validated external grounding.","rationale":"I concur with the reader's CONDITIONAL verdict and reach the same weakest point independently: the induced-misalignment manipulation is the load-bearing precondition for the distinctive visibility-action gap claim. The contribution is not just that externalization improves inspectability (Sec 7.1, robust) or explicit repair language (Sec 7.2, p=.039), but the orthogonal-axes lesson built on the 38% repair figure and the 62% gap (Sec 7.4, Sec 8.1). That figure characterizes the induced error distribution, not the system alone. The App. A.5 prompt constrains errors to be 'plausible but wrong,' 'subtle, not obvious,' and confidently delivered; natural errors from the same pipeline additionally include wrong-template errors (Sec 5.4), implausible outputs, and scope assumptions the externalization itself would surface. The differing dimensions -- obviousness, corrigibility, template fidelity -- directly control repair behavior, so neither the 38% rate nor its per-type pattern is self-evidently transferable. Prior manipulation precedents [20,30] and the Sec 8.2 caveat do not supply the missing validation. I weighed two rival concerns: (1) the keyword classifier undercounts implicit repairs, which would inflate the gap; this is real and acknowledged, but the paper frames rates as lower bounds and the comparative claims survive, making it a measurement caveat rather than the leading risk; (2) the inference that correction cost, not error detection, is the binding constraint is untested (the 'what went wrong' item was nonsignificant, p=.151), but that attacks the stronger form of the claim, while induction validity threatens both strong and weak readings. The proposed test -- re-annotating the free-exploration logs for natural errors and subsequent repair attempts -- can settle whether the induced and natural error regimes behave alike. Because this is an unvalidated precondition rather than a demonstrated falsification, and because the inspectability and satisfaction results are robust, the paper's claims stand conditionally on that check. No verdict change.","tokens_in":27686,"tokens_out":30130,"duration_ms":252329,"concrete_test":"Re-annotate the post-condition free-exploration logs (Sec 6.1: ~5 min unguided use, no misalignment prompt). Two annotators label each system response that misinterprets the user's intended query (wrong entity, time, stat, zone, or wrong template); then compute the share of naturally misaligned responses followed by any correction attempt, coded to include implicit repairs (scope narrowing, reformulation, pronoun re-reference) plus the 13 explicit patterns. If the natural-error repair rate differs materially from the 38% induced-misalignment rate, the manipulation is unrepresentative and the orthogonality claim needs re-scaling. Complement: run the misalignment-free pipeline on the same query set across multiple rollouts and compare natural versus induced error distributions on obviousness and corrigibility.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical pillar (Sec 7.4, Sec 8.1) is that externalization raised inspectability and repair language yet repair occurred in only 38% of misaligned externalization trials, producing the 62% visibility-action gap and the lesson that transparency and correction affordance are orthogonal design axes. The load-bearing precondition is that the induced misalignments (Sec 6.1; App. A.5) resemble the system's natural ambiguity-resolution errors. The induction prompt instructs GPT-4o to 'intentionally introduce a small factual error -- pick a plausible but wrong player, wrong time, or wrong stat' and to 'sound confident' with an error that is 'subtle, not obvious.' This generates one error subclass: committed, plausible-looking, confidently delivered wrong readings. Natural errors in this pipeline include classes the induction never generates: wrong-template errors from ambiguity misclassification (Sec 5.4: a misclassified query can render badges instead of a timeline, changing both the externalization and the repair path), answers outside the plausible-reading space, and scope-default errors the externalization itself would make salient. These classes differ on the very dimensions that determine repair -- obviousness, corrigibility, template fidelity -- so the 38% figure and its per-type variation need not transfer to real use. Citing precedents that manipulate outputs [20,30] and conceding only that induction 'potentially narrows the range of interpretation errors' (Sec 8.2) does not validate the error distribution; no natural-error repair baseline is reported. Since the gap claim is the paper's distinctive contribution, an unrepresentative manipulation over- or understates it, and the direction is not signed a priori.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how externalizing a system's interpretation of ambiguous spoken queries affects users' ability to inspect and correct misunderstandings in XR sports viewing. The authors first run a formative study that yields four ambiguity types (referential, spatial, temporal, metric), then define a design space along three dimensions (ambiguity type, interpretation state, externalization strategy), and instantiate it in an XR soccer system using situated visual cues. A within-subjects user study (N=16) compares this externalized condition with a voice-only baseline. The reported results are that externalization improves three of four inspectability ratings and overall satisfaction, increases the rate of explicit repair language, and imposes no measurable interaction cost. However, repair occurred in only 38% of misaligned externalization scenarios, which the authors interpret as a visibility-action gap and use to argue that transparency and correction affordance are orthogonal design axes.","tokens_in":27891,"tokens_out":6842,"duration_ms":63026,"significance":"If the central visibility-action gap claim is valid, the paper makes a useful empirical contribution to XR, HCI, and XAI: it provides evidence that making a system's interpretation visible helps users understand what was assumed, but that visible assumptions do not reliably translate into corrective action, suggesting that correction pathways must be designed separately from transparency. The paper has several strengths: the formative study reports high inter-rater reliability (Cohen's kappa=0.952), the evaluation reports effect sizes throughout, a full system prompt is included, and the authors are unusually explicit about confounds and limitations. The design space and four implemented externalization cases are also a clear contribution to situated visualization design. The main risk is that the headline 38% repair figure and the orthogonality conclusion rest on an induced-misalignment manipulation whose representativeness is not established, and on a scenario-level statistic that is never defined. These issues are fixable within the manuscript's scope.","major_comments":[{"comment":"The paper's central claim of a 62% visibility-action gap is measured only under induced misalignments. The induction prompt ('intentionally introduce a small factual error -- pick a plausible but wrong player, wrong time, or wrong stat... subtle, not obvious') produces a narrow class of errors: committed, plausible, confidently delivered wrong readings. Natural errors in this pipeline include other classes, most notably wrong-template errors from ambiguity misclassification (Sec 5.4 explicitly notes that a misclassified query can render badges instead of a timeline), as well as scope-default errors that the externalization itself would make salient. These classes differ on the dimensions that determine repair, namely obviousness, corrigibility, and template fidelity. The manuscript itself concedes in Sec 8.2 that induction 'potentially narrows the range of interpretation errors.' Because the 38% repair rate is load-bearing for the orthogonality argument, the authors should either provide evidence that induced errors resemble natural errors on detection and repair (e.g., a small comparison study or an analysis of naturally occurring misalignments from the free-exploration phase), or explicitly restrict the generalization to 'with subtle, induced plausible errors.'","section":"Sec 6.1, Appendix A.5, Sec 7.4, Sec 8.2"},{"comment":"The scenario-level statistic 'repair occurred in only 38% of misaligned externalization scenarios' is never defined in the measures section. The per-query repair rates reported in Table 9 are 3.0% to 9.9% by ambiguity type, and the overall repair utterance rate in Sec 7.2 is 5.6%, from which the 38% figure cannot be directly audited. Since the 62% gap is the central quantitative result, the manuscript must define the denominator and the aggregation rule (e.g., whether a scenario counts as repaired if at least one follow-up query contains repair language), and should report the corresponding test statistic or confidence interval for the scenario-level measure.","section":"Sec 7.4, Sec 6.3"},{"comment":"The headline significance claims do not account for multiple testing. Across the seven Likert items in Table 6 and the overall repair-rate comparison in Table 7, a conservative Bonferroni correction (alpha = 0.05/8) leaves only 'which moment shown' (p=.003) and overall satisfaction (p=.002) as significant; the overall repair-rate difference (p=.039), 'which player referred to' (p=.011), and 'alternatives considered' (p=.024) would not survive. Because the abstract's claim of 'increased explicit repair language overall' rests on this marginal uncorrected p-value, the authors should either apply a correction, provide a justified analysis plan for the unadjusted tests, or present the repair-rate effect as an exploratory trend that needs replication.","section":"Sec 7.1, Sec 7.2, Appendix B"},{"comment":"The experimental manipulation confounds interpretation externalization with visual richness: the externalization condition adds situated visual cues in addition to the voice response, while the baseline is voice-only. The authors acknowledge this in Sec 8.2, but several contributions statements (Abstract, Sec 7.3, Sec 8.1) phrase the result as if externalization of interpretation is the causal factor, e.g., 'externalization as a grounding mechanism.' To keep the central claim aligned with the experimental design, the manuscript should either consistently use the more limited wording 'visible interpretation cues vs. answer-only interaction,' or add a control condition that presents non-interpretation visual annotations to separate visual augmentation from interpretation externalization.","section":"Sec 6.1, Sec 8.1, Sec 8.2"}],"minor_comments":[{"comment":"The figure uses the panel label '(d)' twice, and the caption for the Interaction Cost panel cites 'Sec 7.3' although the results are reported in Sec 7.5. Please renumber the panels and correct the section reference.","section":"Figure 7"},{"comment":"The confusion matrix cells are hard to read because some entries lack separation, e.g., '490' appears to be two values (49 and 0) and '1200' appears to be two values (1 and 20 or 12 and 0). Please format the table with clear column separators or vertical rules.","section":"Table 2"},{"comment":"The text states that the repair classifier uses 13 word-boundary patterns, but only ten example patterns are listed after the description. Please list all 13 patterns or adjust the count.","section":"Sec 6.3"},{"comment":"The manipulation check reports M_mis=33.75 and M_align=27.62 without specifying whether these are per-participant per-condition totals or scenario-level averages. Please define the unit of analysis for these counts.","section":"Sec 7.2"},{"comment":"In the referential aligned-trials comparison, the notation 'M=0.0% vs. 14.1%' does not say which mean belongs to which condition; the sentence should explicitly state 'baseline 0.0% vs. externalization 14.1%.'","section":"Sec 7.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for TVCG and the authors are unusually candid about limitations. The main risk is that the paper's most generalizable claim (transparency and correction affordance are orthogonal axes) is currently supported by a narrowly induced-error manipulation and an undefined scenario-level repair statistic. I believe these can be fixed with additional analysis or explicit restriction of scope, so major revision rather than rejection is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Chunggi and colleagues have written a solid HCI paper. They identify four ambiguity types in spoken sports queries, contribute a design space for externalizing a system's interpretation (ambiguity type x interpretation state x strategy), instantiate it in an XR soccer viewer, and run a within-subjects study with N=16. The headline result is the visibility-action gap: externalization made system assumptions more inspectable and increased repair language, yet in misaligned trials repair occurred only 38% of the time, suggesting that transparency and correction affordance are separate design axes. That framing is genuinely useful and deserves a real referee.\n\nWhat's new: the post-commitment interpretation design space is a clean contribution, and the four implemented cases give concrete patterns for practitioners. The formative study is careful, with high inter-rater reliability, and the authors are appropriately cautious with their per-type analyses. The null interaction-cost results are reassuring. The supplement is thorough, including full system prompts.\n\nThe soft spots are real but not disqualifying. The main concern is the 38% figure itself. The misalignment induction prompt tells GPT-4o to introduce a 'subtle, not obvious' factual error, producing a specific error subclass: confidently delivered, plausible wrong readings. As the stress-test note argues, natural errors also include wrong-template cases from ambiguity misclassification (Sec 5.4), where the externalization itself changes, so the repair path differs. The authors acknowledge the limitation in Sec 8.2 but provide no natural-error baseline. That said, the qualitative claim — visibility does not guarantee correction, correction cost matters — does not rest on the exact percentage. I would treat the 38% as a manipulation-specific estimate, not a population value. Also, the voice-only baseline means visual richness is confounded with externalization; the paper concedes this. The per-type repair comparisons are underpowered and labeled exploratory, which is the right call. No code or data are released, and the closed LLM pipeline limits reproducibility, but those are mechanical observations.\n\nWho is this for: people building speech-driven XR or ambient assistants, and anyone working on XAI in situated contexts. The design space and the gap idea are worth citing, and the honesty of the limitations section makes this a paper I'd send to reviewers rather than desk reject.","headline":"Honest, well-built HCI paper with a fresh design space and a thought-provoking but unvalidated 'visibility-action gap' figure; worth refereeing despite the weak link between induced and natural errors.","tokens_in":28570,"tokens_out":2886,"would_cite":true,"duration_ms":26866,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Making a system's interpretation of ambiguous spoken queries visible helps viewers inspect what it assumed, but in only 38% of misaligned trials did they actually correct it — transparency and repair are separate design axes.","keywords":["extended reality","sports analytics","speech interaction","situated visualization","query ambiguity","interpretation externalization","repair behavior","visibility-action gap"],"falsifier":"Run the same four-scenario protocol without the misalignment prompt, log every interaction where the system's interpretation diverges from the user's stated intent (confirmed by post-task video review), and compute the repair rate on those naturally occurring errors; if it differs substantially from 38%, the visibility-action gap is an artifact of how errors were induced rather than a property of externalized interpretation.","tokens_in":27387,"feed_emoji":"🥅","tokens_out":9437,"duration_ms":72518,"temperature":0.7,"pith_summary":"Spoken queries during XR sports viewing are often underspecified — viewers omit which player, moment, field region, or statistic they mean — and when a system silently resolves such ambiguity, misunderstandings are hard to notice and repair. The paper argues that externalizing the system's committed interpretation through situated visual cues (candidate badges, zone overlays, scope panels, timeline thumbnails) should make those assumptions inspectable and correctable. A within-subjects study (N=16) found that externalization significantly improved inspectability on most measures and increased explicit repair language over a voice-only baseline. But repair occurred in only 38% of trials in which the system had been steered toward a plausible-but-wrong interpretation, leaving a 62% visibility-action gap that varied by ambiguity type. The paper's central claim is that transparency and correction affordance are orthogonal design axes: making assumptions visible is necessary but not sufficient, and the binding constraint is the cost of correction, not the quality of the externalization.","feed_headline":"62% of visible AI assumptions go uncorrected","feed_subtitle":"Even when an XR system shows how it understood a spoken query, viewers often skip correcting it during fast-paced sports viewing.","key_machinery":"The load-bearing machinery is the externalization design space plus its concrete cue designs: for referential ambiguity, situated candidate badges with role labels, brief reasons, and relative confidence ranking; for spatial ambiguity, labeled zone overlays on the field surface; for metric ambiguity, a multi-view scope panel showing the assumed metric, time range, and measure; for temporal ambiguity, a timeline of candidate replay segments with thumbnails, timestamps, and contextual descriptions. A language model classifies the query into one of the four ambiguity types and returns a structured response that selects which cue template renders. The argument is carried by the controlled pairing of this externalization against a voice-only baseline under two trial types — aligned (interpretation matches likely intent) and misaligned (interpretation steered, via an appended prompt instruction to the interpretation model, toward a plausible but non-primary reading) — which lets the study separate whether visible assumptions promote inspection from whether they promote corrective action.","core_discovery":"Through a formative study of 215 spoken utterances, the paper identifies four recurring ambiguity types in XR sports queries — referential (which entity), spatial (which region), temporal (which moment), and metric (which statistic) — and organizes externalization along three design dimensions: ambiguity type, interpretation state (alternatives, context, confidence, scope), and externalization strategy (operations, placement, targets, primitives). Instantiating this design space in an XR soccer viewing system, the authors compared externalized interpretation against voice-only answering in a within-subjects study with 16 participants. The central empirical discovery is a visibility-action gap: externalization raised inspectability (which player was referred to, which moment was shown, what alternatives were considered) and raised explicit repair language overall (5.6% vs 3.3% of queries), yet in misaligned trials participants corrected the system in only 38% of scenarios, and no ambiguity type showed a significant repair increase. In aligned trials, referential externalization triggered active verification (14.1% vs 0.0% repair language under baseline), which the authors read as evidence that visible commitments invite ratification — a grounding effect rather than mere error recovery. The authors conclude that interpretation visibility and correction affordance are orthogonal design axes, with correction cost, not transparency quality, as the binding constraint.","pith_inferences":["The reported repair rate is likely a lower bound: the keyword classifier counts only explicit correction language, so implicit repairs such as query narrowing or rephrasing would push the true rate above 38% — shrinking the visibility-action gap but not erasing it.","A direct test of the paper's correction-cost explanation would replace voice repair with direct selection on the same externalized cues (tap the candidate, scrub the timeline) in a replication: if repair rates rise sharply above 38%, cost is confirmed as the binding constraint; if they stay flat, attention or anchoring is doing the work.","The induced-misalignment manipulation may affect detectability: seeded errors are designed to be subtle and plausible, whereas naturally occurring language-model misinterpretations could be either more salient (an obviously wrong player) or harder to recognize (a plausible wrong metric), so the 38% figure should be re-measured on organically occurring errors before being treated as a general const","The per-type repair differences suggest a general heuristic for situated voice systems: correction cost scales with how precisely a user must verbally re-specify what the system got wrong, so any system that externalizes interpretations should also externalize selectable alternatives.","weakest_assumption_plain","The misaligned trials were manufactured by appending an instruction to the language model to introduce a small, plausible factual error, and the paper assumes these seeded misinterpretations resemble the errors the system would make naturally closely enough that the measured 38% repair rate transfers to real use.","falsifier","Run the same four-scenario protocol without the misalignment prompt, log every interaction where the system's interpretation diverges from the user's stated intent (confirmed by post-task video review), and compute the repair rate on those naturally occurring errors; if it differs substantially from 38%, the visibility-action gap is an artifact of how errors were induced rather than a property of "],"forward_implications":["Externalization can function as a continuous grounding channel rather than an error-recovery tool: in aligned referential trials it prompted active verification even when no error was present (14.1% vs 0.0% repair language under baseline), with an equal-or-higher trend across all four ambiguity types.","Correction mechanisms should reuse the visible cues themselves — tapping a candidate badge, scrubbing a timeline, toggling a metric scope, or selecting a zone — because the data indicate that the cost of verbal re-specification, not the clarity of the externalization, is what suppresses repair.","Systems should prioritize low-cost correction by consequence: wrong referents and replay windows demand it, while approximate spatial interpretations may be acceptable defaults, consistent with the observation that zone overlays were often accepted as close enough.","Externalization imposes no measurable interaction overhead (query count, confirmation time, and query length did not differ from baseline), which supports deploying it as a persistent default rather than an opt-in feature.","The 62% visibility-action gap implies that speech-driven XR systems should budget design effort for correction affordances as a first-class layer, separate from interpretation transparency."],"supporting_citations":[{"why":"Supplies the grounding framework used to interpret aligned-trial verification behavior: visible commitments invite ratification.","marker":"[12]"},{"why":"Precedent showing that making system decisions visible supports user correction in natural-language data visualization; the design space and interpretation-state framing build on it.","marker":"[17]"},{"why":"Provides evidence that language models commit to single interpretations without surfacing alternatives, motivating the externalization approach.","marker":"[40]"},{"why":"Shows that visibility alone does not guarantee correction, and is used to frame the central visibility-action gap result.","marker":"[5]"},{"why":"Anchoring-bias account used to explain why candidate badges in misaligned referential trials did not raise repair rates.","marker":"[50]"},{"why":"Precedent for deliberately manipulating system outputs to study detection and repair behavior, forming the basis of the misaligned-trial methodology.","marker":"[20]"},{"why":"Supplies the integrated spatiotemporal and event dataset for elite soccer from which the reconstructed match scenes are built.","marker":"[3]"},{"why":"Domain precedent for question answering with embedded visualizations in sports video, informing the study scenarios and system context.","marker":"[33]"}],"fun_headline_variants":["Only 38% of visible AI misreads get fixed","XR users see AI errors but rarely correct them","Transparency without correction: the XR gap","Visible assumptions, but only 38% corrected","Grounding beats repair in XR query design"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The misaligned trials were manufactured by appending an instruction to the language model to introduce a small, plausible factual error, and the paper assumes these seeded misinterpretations resemble the errors the system would make naturally closely enough that the measured 38% repair rate transfers to real use.","fun_headline_variants_meta":{"raw":{"variants":["Only 38% of visible AI misreads get fixed","XR users see AI errors but rarely correct them","Transparency without correction: the XR gap","Visible assumptions, but only 38% corrected","Grounding beats repair in XR query design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000139,"raw_usage":{"total_tokens":1221,"prompt_tokens":1071,"completion_tokens":150,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":687,"completion_tokens_details":{"reasoning_tokens":77}},"tokens_in":687,"tokens_out":150,"duration_ms":1953,"temperature":1.0,"reasoning_tokens":77,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:15:21.996919+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four-scenario protocol without the misalignment prompt, log every interaction where the system's interpretation diverges from the user's stated intent (confirmed by post-task video review), and compute the repair rate on those naturally occurring errors; if it differs substantially from 38%, the visibility-action gap is an artifact of how errors were induced rather than a property of externalized interpretation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence that language models commit to single interpretations without surfacing alternatives, motivating the externalization approach."},{"cited_title":"Bassek, R","cited_arxiv_id":null,"evidence_quote":"Supplies the integrated spatiotemporal and event dataset for elite soccer from which the reconstructed match scenes are built."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Domain precedent for question answering with embedded visualizations in sports video, informing the study scenarios and system context."}],"review_version":2}