REVIEW 4 major objections 5 minor 74 references
Who's That Player?: Externalizing Query Interpretation in Spoken XR Sports Interaction
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Making a system's interpretation of ambiguous spoken queries visible helps viewers inspect what it assumed, but in only 38% of misaligned trials did they actually correct it — transparency and repair are separate design axes.
desk verdict Honest, well-built HCI paper with a fresh design space and a thought-provoking but unvalidated 'visibility-action gap' figure; worth refereeing despite the weak link between induced and natural errors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the externalization design space plus its concrete cue designs: for referential ambiguity, situated candidate badges with role labels, brief reasons, and relative confidence ranking; for spatial ambiguity, labeled zone overlays on the field surface; for metric ambiguity, a multi-view scope panel showing the assumed metric, time range, and measure; for temporal ambiguity, a timeline of candidate replay segments with thumbnails, timestamps, and contextual descriptions. A language model classifies the query into one of the four ambiguity types and returns a structured response that selects which cue template renders. The argument is carried by the controlled pairing of this externalization against a voice-only baseline under two trial types — aligned (interpretation matches likely intent) and misaligned (interpretation steered, via an appended prompt instruction to the interpretation model, toward a plausible but non-primary reading) — which lets the study separate whether visible assumptions promote inspection from whether they promote corrective action.
What would settle it
Run the same four-scenario protocol without the misalignment prompt, log every interaction where the system's interpretation diverges from the user's stated intent (confirmed by post-task video review), and compute the repair rate on those naturally occurring errors; if it differs substantially from 38%, the visibility-action gap is an artifact of how errors were induced rather than a property of externalized interpretation.
Extended reading notes
Core claim
Through a formative study of 215 spoken utterances, the paper identifies four recurring ambiguity types in XR sports queries — referential (which entity), spatial (which region), temporal (which moment), and metric (which statistic) — and organizes externalization along three design dimensions: ambiguity type, interpretation state (alternatives, context, confidence, scope), and externalization strategy (operations, placement, targets, primitives). Instantiating this design space in an XR soccer viewing system, the authors compared externalized interpretation against voice-only answering in a within-subjects study with 16 participants. The central empirical discovery is a visibility-action gap: externalization raised inspectability (which player was referred to, which moment was shown, what alternatives were considered) and raised explicit repair language overall (5.6% vs 3.3% of queries), yet in misaligned trials participants corrected the system in only 38% of scenarios, and no ambiguity type showed a significant repair increase. In aligned trials, referential externalization triggered active verification (14.1% vs 0.0% repair language under baseline), which the authors read as evidence that visible commitments invite ratification — a grounding effect rather than mere error recovery. The authors conclude that interpretation visibility and correction affordance are orthogonal design axes, with correction cost, not transparency quality, as the binding constraint.
Load-bearing premise
The misaligned trials were manufactured by appending an instruction to the language model to introduce a small, plausible factual error, and the paper assumes these seeded misinterpretations resemble the errors the system would make naturally closely enough that the measured 38% repair rate transfers to real use.
Editorial extensions
If this is right
- Externalization can function as a continuous grounding channel rather than an error-recovery tool: in aligned referential trials it prompted active verification even when no error was present (14.1% vs 0.0% repair language under baseline), with an equal-or-higher trend across all four ambiguity types.
- Correction mechanisms should reuse the visible cues themselves — tapping a candidate badge, scrubbing a timeline, toggling a metric scope, or selecting a zone — because the data indicate that the cost of verbal re-specification, not the clarity of the externalization, is what suppresses repair.
- Systems should prioritize low-cost correction by consequence: wrong referents and replay windows demand it, while approximate spatial interpretations may be acceptable defaults, consistent with the observation that zone overlays were often accepted as close enough.
- Externalization imposes no measurable interaction overhead (query count, confirmation time, and query length did not differ from baseline), which supports deploying it as a persistent default rather than an opt-in feature.
- The 62% visibility-action gap implies that speech-driven XR systems should budget design effort for correction affordances as a first-class layer, separate from interpretation transparency.
Reading between the lines
- The reported repair rate is likely a lower bound: the keyword classifier counts only explicit correction language, so implicit repairs such as query narrowing or rephrasing would push the true rate above 38% — shrinking the visibility-action gap but not erasing it.
- A direct test of the paper's correction-cost explanation would replace voice repair with direct selection on the same externalized cues (tap the candidate, scrub the timeline) in a replication: if repair rates rise sharply above 38%, cost is confirmed as the binding constraint; if they stay flat, attention or anchoring is doing the work.
- The induced-misalignment manipulation may affect detectability: seeded errors are designed to be subtle and plausible, whereas naturally occurring language-model misinterpretations could be either more salient (an obviously wrong player) or harder to recognize (a plausible wrong metric), so the 38% figure should be re-measured on organically occurring errors before being treated as a general const
- The per-type repair differences suggest a general heuristic for situated voice systems: correction cost scales with how precisely a user must verbally re-specify what the system got wrong, so any system that externalizes interpretations should also externalize selectable alternatives.
- weakest_assumption_plain
- The misaligned trials were manufactured by appending an instruction to the language model to introduce a small, plausible factual error, and the paper assumes these seeded misinterpretations resemble the errors the system would make naturally closely enough that the measured 38% repair rate transfers to real use.
- falsifier
- Run the same four-scenario protocol without the misalignment prompt, log every interaction where the system's interpretation diverges from the user's stated intent (confirmed by post-task video review), and compute the repair rate on those naturally occurring errors; if it differs substantially from 38%, the visibility-action gap is an artifact of how errors were induced rather than a property of
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates how externalizing a system's interpretation of ambiguous spoken queries affects users' ability to inspect and correct misunderstandings in XR sports viewing. The authors first run a formative study that yields four ambiguity types (referential, spatial, temporal, metric), then define a design space along three dimensions (ambiguity type, interpretation state, externalization strategy), and instantiate it in an XR soccer system using situated visual cues. A within-subjects user study (N=16) compares this externalized condition with a voice-only baseline. The reported results are that externalization improves three of four inspectability ratings and overall satisfaction, increases the rate of explicit repair language, and imposes no measurable interaction cost. However, repair occurred in only 38% of misaligned externalization scenarios, which the authors interpret as a visibility-action gap and use to argue that transparency and correction affordance are orthogonal design axes.
Significance. If the central visibility-action gap claim is valid, the paper makes a useful empirical contribution to XR, HCI, and XAI: it provides evidence that making a system's interpretation visible helps users understand what was assumed, but that visible assumptions do not reliably translate into corrective action, suggesting that correction pathways must be designed separately from transparency. The paper has several strengths: the formative study reports high inter-rater reliability (Cohen's kappa=0.952), the evaluation reports effect sizes throughout, a full system prompt is included, and the authors are unusually explicit about confounds and limitations. The design space and four implemented externalization cases are also a clear contribution to situated visualization design. The main risk is that the headline 38% repair figure and the orthogonality conclusion rest on an induced-misalignment manipulation whose representativeness is not established, and on a scenario-level statistic that is never defined. These issues are fixable within the manuscript's scope.
major comments (4)
- [Sec 6.1, Appendix A.5, Sec 7.4, Sec 8.2] The paper's central claim of a 62% visibility-action gap is measured only under induced misalignments. The induction prompt ('intentionally introduce a small factual error -- pick a plausible but wrong player, wrong time, or wrong stat... subtle, not obvious') produces a narrow class of errors: committed, plausible, confidently delivered wrong readings. Natural errors in this pipeline include other classes, most notably wrong-template errors from ambiguity misclassification (Sec 5.4 explicitly notes that a misclassified query can render badges instead of a timeline), as well as scope-default errors that the externalization itself would make salient. These classes differ on the dimensions that determine repair, namely obviousness, corrigibility, and template fidelity. The manuscript itself concedes in Sec 8.2 that induction 'potentially narrows the range of interpretation errors.' Because the 38% repair rate is load-bearing for the orthogonality argument, the authors should either provide evidence that induced errors resemble natural errors on detection and repair (e.g., a small comparison study or an analysis of naturally occurring misalignments from the free-exploration phase), or explicitly restrict the generalization to 'with subtle, induced plausible errors.'
- [Sec 7.4, Sec 6.3] The scenario-level statistic 'repair occurred in only 38% of misaligned externalization scenarios' is never defined in the measures section. The per-query repair rates reported in Table 9 are 3.0% to 9.9% by ambiguity type, and the overall repair utterance rate in Sec 7.2 is 5.6%, from which the 38% figure cannot be directly audited. Since the 62% gap is the central quantitative result, the manuscript must define the denominator and the aggregation rule (e.g., whether a scenario counts as repaired if at least one follow-up query contains repair language), and should report the corresponding test statistic or confidence interval for the scenario-level measure.
- [Sec 7.1, Sec 7.2, Appendix B] The headline significance claims do not account for multiple testing. Across the seven Likert items in Table 6 and the overall repair-rate comparison in Table 7, a conservative Bonferroni correction (alpha = 0.05/8) leaves only 'which moment shown' (p=.003) and overall satisfaction (p=.002) as significant; the overall repair-rate difference (p=.039), 'which player referred to' (p=.011), and 'alternatives considered' (p=.024) would not survive. Because the abstract's claim of 'increased explicit repair language overall' rests on this marginal uncorrected p-value, the authors should either apply a correction, provide a justified analysis plan for the unadjusted tests, or present the repair-rate effect as an exploratory trend that needs replication.
- [Sec 6.1, Sec 8.1, Sec 8.2] The experimental manipulation confounds interpretation externalization with visual richness: the externalization condition adds situated visual cues in addition to the voice response, while the baseline is voice-only. The authors acknowledge this in Sec 8.2, but several contributions statements (Abstract, Sec 7.3, Sec 8.1) phrase the result as if externalization of interpretation is the causal factor, e.g., 'externalization as a grounding mechanism.' To keep the central claim aligned with the experimental design, the manuscript should either consistently use the more limited wording 'visible interpretation cues vs. answer-only interaction,' or add a control condition that presents non-interpretation visual annotations to separate visual augmentation from interpretation externalization.
minor comments (5)
- [Figure 7] The figure uses the panel label '(d)' twice, and the caption for the Interaction Cost panel cites 'Sec 7.3' although the results are reported in Sec 7.5. Please renumber the panels and correct the section reference.
- [Table 2] The confusion matrix cells are hard to read because some entries lack separation, e.g., '490' appears to be two values (49 and 0) and '1200' appears to be two values (1 and 20 or 12 and 0). Please format the table with clear column separators or vertical rules.
- [Sec 6.3] The text states that the repair classifier uses 13 word-boundary patterns, but only ten example patterns are listed after the description. Please list all 13 patterns or adjust the count.
- [Sec 7.2] The manipulation check reports M_mis=33.75 and M_align=27.62 without specifying whether these are per-participant per-condition totals or scenario-level averages. Please define the unit of analysis for these counts.
- [Sec 7.3] In the referential aligned-trials comparison, the notation 'M=0.0% vs. 14.1%' does not say which mean belongs to which condition; the sentence should explicitly state 'baseline 0.0% vs. externalization 14.1%.'
Circularity Check
No circular derivation: the paper's empirical claims are measured behavioral outcomes, and its acknowledged limitations are validity threats, not circularity.
full rationale
The paper contains no derivation chain of the kind that could be circular. Its central quantitative claims—higher inspectability, higher explicit repair language, and the 38% repair rate / 62% visibility-action gap in misaligned trials—are measured behavioral outcomes from a within-subjects study (Sec 7.1–7.4), not quantities derived from fitted parameters or from definitions. The misalignment induction in Sec 6.1 and App. A.5 is an experimental manipulation (a prompt override instructing GPT-4o to make a plausible but subtle error), and the paper itself flags the representativeness concern in Sec 8.2 as 'potentially narrowing the range of interpretation errors'; that is a validity threat, not circularity. Self-citations to prior systems (Sportify [33], Omnioculars [39], iBall [63]) are used as methodological precedents for reconstructed match data and study design, not as load-bearing justification for the empirical claims. The taxonomy and design space are author-generated constructs, but they are instantiated and evaluated through independent behavioral data, and the acknowledged GPT-4o classification/interpretation coupling (Sec 8.2) is a system limitation rather than a case of a prediction being equivalent to its input. No quoted equation or definition makes any claimed result true by construction; the empirical measures are separable from the constructs being investigated. Therefore no significant circularity is present.
Assumptions & free parameters
free parameters (1)
- Repair classifier keyword patterns
assumptions (7)
- domain assumption The four ambiguity types (referential, spatial, temporal, metric) identified in the N=5 formative study cover the space of underspecification in spoken XR sports queries.
- domain assumption GPT-4o classification (86.5% accuracy, kappa=0.815) is reliable enough that the wrong externalization template is rarely triggered in the user study.
- domain assumption The keyword-based repair classifier (13 patterns) measures repair behavior comparably across conditions.
- domain assumption Prompt-induced misalignment yields interpretations that are plausible but non-primary and representative of genuine system errors.
- domain assumption Voice-only is an adequate baseline for testing externalization despite differing in visual richness.
- domain assumption Self-reported Likert items adapted from XAI transparency frameworks are valid measures of inspectability.
- standard math Wilcoxon signed-rank tests on per-participant paired means are appropriate for N=16 within-subjects data.
invented entities (1)
-
Design space for externalizing interpretations (ambiguity type x interpretation state x externalization strategy)
Cite this review
Pith. "Pith review of Who's That Player?: Externalizing Query Interpretation in Spoken XR Sports Interaction." pith.science (2026). https://pith.science/paper/RFZRTUHW
@misc{pith2026260800876,
author = {Pith},
title = {Pith review of: Who's That Player?: Externalizing Query Interpretation in Spoken XR Sports Interaction},
year = {2026},
howpublished = {\url{https://pith.science/paper/RFZRTUHW}},
note = {Machine review of arXiv:2608.00876}
}
read the original abstract
XR sports viewing enables spectators to follow play from immersive, spatially anchored perspectives while accessing contextual analytics directly within the scene. In such settings, speech offers a practical interaction modality because text entry and menu navigation can interrupt attention during fast-paced gameplay. However, spoken queries are often underspecified: viewers may omit which player, time period, field location, or metric they intend. When systems resolve these ambiguities implicitly, their assumptions remain hidden, making misinterpretations difficult to notice and correct (repair). We investigate how externalizing a system's interpretation of spoken queries can support inspection and correction of such misunderstandings in XR sports viewing. Through a formative study, we identified four recurring ambiguity types (referential, spatial, temporal, and metric) that characterize ambiguous spoken queries in this context. We develop a design space that organizes externalization along three dimensions (ambiguity type, interpretation state, externalization strategy) and instantiate it in an interactive XR soccer viewing system that combines situated visual cues with supporting analytic views. A within-subjects user study (N=16) comparing externalized interpretation against a voice-only baseline reveals that externalization is associated with higher inspectability on most measured dimensions and increased explicit repair language overall. However, repair occurred in only 38% of misaligned externalization trials, and this visibility-action gap varied by ambiguity type, indicating that transparency and correction affordance are orthogonal design axes.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
P. Andrews, O. E. Nordberg, S. Zubicueta Portales, N. Borch, F. Guribye, K. Fujita et al. Aicommentator: A multimodal conversational agent for embedded visualization in football viewing. InProceedings of the 29th International Conference on Intelligent User Interfaces, pp. 14–34, 2024. doi: 10.1145/3640543.3645197 2
arXiv 2024
-
[2]
G. Bansal, T. Wu, J. Zhou, R. Fok, B. Nushi, E. Kamar et al. Does the whole exceed its parts? the effect of ai explanations on complementary team performance. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems, pp. 1–16, 2021. doi: 10.1145/3411764. 3445717 3
doi:10.1145/3411764 2021
- [3]
-
[4]
V . Braun and V . Clarke. Using thematic analysis in psychology.Qual- itative Research in Psychology, 3(2):77–101, 2006. doi: 10.1191/ 1478088706qp063oa 7
work page 2006
-
[5]
Z. Buçinca, M. B. Malaya, and K. Z. Gajos. To trust or to think: Cog- nitive forcing functions can reduce overreliance on AI in AI-assisted decision-making.Proceedings of the ACM on Human-Computer Interac- tion, 5(CSCW1):1–21, 2021. doi: 10.1145/3449287 3, 7, 9
doi:10.1145/3449287 2021
-
[6]
K. B. Buldu, S. Özdel, K. H. C. Lau, M. Wang, D. Saad, S. Schönborn et al. Cuify the xr: An open-source package to embed llm-powered conversational agents in xr. In2025 IEEE International Conference on Artificial Intelligence and eXtended and Virtual Reality (AIxVR), pp. 192– 197, 2025. doi: 10.1109/AIxVR63409.2025.00037 1, 2
arXiv 2025
-
[8]
Z. Chen, S. Ye, X. Chu, H. Xia, H. Zhang, H. Qu et al. Augmenting sports videos with viscommentator.IEEE Transactions on Visualization and Computer Graphics, 28(1):824–834, 2021. doi: 10.1109/TVCG.2021. 3114806 1, 2
- [9]
Show all 74 references
-
[10]
Chromik and A
M. Chromik and A. Butz. Human-xai interaction: A review and design principles for explanation user interfaces. InHuman-Computer Interaction – INTERACT 2021, pp. 619–640. Springer, 2021. doi: 10.1007/978-3-030 -85616-8_36 3, 9
2021 doi
-
[11]
X. Chu, X. Xie, S. Ye, H. Lu, H. Xiao, Z. Yuan et al. Tivee: Visual exploration and explanation of badminton tactics in immersive visual- izations.IEEE Transactions on Visualization and Computer Graphics, 28(1):118–128, 2021. doi: 10.1109/TVCG.2021.3114861 2
2021
-
[12]
H. H. Clark and S. E. Brennan. Grounding in communication.American Psychological Association, 1991. doi: 10.1037/10096-006 2, 3, 4, 7, 8, 9
1991 doi
-
[13]
Corbin and A
J. Corbin and A. Strauss.Basics of Qualitative Research: Tech- niques and Procedures for Developing Grounded Theory. Sage Pub- lications, 4th ed., 2014. https://us.sagepub.com/en-us/nam/ basics-of-qualitative-research/book235578. 3
2014
-
[14]
A. K. Dey and J. Mankoff. Designing mediation for context-aware applica- tions.ACM Transactions on Computer-Human Interaction, 12(1):53–80,
-
[15]
Y . Feng, X. Wang, B. Pan, K. K. Wong, Y . Ren, S. Liu et al. Xnli: Ex- plaining and diagnosing nli-based visual data analysis.IEEE Transactions on Visualization and Computer Graphics, 30(7):3813–3827, 2024. doi: 10 .1109/TVCG.2023.3240003 2, 4
2024
-
[16]
Fu and J
Y . Fu and J. T. Stasko. Hoopinsight: Analyzing and comparing basketball shooting performance through visualization.IEEE Transactions on Visual- ization and Computer Graphics, 30:858–868, 2023. doi: 10.1109/TVCG. 2023.3326910 2
2023
-
[17]
T. Gao, M. Dontcheva, E. Adar, Z. Liu, and K. G. Karahalios. Datatone: Managing ambiguity in natural language interfaces for data visualization. InProceedings of the 28th annual acm symposium on user interface soft- ware & technology, pp. 489–500, 2015. doi: 10.1145/2807442.28...
2015
-
[18]
Y . Gao, H. Zhu, P. Ng, C. N. dos Santos, Z. Wang, F. Nan et al. Answering ambiguous questions through generative evidence fusion and round-trip prediction. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joi...
2021 doi
-
[19]
Giorgio, D
P. Giorgio, D. Jarvis, B. Auxier, H. Bobich, and K. Harwood. sports fan in- sights: The beginning of the immersive sports era.Deloitte Insights,
-
[20]
Heuer and A
H. Heuer and A. Breiter. More than accuracy: towards trustworthy machine learning interfaces for object recognition. InProceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization, pp. 298– 302, 2020. doi: 10.1145/3340631.3394873 7, 12
2020
-
[21]
J. D. Hincapié-Ramos, X. Guo, P. Moghadasian, and P. Irani. Consumed endurance: A metric to quantify arm fatigue of mid-air interactions. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI ’14, pp. 1063–1072. ACM, 2014. doi: 10.1145/2556288. 2557130 1
2014 doi
-
[22]
R. R. Hoffman, S. T. Mueller, G. Klein, and J. Litman. Measures for explainable ai: Explanation goodness, user satisfaction, mental models, curiosity, trust, and human-ai performance.Frontiers in Computer Science, 5:1096257, 2023. doi: 10.3389/fcomp.2023.1096257 3, 7, 8
2023
-
[23]
E. Horvitz. Principles of mixed-initiative user interfaces. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems, pp. 159–166, 1999. doi: 10.1145/302979.303030 3
1999
-
[24]
S. Jang, W. Stuerzlinger, S. Ambike, and K. Ramani. Modeling cumu- lative arm fatigue in mid-air interaction based on perceived exertion and kinetics of arm motion. InProceedings of the 2017 CHI Conference on Human Factors in Computing Systems, pp. 3328–3339, 2017. doi: 10. 11...
2017
-
[25]
J.-Y . Jian, A. M. Bisantz, and C. G. Drury. Foundations for an em- pirically determined scale of trust in automated systems.International Journal of Cognitive Ergonomics, 4(1):53–71, 2000. doi: 10.1207/ S15327566IJCE0401_04 7
-
[26]
F. Kern, F. Niebling, and M. E. Latoschik. Text input for non-stationary XR workspaces: Investigating tap and word-gesture keyboards in virtual and augmented reality.IEEE Transactions on Visualization and Computer Graphics, 29(5):2658–2669, 2023. doi: 10.1109/TVCG.2023.3247098 1
2023
-
[27]
Kiesel, A
J. Kiesel, A. Bahrami, B. Stein, A. Anand, and M. Hagen. Toward voice query clarification. InThe 41st International ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 1257–1260,
-
[28]
Knierim, T
P. Knierim, T. Kosch, J. Groschopp, and A. Schmidt. Opportunities and challenges of text input in portable virtual reality. InExtended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems, pp. 1–8, 2020. doi: 10.1145/3334480.3382920 1
2020
-
[29]
Lai and C
V . Lai and C. Tan. On human predictions with explanations and predictions of machine learning models: A case study on deception detection. InPro- ceedings of the Conference on Fairness, Accountability, and Transparency, pp. 29–38, 2019. doi: 10.1145/3287560.3287590 3, 7
2019
-
[30]
how do i fool you?
H. Lakkaraju and O. Bastani. "how do i fool you?": Manipulating user trust via misleading black box explanations. InProceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pp. 79–85. ACM, 2020. doi: 10. 1145/3375627.3375833 7, 12
2020
-
[31]
J. R. Landis and G. G. Koch. The measurement of observer agreement for categorical data.Biometrics, 33(1):159–174, 1977. doi: 10.2307/2529310 3
1977 doi
-
[32]
C. Lee, U. Gong, T. Lin, S. Zollmann, S. A. Epsley, A. Petway et al. Vair: Visual analytics for injury risk exploration in sports. In2025 IEEE 16th Workshop on Visual Analytics in Healthcare (VAHC), pp. 22–28. IEEE,
-
[33]
C. Lee, T. Lin, H. Pfister, and C. Zhu-Tian. Sportify: Question answering with embedded visualizations and personified narratives for sports video. IEEE Transactions on Visualization and Computer Graphics, 31(1):12–22,
-
[34]
C. Lee, H. Saiki, T. Lin, E. Ikeda, K. Suzuki, C. Zhu-Tian et al. Vistar: Virtual skill training with augmented reality with 3d avatars and llm coaching agent. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, pp. 1–16, 2026. doi: 10.1145/3772318....
2026 doi
-
[35]
J. Lee, S. S. Rodriguez, R. Natarrajan, J. Chen, H. Deep, and A. Kir- lik. What’s this? a voice and touch multimodal approach for ambiguity 10 © 2026 IEEE. This is the author’s version of the article that has been published in IEEE Transactions on Visualization and Computer Gr...
2026
-
[36]
J. Lee, J. Wang, E. Brown, L. Chu, S. S. Rodriguez, and J. E. Froehlich. Gazepointar: A context-aware multimodal voice assistant for pronoun disambiguation in wearable augmented reality. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pp. 1–20, ...
2024
-
[37]
doi: 10.1109/V AHC69430.2025.00008 2
2025
-
[38]
T. Lin, A. Aouididi, Z. Chen, J. Beyer, H. Pfister, and J.-H. Wang. VIRD: Immersive match video analysis for high-performance badminton coach- ing.IEEE Transactions on Visualization and Computer Graphics, 2023. doi: 10.1109/TVCG.2023.3327161 1, 2
2023
-
[40]
A. Liu, Z. Wu, J. Michael, A. Suhr, P. West, A. Koller et al. We’re afraid language models aren’t modeling ambiguity. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 790–807. Association for Computational Linguistics, 2023. doi: 10...
2023
-
[41]
Z. Liu, X. Xie, M. He, W. Zhao, Y . Wu, L. Cheng et al. Smartboard: Visual exploration of team tactics with llm agent.IEEE Transactions on Visualization and Computer Graphics, 31(1):23–33, 2025. doi: 10. 1109/TVCG.2024.3456200 2
2025
-
[42]
W. H. Lo, S. Zollmann, and H. Regenbrecht. Xrspectator: Immersive, augmented sports spectating. InProceedings of the 27th ACM Symposium on Virtual Reality Software and Technology, pp. 109:1–109:2, 2021. doi: 10.1145/3489849.3489930 1, 2, 4
2021
-
[43]
Q. V . Liao, H. Subramonyam, J. Wang, and J. W. Vaughan. Designerly un- derstanding: Information needs for model transparency to support design ideation for ai-powered user experience. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–21, 20...
2023
-
[44]
J. Ma, L. Shi, K. A. Robertsen, and P. Chi. Ambigchat: Interactive hierarchical clarification for ambiguous open-domain question answering. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology, pp. 1–18, 2025. doi: 10.1145/3746059.3747686 3
2025
-
[45]
Maathuis, M
C. Maathuis, M. A. Cidota, D. Datcu, and L. Marin. Integrating explain- able artificial intelligence in extended reality environments: A systematic survey.Mathematics, 13(2):290, 2025. doi: 10.3390/math13020290 3
2025 doi
-
[46]
S. Min, J. Michael, H. Hajishirzi, and L. Zettlemoyer. Ambigqa: Answer- ing ambiguous open-domain questions. InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pp. 5783–5797, 2020. doi: 10.18653/v1/2020.emnlp-main.466 1, 2, 4
2020 doi
-
[47]
Mohseni, N
S. Mohseni, N. Zarei, and E. D. Ragan. A multidisciplinary survey and framework for design and evaluation of explainable AI systems.ACM Transactions on Interactive Intelligent Systems, 11(3-4):1–45, 2021. doi: 10.1145/3387166 7, 8
2021 doi
-
[48]
M. J. Muller. Grounded theory method in hci and cscw. InCambridge Handbook of Research Methods in Human-Computer Interaction. Cam- bridge University Press, 2014. doi: 10.1201/b11963 3
2014 doi
-
[49]
A. G. Losada, R. Therón, and A. Benito. Bkviz: A basketball visual analysis tool.IEEE Computer Graphics and Applications, 36:58–68, 2016. doi: 10.1109/MCG.2016.124 2
2016 doi
-
[50]
Nourani, C
M. Nourani, C. Roy, J. E. Block, D. R. Honeycutt, T. Rahman, E. Ragan et al. Anchoring bias affects mental model formation and user reliance in explainable AI systems. InProceedings of the 26th International Conference on Intelligent User Interfaces, pp. 340–350. ACM, 2021. do...
2021
-
[51]
Y . Rong, T. Leemann, T.-T. Nguyen, L. Fiedler, P. Qian, V . Unhelkar et al. Towards human-centered explainable AI: A survey of user studies for model explanations.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(4):2104–2122, 2024. doi: 10.1109/TPAMI.2023.33...
2024
-
[52]
Saiki, C
H. Saiki, C. Lee, H. Takahashi, T. Lin, H. Kishi, K. Tachibana et al. Bridge: Borderless reconfiguration for inclusive and diverse gameplay experience via embodiment transformation. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems, pp. 1–17, 2026....
2026
-
[53]
Scharowski, S
N. Scharowski, S. A. C. Perrig, L. F. Aeschbach, N. von Felten, K. Opwis, P. Wintersberger et al. To trust or distrust AI: A questionnaire validation study. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. ACM, 2025. doi: 10.1145/3715275.3732025 7
2025
-
[54]
Setlur, M
V . Setlur, M. Tory, and A. Djalali. Inferencing underspecified natural language utterances in visual analysis. InProceedings of the 24th Interna- tional Conference on Intelligent User Interfaces, pp. 40–51, 2019. doi: 10. 1145/3301275.3302270 1, 2, 3, 4, 9
2019
-
[55]
Munzner.Visualization Analysis and Design
T. Munzner.Visualization Analysis and Design. CRC Press, 2014. doi: 10 .1145/3721241.3733989 5
2014
-
[56]
Y . Tang, J. Situ, A. Y . Cui, M. Wu, and Y . Huang. Llm integration in extended reality: A comprehensive review of current trends, challenges, and future perspectives. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 2025. doi: 10.1145/3706598. ...
2025 doi
-
[57]
Vasconcelos, M
H. Vasconcelos, M. Jörke, M. Grunde-McLaughlin, T. Gerstenberg, M. S. Bernstein, and R. Krishna. Explanations can reduce overreliance on ai systems during decision-making.Proceedings of the ACM on Human- Computer Interaction, 7(CSCW1):1–38, 2023. doi: 10.1145/3579605 3
2023 doi
-
[58]
Willett, Y
W. Willett, Y . Jansen, and P. Dragicevic. Embedded data representations. IEEE transactions on visualization and computer graphics, 23(1):461–470,
-
[59]
Y . Wu, D. Deng, X. Xie, M. He, J. Xu, H. Zhang et al. Obtracker: Visual analytics of off-ball movements in basketball.IEEE Transactions on Visualization and Computer Graphics, 29:929–939, 2022. doi: 10.1109/ TVCG.2022.3209373 2
2022
-
[60]
X. Xu, A. Yu, T. R. Jonker, K. Todi, F. Lu, X. Qian et al. Xair: A framework of explainable ai in augmented reality. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pp. 1–30,
2023
-
[61]
H. Song, K. Whitley, E. Krokos, and A. Varshney. Sia: A framework for context-aware intent clarification in speech-driven immersive analytics. InProceedings of the 31st International Conference on Intelligent User Interfaces, pp. 1687–1703, 2026. doi: 10.1145/3742413.3789063 1, 2, 3
2026
-
[62]
Zhi and R
Q. Zhi and R. A. Metoyer. Gamebot: A visualization-augmented chatbot for sports game. InExtended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems, pp. 1–7, 2020. doi: 10.1145/ 3334480.3382794 2
2020
-
[63]
Zhu-Tian, Q
C. Zhu-Tian, Q. Yang, J. Shan, T. Lin, J. Beyer, H. Xia et al. iball: Aug- menting basketball videos with gaze-moderated embedded visualizations. InProceedings of the 2023 CHI Conference on Human Factors in Com- puting Systems, pp. 841:1–841:18, 2023. doi: 10.1145/3544548.3581...
2023
-
[64]
Zollmann, T
S. Zollmann, T. Langlotz, M. Loos, W. H. Lo, and L. Baker. Arspectator: Exploring augmented reality for sport events. InSIGGRAPH Asia 2019 technical briefs, pp. 75–78. 2019. doi: 10.1145/3355088.3365162 2
2019
-
[65]
O” indicates the region is used; “X
S. Zollmann, W. H. Lo, T. Langlotz, and H. Regenbrecht. Stats on- site—sports spectator experience through situated visualizations.Comput- ers & Graphics, 99:13–27, 2021. doi: 10.1016/j.cag.2021.12.009 2 11 A SUPPLEMENTARYMATERIAL A.1 Survey of Existing XR Sports Viewing Appli...
2021 doi
-
[68]
doi: 10.1145/3544548.3581500 3, 4, 5, 9
-
[69]
Zhi and R
Q. Zhi and R. Metoyer. Gameviews: Understanding and supporting data-driven sports storytelling. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology, pp. 471–483,
-
[75]
[UI State] - Real-time match context: match_time, half, score, possession, clip_events (chronological order, last = most recent, with player_id = spawn ID), team rosters (each player has name, num, pos, and id = spawn ID)
-
[76]
[Reference Data] - Relevant context selected per question: - Match Detail: lineups, goals, cards, substitutions, fouls, set pieces, individual player stats, match ratings - Player Profiles: bio, nationality, height, position, career stats, scouting notes, debut info, youth car...
2020
-
[2005]
doi: 10.1145/1057237.1057241 3
-
[2016]
doi: 10.1109/TVCG.2016.2598608 5
2016
-
[2018]
doi: 10.1145/3209978.3210160 2
-
[2019]
doi: 10.1145/3290605.3300499 1, 2
-
[2023]
https://www.deloitte.com/us/en/insights/industry/ sports/immersive-sports-fandom.html. 1, 3
-
[2024]
doi: 10.1109/TVCG.2024.3456332 1, 2, 3, 6, 12
2024
-
[2025]
doi: 10.1038/s41597-025-04505-y 1, 6
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.