Pith. sign in

REVIEW 1 major objections 5 minor 115 references

Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap

T0 review · 1 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read An AR X-ray that embeds AI-suggested targets into the real world led people to over-rely on the AI and choose worse than they did with a 2D minimap, while still improving their spatial mapping.

desk verdict Genuinely new empirical result on AR + AI reliance, but the headline conclusion is muddied by a legibility confound and some unbalanced exclusions. read the letter →

arxiv 2507.14316 v5 pith:IWBWEEZU submitted 2025-07-18 cs.HC

classification cs.HC
keywords AugmentedRealityHuman-AIcollaborationAppropriaterelianceOver-relianceEmbeddedvisualizationsSpatialdecision-makingX-rayvisualizationEmpiricalstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests a widely held assumption: that embedding AI suggestions directly into the physical world with augmented reality will help people rely on the AI more appropriately when making quick spatial choices. In a controlled study of 32 people choosing among four coffee machines in a two-floor building under a 20-second deadline, the AR X-ray view—which drew targets and AI cues in place in the environment—actually produced lower decision accuracy than a 2D minimap, driven mainly by over-reliance on bad AI suggestions. The same X-ray view did improve spatial mapping, since participants made far fewer errors when pointing to the target they had chosen. If these results hold, they imply that embedding AI cues into the real world can miscalibrate trust and undercut decision quality even while improving spatial awareness, and that AR designers should not assume that spatial embedding is inherently beneficial for human-AI collaboration.

What carries the argument

The load-bearing machinery is the experimental contrast between two visualizations of identical decision data—an AR X-ray that overlays a 1:1 scale three-dimensional digital twin of the building onto the real world, and a 2D Minimap that shows a top-down abstraction—combined with a simulated AI that suggests a single optimal target at 75% accuracy. The study uses a within-subjects 2×2 design (visualization × AI availability) with 32 participants and 1024 trials, a 20-second time limit, and a task in which participants choose among four coffee machines by trading off walking distance and queue length. The key metric is the reliance classification: appropriate reliance (following correct or overriding incorrect AI), over-reliance (accepting suboptimal or worst suggestions), and under-reliance (rejecting correct suggestions for a worse choice), with over-reliance counts compared against a random-selection baseline of 0.75 trials per block. The argument runs through the quantitative contrast in over-reliance between conditions, supported by post-trial pointing error rates and qualitative interview themes.

What would settle it

A replication that makes the X-ray as legible as the minimap—numeric queue counts, no occluded queues, all four targets visible from the start—would settle the interpretation: if over-reliance disappears when verification is easy, the effect is a rational response to noisier evidence; if it persists, the embeddedness account holds.

Watch

Extended reading notes

Core claim

The study's central claim is that in time-critical spatial target selection with imperfect AI support, the AR X-ray embedded visualization led to greater inappropriate reliance on AI, primarily over-reliance, compared with a 2D Minimap, contradicting all three of the authors' hypotheses. Participants using the X-ray accepted suboptimal or worst AI suggestions far more often—on average 1.41 over-reliance trials per block versus 0.72 under the Minimap, and above the 0.75 random-choice baseline—while showing fewer under-reliance trials. The authors attribute this to occlusion in the see-through view, difficulty estimating walking distances and queue lengths in a large environment, a visual proximity illusion in which targets behind walls appear closer than the walking path actually is, and to heightened trust in the realistic, embodied presentation of AI suggestions. The X-ray did, however, significantly reduce post-trial pointing errors, demonstrating a real benefit in spatial mapping. The authors conclude that embedding AI cues into physical space is not inherently beneficial and can miscalibrate reliance, and they argue the X-ray's strength lies in action-oriented spatial tasks rather than isolated selection decisions.

Load-bearing premise

The study's interpretation assumes the over-reliance comes from the act of embedding, rather than from the X-ray making the underlying data (distances and queue lengths) harder to verify than the minimap did; the paper's own results note that targets were often occluded and distances and queues were hard to judge.

Editorial extensions

If this is right

  • Designers should not assume that embedding AI cues in AR is inherently beneficial; perceptual challenges such as occlusion and distance misjudgment make it harder to verify the AI, so embedded cues need explicit verification support.
  • Less embodied or less realistic representations of AI suggestions may reduce over-reliance, because several participants reported trusting the AI more precisely because its suggestion was rendered as a realistic object in the scene.
  • AR X-ray views appear better matched to action-oriented spatial tasks such as navigation, evacuation, and first response, where improved spatial mapping directly matters, than to isolated selection decisions.
  • Faster decisions under X-ray+AI after removing initial search time suggest embedding may lower deliberative engagement, so deliberation prompts such as cognitive forcing functions may be needed in AR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the over-reliance is largely a response to the higher cost of verifying the AI in the X-ray view, then a decision-theoretic measure that separates cognitive limits from reliance behavior would reframe much of the observed 'inappropriate' reliance as a rational adjustment to noisier evidence rather than a distinct bias caused by embeddedness.
  • A testable extension the paper does not run: increasing X-ray legibility (numeric queue counts, unobstructed views, explicit path overlays) should reduce over-reliance, which would isolate the embeddedness effect from the verification-cost effect.
  • The spatial-mapping advantage hints at a hybrid division of labor: use the X-ray for the action phase (walking, pointing, executing) and keep a map-like verification panel for the decision phase, potentially combining the X-ray's mapping benefit with the minimap's better scrutiny of the AI.
  • Because participants reported that realistic, embodied AI suggestions felt more trustworthy, the finding likely extends to other high-fidelity AR presentations, such as annotations anchored to physical objects, suggesting AR may amplify the persuasive weight of AI output compared with flat screens.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. This paper reports a within-subjects user study (N=32) comparing an AR X-ray embedded visualization against a 2D Minimap in an AI-assisted spatial decision-making task. Participants selected one of four coffee machines in a two-floor building under a 20-second time limit, balancing walking distance and queue length, with and without AI suggestions (simulated AI with 75% accuracy). The authors hypothesized that the embedded X-ray would improve decision accuracy, promote more appropriate reliance, and shorten response times. All three hypotheses were contradicted: X-ray led to lower decision accuracy, greater inappropriate reliance on AI (mostly over-reliance), and slower raw response times, while also improving spatial mapping as measured by pointing errors. Qualitative interviews attribute the over-reliance to perceptual challenges (occlusion, distance estimation), visual proximity illusions, and heightened trust in embodied AI suggestions. The authors discuss design implications and call for better AR-AI integration.

Significance. If the central claim holds, the paper provides an important, counter-intuitive empirical result that challenges the common assumption that embedding AI suggestions directly into the physical environment reduces cognitive load and fosters appropriate reliance. The study is methodologically careful in several respects: a real two-floor environment with an Apple Vision Pro, a counterbalanced within-subjects design, a pre-registered-style hypothesis set that was openly contradicted, simulated AI with controlled accuracy, high inter-rater reliability for video annotations, and a public repository for the system. The finding that the X-ray improves spatial mapping while impairing decision accuracy is a useful dual-effect result for the AR and human-AI collaboration communities. The paper is honest about its limitations and provides a reasonable foundation for future work, provided the central interpretation is appropriately qualified.

major comments (1)
  1. [Section 5.1, first paragraph] The exclusion of 55 trials (5.4%) due to 20-second timeouts is not neutral across conditions: 19 trials were excluded in each X-ray condition (X-ray+AI and X-ray+NoAI) versus only 8 and 9 in the Minimap conditions. The authors do not provide any sensitivity analysis. If timeouts are more likely when participants are confused or unable to verify the AI suggestion, excluding them could systematically bias accuracy and reliance estimates in the X-ray conditions. For example, if participants who timed out were more likely to have been deliberating between the AI suggestion and a better alternative, their exclusion could inflate or deflate the measured over-reliance rate. I request a robustness check that codes timed-out trials under alternative assumptions (e.g., as errors, as accepting the AI suggestion, as rejecting it) and reports whether the key conclusions in Sections 5.1.1 and 5.1.2 change. Without this, the main statistical results rest on a potentially non-ignorable missingness mechanism.
minor comments (5)
  1. [Section 5.1.2] The random-selection baseline for under-reliance is reported as 4.25 expected trials per block, but the formal definition in Section 4.6 defines under-reliance as rejecting correct AI suggestions in favor of a worse option. Under that definition, with 5 optimal, 2 suboptimal, and 1 worst suggestions per block, the expected number of under-reliance trials under random choice would be 5 × 3/4 = 3.75, not 4.25. The value 4.25 appears to count any choice worse than the AI suggestion, even when the AI itself is suboptimal. Please align the baseline calculation with the stated definition or clarify the discrepancy.
  2. [Section 5.1.2] The claim that inappropriate reliance in the X-ray+AI condition was 'primarily driven by over-reliance' rests on a descriptive comparison to the random baselines (over-reliance mean 1.41 vs. baseline 0.75; under-reliance mean 1.63 vs. baseline 4.25). The raw mean under-reliance (1.63) is actually larger than the mean over-reliance (1.41). The conclusion would be stronger if the authors reported a formal test of whether over-reliance exceeds under-reliance or at least discussed this apparent tension explicitly.
  3. [Section 5.1.1, Figure 7] Several post-hoc contrasts are reported as F values with negative magnitudes, e.g., F(1,155) = -3.37 and F(1,31) = -3.27. Since F is non-negative, these appear to be t- or z-values mislabeled as F. Please correct the notation and report the appropriate test statistic.
  4. [Section 5.1.3] The response-time analysis that excludes initial search time uses the assumption that 'meaningful decision-making starts only after at least two options have been identified.' This assumption, though referenced to prior work, is debatable and the manual annotation of S1/S2 segments could introduce bias. The claim that 'with AI assistance, decisions were made more quickly with the X-ray' is used in Section 6.4 to support a cognitive-engagement argument. Given the ad hoc nature of the exclusion, this finding should be labeled as exploratory or supported by an alternative analysis (e.g., using total response time as a conservative bound).
  5. [Section 6.5] The limitations section does not explicitly acknowledge the confound between the visual embedding and the lower legibility/accessibility of decision data in the X-ray condition. Adding a sentence that this confound limits the generalizability of the over-reliance conclusion would be important for readers.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical study, hypotheses contradicted by results, reliance measures pre-defined, and self-citations are not load-bearing.

full rationale

This paper is an empirical user study; there is no derivation chain whose conclusion is equivalent to its premises. The reliance taxonomy (Sec 4.6 and Figure 6) defines appropriate, over-, and under-reliance from the alignment between human choices and the simulated AI's suggestions; this is a measurement definition, not a fitted prediction. The AI suggestion schedule (5 optimal, 2 suboptimal, 1 worst per block) was fixed before data collection (Sec 4.7), and the random-selection baselines (0.75 over-reliance, 4.25 under-reliance per block) are computed from that fixed schedule, so the comparison is not a post-hoc fit. The quantitative hypotheses H1-H3 predicted the opposite of the observed outcome, making a fit-to-result explanation implausible. Self-citations (e.g., [56], [86], [108]-[110]) are used for design background, apparatus, or prior art; none carries the central claim. The acknowledged limitations in Sec 6.5 and the qualitative reports of occlusion and perceptual difficulty (Sec 5.2.1) are validity or confound concerns about whether 'embedding' rather than reduced information legibility drives over-reliance; they do not reduce the result to its inputs by definition. No circular step was found.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The central claim rests on several design choices and analytical rules. The AI accuracy (0.75) and task constants (SPEED, SERVICE_TIME, time limit) are control parameters set from prior work and pilot studies; they determine which target is optimal and thus which choices count as over- or under-reliance. The more fragile entries are the post-hoc S1/S2 search-time exclusion and the implicit assumption that the X-ray and Minimap provide equally accessible decision information. These are noted in the axioms above.

free parameters (5)
  • AI suggestion distribution = 5 optimal / 2 suboptimal / 1 worst per 8 trials (75% accuracy)
    Chosen to fix AI accuracy at 0.75, following prior work and pilot; participants never saw the distribution, so it is an experimental control rather than fit to outcomes.
  • SPEED (walking speed constant) = 1 m/s
    Empirically set from pilot studies; used in the T_coffee model to determine target optimality.
  • SERVICE_TIME (per-person waiting time) = 15 s
    Empirically set from pilot studies; used to compute total coffee time and target labels.
  • Decision time limit = 20 s
    Fixed decision window; chosen to create time pressure.
  • Initial search exclusion rule (S1/S2) = Time until first and second targets appear in field of view
    Hand-coded post-hoc segmentation that removes search time from response time in X-ray conditions; directly affects the claim that X-ray+AI decisions are faster.
assumptions (6)
  • domain assumption The coffee-machine selection task validly abstracts real-world time-pressured spatial decision-making (emergency evacuation, security, navigation).
    Section 3.1 defines the task abstraction from literature; the generalization of findings depends on this mapping.
  • domain assumption The taxonomy of reliance (appropriate, over-, under-reliance) based on agreement between human choice and AI suggestion optimality is a valid operationalization of AI reliance.
    Section 4.6 defines the classification following [103]; all reliance conclusions depend on it.
  • domain assumption A simulated AI with 75% accuracy behaves like a real AI assistant for the purpose of this study.
    Section 4.7 justifies simulation following [10, 85]; no feedback or accuracy information was given to participants.
  • ad hoc to paper Decision-making begins only once at least two candidate targets have been visually identified, so response time can be split into search time and decision time.
    Section 5.1.3 introduces the S1/S2 exclusion rule after observing raw response times; it reverses the response-time finding and is acknowledged in Section 6.5 as an approximation.
  • domain assumption The two visualizations present the same decision-relevant information with comparable accessibility, so observed differences in reliance are attributable to embeddedness.
    Sections 3.2 and 4.1 describe the two displays; qualitative results in Sections 5.2.1 and 6.1 show the X-ray was harder to perceive, which challenges this comparability assumption.
  • domain assumption Dual-process theory (Type 1 heuristic vs Type 2 analytical thinking) appropriately explains the response time, confidence, and reliance patterns.
    Used in Section 4.8 to motivate hypotheses and in Section 6.4 to interpret lower cognitive engagement; a standard but non-trivial psychological model.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap." pith.science (2026). https://pith.science/paper/IWBWEEZU

@misc{pith2026250714316,
  author       = {Pith},
  title        = {Pith review of: Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IWBWEEZU}},
  note         = {Machine review of arXiv:2507.14316}
}
read the original abstract

Artificial Intelligence (AI) and indoor sensing increasingly support decision-making in spatial environments. However, traditional visualization methods impose a substantial mental workload when viewers translate this digital information into real-world spaces, leading to inappropriate reliance on AI. Embedded visualizations in Augmented Reality (AR), by integrating information into physical environments, may reduce this workload and foster more appropriate reliance on AI. To assess this, we conducted an empirical study (N = 32) comparing an AR embedded visualization (X-ray) and 2D Minimap in AI-assisted, time-critical spatial target selection tasks. Surprisingly, evidence shows that the embedded visualization led to greater inappropriate reliance on AI, primarily as over-reliance, due to factors like perceptual challenges, visual proximity illusions, and highly realistic visual representations. Nonetheless, the embedded visualization demonstrated benefits in spatial mapping. We conclude by discussing empirical insights, design implications, and directions for future research on human-AI collaborative decision in AR.

Figures

Figures reproduced from arXiv: 2507.14316 by the authors.

Figure 1
Figure 1. A participant wearing an Apple Vision Pro performs spatial decision-making tasks using two different visualizations: [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. An example spatial arrangement of the coffee ma￾chine targets locations for a single participant location, illus￾trating four distinct difficulty levels: same+close, same+far, cross+close, and cross+far. For each participant location, there are 4 trials, one difficulty level per trial. The partici￾pant’s location is marked in red, and colored dots indicate target locations in each difficulty level. The white dashed … view at source ↗
Figure 3
Figure 3. The 8 participant locations in the study, distributed [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: A stimulus workflow from an actual trial. Each trial began with a 20-second countdown. Participants viewed either the X-ray or Minimap visualization while searching for the optimal target. After identifying their choice, participants verbally announced the selected tar…
Figure 5
Figure 5. Figure 5: One example of what a participant experienced dur [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Classification of reliance as appropriate, over-, and [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Summary of quantitative results across conditions. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 9
Figure 9. Figure 9: (a) and (b) show the direct and walking distance to [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Example of occlusion in the AR X-ray visualiza￾tion, where overlapping avatars make it difficult to estimate the number of people in the queue, leading to potential mis￾judgment of wait time. Summary. Compared to the Minimap conditions, participants in the X-ray condi…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

115 extracted references · 46 canonical work pages

  1. [1]

    Fadel Adib and Dina Katabi. 2013. See through walls with WiFi!SIGCOMM Comput. Commun. Rev.43, 4 (Aug. 2013), 75–86. doi:10.1145/2534169.2486039

  2. [2]

    Melanie Bancilhon, Lace Padilla, and Alvitta Ottley. 2023. Evaluating Visu- alization Decision-Making with Cognitive Models.Visualization Psychology (2023)

  3. [3]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld. 2021. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proc. of CHI (CHI ’21). ACM, Article 81, 16 pages. doi:10.1145/3411764.3445717

  4. [4]

    Cindy Xiong Bearfield, Lisanne Van Weelden, Adam Waytz, and Steven Fran- coneri. 2024. Same data, diverging perspectives: The power of visualizations to elicit competing interpretations.IEEE TVCG(2024)

  5. [5]

    Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis. 2020. A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy. InProc. of CHI

  6. [6]

    2015.A Survey of Augmented Reality

    Mark Billinghurst, Adrian Clark, and Gun Lee. 2015.A Survey of Augmented Reality. now. doi:10.1561/1100000049

  7. [7]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis.Qualitative research in sport, exercise and health11, 4 (2019), 589–597

  8. [9]

    Brumar, Sam Molnar, Gabriel Appleby, Kristi Potter, and Remco Chang

    Camelia D. Brumar, Sam Molnar, Gabriel Appleby, Kristi Potter, and Remco Chang. 2024. A Typology of Decision-Making Tasks for Visualization. arXiv:2404.08812 [cs.HC] https://arxiv.org/abs/2404.08812

Show all 115 references
  1. [10]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.Proc. ACM Hum.-Comput. Interact.(2021). doi:10. 1145/3449287

  2. [11]

    2003.Virtual reality technology

    Grigore C Burdea and Philippe Coiffet. 2003.Virtual reality technology. John Wiley & Sons

  3. [12]

    Samuel Carton, Qiaozhu Mei, and Paul Resnick. 2020. Feature-based explana- tions don’t help people detect misclassifications of online toxicity. InProc. of AAAI

  4. [13]

    Harris Chaiklin. 2012. Thinking Fast and Slow.Journal of Nervous and Mental Disease200 (2012), 826. https://api.semanticscholar.org/CorpusID:168706774

  5. [14]

    Remco Chang, Mohammad Ghoniem, Robert Kosara, William Ribarsky, Jing Yang, Evan Suma, Caroline Ziemkiewicz, Daniel Kern, and Agus Sudjianto

  6. [15]

    Kurtis Danyluk, Barrett Ens, Bernhard Jenny, and Wesley Willett. 2021. A Design Space Exploration of Worlds in Miniature. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems(Yokohama, Japan)(CHI ’21). Association for Computing Machinery, New York, NY...

  7. [16]

    Dietvorst, Joseph P

    Berkeley J. Dietvorst, Joseph P. Simmons, and Cade Massey. 2014. Algorithm Aversion: People Erroneously Avoid Algorithms after Seeing Them Err.CSN: Business(2014). https://api.semanticscholar.org/CorpusID:1646733

  8. [17]

    Evanthia Dimara, Steven Franconeri, Catherine Plaisant, Anastasia Bezerianos, and Pierre Dragicevic. 2018. A task-based taxonomy of cognitive biases for information visualization.IEEE TVCG(2018)

  9. [18]

    Evanthia Dimara and John Stasko. 2022. A Critical Reflection on Visualization Research: Where Do Decision Making Tasks Hide?IEEE TVCG(2022). doi:10. 1109/IEEETVCG.2021.3114813

  10. [19]

    Weihua Dong, Yulin Wu, Tong Qin, Xinran Bian, Yan Zhao, Yanrou He, Yawei Xu, and Cheng Yu. 2021. What is the difference between augmented reality and 2D navigation electronic maps in pedestrian wayfinding?Cartography and Geographic Information Science(2021)

  11. [20]

    David Drascic and Paul Milgram. 1996. Perceptual issues in augmented reality. InElectronic imaging. https://api.semanticscholar.org/CorpusID:7734379

  12. [21]

    Dzindolet, Linda G

    Mary T. Dzindolet, Linda G. Pierce, Hall P. Beck, and Lloyd A. Dawe. 2002. The Perceived Utility of Human and Automated Aids in a Visual Detection Task.Hu- man Factors44, 1 (2002), 79–94. arXiv:https://doi.org/10.1518/0018720024494856 doi:10.1518/0018720024494856 PMID: 12118875

  13. [22]

    Elkin, Matthew Kay, James J

    Lisa A. Elkin, Matthew Kay, James J. Higgins, and Jacob O. Wobbrock. 2021. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. InUIST. ACM. doi:10.1145/3472749.3474784

  14. [23]

    Beatrix Emo, Christoph Hoelscher, Jan Malte Wiener, and Ruth Conroy Dalton

  15. [24]

    Jonathan St. B. T. Evans and Keith E. Stanovich. 2013. Dual-Process Theories of Higher Cognition: Advancing the Debate.Perspectives on Psychological Science (2013). doi:10.1177/1745691612460685

  16. [25]

    Michael Fernandes, Logan Walls, Sean Munson, Jessica Hullman, and Matthew Kay. 2018. Uncertainty displays using quantile dotplots or cdfs improve transit decision-making. InProc. of CHI

  17. [26]

    Hall, Yuriy Brun, and Cindy Xiong Bearfield

    Aimen Gaba, Zhanna Kaufman, Jason Cheung, Marie Shvakel, Kyle Wm. Hall, Yuriy Brun, and Cindy Xiong Bearfield. 2024. My Model is Unfair, Do People Even Care? Visual Design Affects Trust and Perceived Bias in Machine Learning. IEEE TVCG(2024). doi:10.1109/IEEETVCG.2023.3327192

  18. [27]

    Romeo Giuliano, Franco Mazzenga, Marco Petracca, and Marco Vari. 2013. Indoor localization system for first responders in emergency scenario.IWCMC (2013)

  19. [28]

    Junglas, Blake Ives, and Norman A

    Lakshmi Goel, Iris A. Junglas, Blake Ives, and Norman A. Johnson. 2012. Decision-making in-socio and in-situ: Facilitation in virtual worlds.Decis. Support Syst.52 (2012), 342–352. https://api.semanticscholar.org/CorpusID: 46128169

  20. [29]

    Guerrieri, Michael H

    Jeffrey R. Guerrieri, Michael H. Francis, P. F. Wilson, T. Kos, L. E. Miller, Nelson P. Bryner, D. W. Stroup, and L T. Klein-berndt. 2006. RFID-assisted indoor localiza- tion and communication for first responders.2006 First European Conference on Antennas and Propagation(2006...

  21. [30]

    Ziyang Guo, Yifan Wu, Jason D Hartline, and Jessica Hullman. 2024. A decision theoretic framework for measuring AI reliance. InProc. of FAccAT

  22. [31]

    Sunwoo Ha, Shayan Monadjemi, and Alvitta Ottley. 2024. Guided By AI: Nav- igating Trust, Bias, and Data Exploration in AI-Guided Visual Analytics. In Computer Graphics Forum. Wiley Online Library

  23. [32]

    Newcombe

    Justin Harris, Kathy Hirsh-Pasek, and Nora S. Newcombe. 2013. Understanding spatial transformations: similarities and differences between mental rotation and mental folding.Cognitive Processing(2013). https://api.semanticscholar. org/CorpusID:6072708

  24. [33]

    Mordechai I Henig and John T Buchanan. 1996. Solving MCDM problems: Process concepts.Journal of multi-criteria decision analysis(1996)

  25. [34]

    Teresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz, Pascal Knierim, and Kasper Hornbæk. 2023. When XR and AI Meet - A Scoping Review on Extended Reality and Artificial Intelligence. InProc. of CHI. ACM. doi:10.1145/ 3544548.3581072

  26. [35]

    Kevin Anthony Hoff and Masooda Bashir. 2015. Trust in Automation: Integrating Empirical Evidence on Factors That Influence Trust.Human Factors57, 3 (2015), 407–434. arXiv:https://doi.org/10.1177/0018720814547570 doi:10.1177/ 0018720814547570 PMID: 25875432

  27. [36]

    Sathaporn Hu, Joseph Malloch, and Derek Reilly. 2021. A Comparative Eval- uation of Techniques for Locating Out of View Targets in Virtual Reality. In Graphics Interface 2021

  28. [37]

    Petra Isenberg, Lionel Reveret, and Romain Vuillemot. 2024. HomeOlympics: Turning Everyday Outdoor Physical Activities into Olympic-Style Experience. In Proc. of the IEEE VIS Workshop on First-Person Visualizations for Outdoor Physical Activities. Virtual, Unknown Region. http...

  29. [38]

    Jakobsen and Kasper Hornbæk

    Mikkel R. Jakobsen and Kasper Hornbæk. 2013. Interactive Visualizations on Large and Small Displays: The Interrelation of Display Size, Information Space, and Scale.IEEE Transactions on Visualization and Computer Graphics19, 12 (2013), 2336–2345. doi:10.1109/TVCG.2013.170

  30. [39]

    Yvonne Jansen, Pierre Dragicevic, Petra Isenberg, Jason Alexander, Abhijit Karnik, Johan Kildal, Sriram Subramanian, and Kasper Hornbæk. 2015. Op- portunities and Challenges for Data Physicalization. InProceedings of the 33rd Annual ACM Conference on Human Factors in Computing...

  31. [40]

    Guan, and Maya Gupta

    Heinrich Jiang, Been Kim, Melody Y. Guan, and Maya Gupta. 2018. To trust or not to trust a classifier. InProceedings of the 32nd International Conference on Neural Information Processing Systems(Montréal, Canada)(NIPS’18). Curran Associates Inc., 5546–5557. CHI ’26, April 13–1...

  32. [41]

    Alex Kale, Matthew Kay, and Jessica Hullman. 2020. Visual reasoning strategies for effect size judgments and decisions.IEEE TVCG(2020)

  33. [42]

    Denis Kalkofen, Erick Mendez, and Dieter Schmalstieg. 2009. Comprehensible Visualization for Augmented Reality.IEEE Transactions on Visualization and Computer Graphics15, 2 (2009), 193–204. doi:10.1109/TVCG.2008.96

  34. [43]

    Kangsoo Kim, Luke Boelling, Steffen Haesler, Jeremy Bailenson, Gerd Bruder, and Greg F. Welch. 2018. Does a Digital Assistant Need a Body? The Influence of Visual Embodiment and Social Behavior on the Perception of Intelligent Virtual Agents in AR. InProc. ISMAR. doi:10.1109/I...

  35. [44]

    Roberta L. Klatzky. 1998. Allocentric and Egocentric Spatial Representations: Definitions, Distinctions, and Interconnections. InSpatial Cognition. https: //api.semanticscholar.org/CorpusID:7264069

  36. [45]

    Terry K Koo and Mae Y Li. 2016. A guideline of selecting and reporting intraclass correlation coefficients for reliability research.Journal of chiropractic medicine 15, 2 (2016), 155–163

  37. [46]

    Julian Kreiser, Alexander Hann, Eugen Zizer, and Timo Ropinski. 2017. Decision graph embedding for high-resolution manometry diagnosis.IEEE TVCG(2017)

  38. [47]

    J Richard Landis and Gary G Koch. 1977. The measurement of observer agree- ment for categorical data.biometrics(1977), 159–174

  39. [48]

    David Launder and Chad Perry. 2014. A study identifying factors influencing decision making in dynamic emergencies like urban fire and rescue settings. In A study identifying factors influencing decision making in dynamic emergencies like urban fire and rescue settings. https:...

  40. [49]

    I Want It That Way

    Connor Lawless, Jakob Schoeffer, Lindy Le, Kael Rowan, Shilad Sen, Cristina St. Hill, Jina Suh, and Bahareh Sarrafzadeh. 2024. “I Want It That Way”: Enabling Interactive Decision Support Using Large Language Models and Constraint Programming.TIIS(2024). doi:10.1145/3685053

  41. [50]

    Benjamin Lee, Michael Sedlmair, and Dieter Schmalstieg. 2023. Design patterns for situated visualization in augmented reality.IEEE TVCG(2023)

  42. [51]

    Chunggi Lee, Tica Lin, Hanspeter Pfister, and Chen Zhu-Tian. 2024. Sportify: question answering with embedded visualizations and personified narratives for sports video.IEEE Transactions on Visualization and Computer Graphics (2024)

  43. [52]

    Jaewook Lee, Fanjie Jin, Younsoo Kim, and David Lindlbauer. 2022. User Preference for Navigation Instructions in Mixed Reality. InIEEE VR. 802–811. doi:10.1109/VR51125.2022.00102

  44. [53]

    Jiahao Nick Li, Yan Xu, Tovi Grossman, Stephanie Santosa, and Michelle Li. 2024. OmniActions: Predicting Digital Actions in Response to Real-World Multimodal Sensory Inputs with LLMs. InProc. of CHI. ACM. doi:10.1145/3613904.3642068

  45. [54]

    Zhenlong Li and Huan Ning. 2023. Autonomous GIS: the next-generation AI-powered GIS. arXiv:2305.06453 [cs.AI] https://arxiv.org/abs/2305.06453

  46. [55]

    Tica Lin, Alexandre Aouididi, Chen Zhu-Tian, Johanna Beyer, Hanspeter Pfister, and Jui-Hsien Wang. 2023. VIRD: Immersive Match Video Analy- sis for High-Performance Badminton Coaching.IEEE TVCG(2023). https: //api.semanticscholar.org/CorpusID:260125958

  47. [56]

    Xiaoan Liu, Difan Jia, Xianhao Carton Liu, Mar Gonzalez-Franco, and Chen Zhu-Tian. 2025. Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations.arXiv preprint arXiv:2507.17248(2025)

  48. [57]

    X-Ray Vision

    Mark A. Livingston, Arindam Dey, Christian Sandor, and B. Thomas. 2013. Pursuit of “X-Ray Vision” for Augmented Reality. InPursuit of “X-Ray Vision” for Augmented Reality. https://api.semanticscholar.org/CorpusID:16721214

  49. [58]

    Helena Löfström. 2023. On the Definition of Appropriate Trust and the Tools that Come with it.2023 Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE)(2023), 1555–1562. https://api.semanticscholar.org/ CorpusID:262084200

  50. [59]

    Shuai Ma, Ying Lei, Xinru Wang, Chengbo Zheng, Chuhan Shi, Ming Yin, and Xiaojuan Ma. 2023. Who should i trust: Ai or myself? leveraging human and ai correctness likelihood to promote appropriate trust in ai-assisted decision- making. InProc. of CHI

  51. [60]

    Manfredi, Nicola Felice Capece, R

    G. Manfredi, Nicola Felice Capece, R. P. Di, and Ugo Erra. 2024. A Mixed Reality Application for Multi-Floor Building Evacuation Drills using Real-Time Pathfind- ing and Dynamic 3D Modeling. https://api.semanticscholar.org/CorpusID: 273970950

  52. [61]

    Katelyn Morrison, Donghoon Shin, Kenneth Holstein, and Adam Perer. 2023. Evaluating the impact of human explanation strategies on human-AI visual decision-making.Proc. of HCI(2023)

  53. [62]

    R Nydegger, L Nydegger, and F Basile. 2011. Post-traumatic stress disorder and coping among career professional firefighters.American Journal of Health Sciences(2011)

  54. [63]

    Emre Oral, Ria Chawla, Michel Wijkstra, Narges Mahyar, and Evanthia Dimara

  55. [64]

    Lace M Padilla, Sarah H Creem-Regehr, Mary Hegarty, and Jeanine K Stefanucci

  56. [65]

    Lace M. K. Padilla, Sarah H. Creem-Regehr, Mary Hegarty, and Jeanine K. Ste- fanucci. 2018. Decision making with visualizations: a cognitive framework across disciplines.Cognitive Research: Principles and Implications(2018)

  57. [66]

    Stephan Pajer, Marc Streit, Thomas Torsney-Weir, Florian Spechtenhauser, Torsten Möller, and Harald Piringer. 2016. Weightlifter: Visual weight space exploration for multi-criteria decision making.IEEE TVCG(2016)

  58. [67]

    Greg Penney, David Launder, Joe Cuthbertson, and Matthew B Thompson

  59. [68]

    Leon Pietschmann, Paul-David Joshua Zuercher, Erik Bub’ik, Chen Zhu-Tian, Hanspeter Pfister, and Thomas Bohné. 2023. Quantifying the Impact of XR Visual Guidance on User Performance Using a Large-Scale Virtual Assembly Ex- periment.IEEE VIS(2023). https://api.semanticscholar.o...

  60. [69]

    George Pu, Paul Wei, Amanda Aribe, James Boultinghouse, Nhi Dinh, Fang Xu, and Jing Du. 2021. Seeing Through Walls: Real-Time Digital Twin Modeling of Indoor Spaces. In2021 Winter Simulation Conference (WSC). doi:10.1109/ WSC52266.2021.9715364

  61. [70]

    Carlos Quijano-Chavez, Nina Doerr, Benjamin Lee, Dieter Schmalstieg, and Michael Sedlmair. 2024. Brushing and Linking for Situated Analytics. In2024 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). IEEE, 597–603

  62. [71]

    Priya Raghubir and Aradhna Krishna. 1996. As the crow flies: Bias in consumers’ map-based distance judgments.Journal of Consumer Research(1996)

  63. [72]

    Nils Reisen, Ulrich Hoffrage, and Fred W Mast. 2008. Identifying decision strategies in a consumer choice situation.Judgment and decision making(2008)

  64. [73]

    Jonatan Reyes, Anil Ufuk Batmaz, and Marta Kersten-Oertel. 2025. Trusting AI: does uncertainty visualization affect decision-making?Frontiers in Computer Science(2025)

  65. [74]

    J Edward Russo and France Leclerc. 1994. An eye-fixation analysis of choice processes for consumer nondurables.Journal of consumer research(1994)

  66. [75]

    David Saffo, Sara Di Bartolomeo, Tarik Crnovrsanin, Laura South, Justin Raynor, Caglar Yildirim, and Cody Dunne. 2024. Unraveling the Design Space of Im- mersive Analytics: A Systematic Review.IEEE TVCG30, 1 (2024), 495–506. doi:10.1109/IEEETVCG.2023.3327368

  67. [76]

    Sara Salimzadeh, Gaole He, and Ujwal Gadiraju. 2024. Dealing with Uncertainty: Understanding the Impact of Prognostic Versus Diagnostic Tasks on Trust and Reliance in Human-AI Decision Making. InProc. of CHI. ACM. doi:10.1145/ 3613904.3641905

  68. [77]

    Max Schemmer, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger

  69. [78]

    Sharad Sharma, James Stigall, and Sri Teja Bodempudi. 2020. Situational Awareness-based Augmented Reality Instructional (ARI) Module for Building Evacuation. InVRW. doi:10.1109/VRW50115.2020.00020

  70. [79]

    Ritsos, and Niklas Elmqvist

    Sungbok Shin, Andrea Batch, Peter William Scott Butcher, Panagiotis D. Ritsos, and Niklas Elmqvist. 2024. The Reality of the Situation: A Survey of Situated Analytics.IEEE TVCG(Aug 2024). doi:10.1109/IEEETVCG.2023.3285546

  71. [80]

    Siegel and Sheldon H

    Alexander W. Siegel and Sheldon H. White. 1975. The development of spatial representations of large-scale environments.Advances in child development and behavior(1975). https://api.semanticscholar.org/CorpusID:28635993

  72. [81]

    Herbert A. Simon. 1977.The Logic of Heuristic Decision Making. Springer Netherlands. doi:10.1007/978-94-010-9521-1_10

  73. [82]

    Ortega-Alvarado, and Francisco Ramón Feito- Higueruela

    Gregorio Soria, Lidia M. Ortega-Alvarado, and Francisco Ramón Feito- Higueruela. 2018. Augmented and Virtual Reality for Underground Facilities Management.J. Comput. Inf. Sci. Eng.(2018). https://api.semanticscholar.org/ CorpusID:116308226

  74. [83]

    Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations. InProc. of IUI. ACM. doi:10.1145/3581641.3584066

  75. [85]

    Gajos, and Finale Doshi-Velez

    Siddharth Swaroop, Zana Buçinca, Krzysztof Z. Gajos, and Finale Doshi-Velez

  76. [86]

    Haoyu Tan, Tongyu Nie, and Evan Suma Rosenberg. 2024. Invisible Mesh: Effects of X-Ray Vision Metaphors on Depth Perception in Optical-See-Through Augmented Reality.IEEE VR(2024). https://api.semanticscholar.org/CorpusID: 269175599

  77. [87]

    Barbara Tversky. 2003. Structures Of Mental Spaces: How People Think About Space.Environment and Behavior(2003). doi:10.1177/0013916502238865

  78. [88]

    2008.Spatial Cognition: Embodied and Situated

    Barbara Tversky. 2008.Spatial Cognition: Embodied and Situated. Cambridge University Press

  79. [89]

    Conway, and Randy Pausch

    Richard Stoakley, Matthew J. Conway, and Randy Pausch. 1995. Virtual reality on a WIM: interactive worlds in miniature. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Denver, Colorado, USA)(CHI ’95). ACM Press/Addison-Wesley Publishing Co., USA, 2...

  80. [90]

    Emily Wall, Subhajit Das, Ravish Chawla, Bharath Kalidindi, Eli T Brown, and Alex Endert. 2017. Podium: Ranking data using mixed-initiative visual analytics. Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? CHI ’26, April 13–17, 2026...

  81. [91]

    Emily Wall, Laura Matzen, Mennatallah El-Assady, Peta Masters, Helia Hossein- pour, Alex Endert, Rita Borgo, Polo Chau, Adam Perer, Harald Schupp, et al. 2024. Trust Junk and Evil Knobs: Calibrating Trust in AI Visualization. InPacificVis. IEEE

  82. [92]

    Emily Wall, Arpit Narechania, Adam Coscia, Jamal Paden, and Alex Endert

  83. [93]

    Qianwen Wang, Kexin Huang, Payal Chandak, Marinka Zitnik, and Nils Gehlen- borg. 2022. Extending the nested model for user-centric xai: A design study on gnn-based drug repurposing.IEEE TVCG(2022)

  84. [94]

    Amelia C Warden, Christopher D Wickens, Domenick Mifsud, Shannon Ourada, Benjamin A Clegg, and Francisco R Ortega. 2022. Visual search in augmented reality: Effect of target cue type and location. InProceedings of the human factors and ergonomics society annual meeting, Vol. 6...

  85. [95]

    Di Weng, Heming Zhu, Jie Bao, Yu Zheng, and Yingcai Wu. 2018. Homefinder revisited: Finding ideal homes with reachability-centric multi-criteria decision making. InProc. of CHI

  86. [96]

    Bernstein, and Ranjay Krishna

    Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making.Proc. ACM Hum.-Comput. Interact.CSCW1 (2023). doi:10.1145/3579605

  87. [97]

    Wesley Willett, Yvonne Jansen, and Pierre Dragicevic. 2017. Embedded Data Representations.IEEE TVCG(2017). doi:10.1109/IEEETVCG.2016.2598608

  88. [98]

    Wobbrock, Leah Findlater, Darren Gergle, and James J

    Jacob O. Wobbrock, Leah Findlater, Darren Gergle, and James J. Higgins. 2011. The aligned rank transform for nonparametric factorial analyses using only anova procedures. InProc. of CHI. ACM. doi:10.1145/1978942.1978963

  89. [99]

    Tim Wächter, Jan Rexilius, and Matthias König. 2022. Interactive evacuation in intelligent buildings assisted by mixed reality.Journal of Smart Cities and Society(2022). arXiv:https://journals.sagepub.com/doi/pdf/10.3233/SCS-220009 doi:10.3233/SCS-220009

  90. [100]

    Fang Xu, Tianyu Zhou, Tri Nguyen, and Jing Du. 2024. Augmented reality in team-based search and rescue: Exploring spatial perspectives for enhanced navigation and collaboration.Safety Science(2024). doi:10.1016/j.ssci.2024. 106556

  91. [101]

    Fang Xu, Tianyu Zhou, Hengxu You, and Jing Du. 2024. Improving indoor wayfinding with AR-enabled egocentric cues: A comparative study.Advanced Engineering Informatics59 (2024), 102265. doi:10.1016/j.aei.2023.102265

  92. [102]

    Liuchang Xu, Shuo Zhao, Qingming Lin, Luyao Chen, Qianqian Luo, Sensen Wu, Xinyue Ye, Hailin Feng, and Zhenhong Du. 2024. Evaluating Large Language Models on Spatial Tasks: A Multi-Task Benchmarking Study.ArXiv(2024). https://api.semanticscholar.org/CorpusID:271956990

  93. [103]

    Fumeng Yang, Zhuanyi Huang, Jean Scholtz, and Dustin L Arendt. 2020. How do visual explanations foster end users’ appropriate trust in machine learning?. InProc. of IUI

  94. [104]

    Büchner, and Christoph Hölscher

    Jan Malte Wiener, Simon J. Büchner, and Christoph Hölscher. 2009. Taxonomy of Human Wayfinding Tasks: A Knowledge-Based Approach.Spatial Cognition & Computation(2009). https://api.semanticscholar.org/CorpusID:16529538

  95. [105]

    Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach. 2019. Understanding the Effect of Accuracy on Trust in Machine Learning Models. InProc. of CHI. ACM. doi:10.1145/3290605.3300509

  96. [106]

    Kexin Zhang, Brianna R Cochran, Ruijia Chen, Lance Hartung, Bryce Sprecher, Ross Tredinnick, Kevin Ponto, Suman Banerjee, and Yuhang Zhao. 2024. Ex- ploring the Design Space of Optical See-through AR Head-Mounted Displays to Support First Responders in the Field. InProc. of CH...

  97. [107]

    Jieqiong Zhao, Yixuan Wang, Michelle V Mancenido, Erin K Chiou, and Ross Maciejewski. 2023. Evaluating the impact of uncertainty visualization on model reliance.IEEE TVCG(2023)

  98. [108]

    Chen Zhu-Tian, Daniele Chiappalupi, Tica Lin, Yalong Yang, Johanna Beyer, and Hanspeter Pfister. 2024. RL-LABEL: A Deep Reinforcement Learning Approach Intended for AR Label Placement in Dynamic Scenarios.TVCG(2024). doi:10. 1109/IEEETVCG.2023.3326568

  99. [109]

    Chen Zhu-Tian, Qisen Yang, Jiarui Shan, Tica Lin, Johanna Beyer, Haijun Xia, and Hanspeter Pfister. 2023. iBall: Augmenting Basketball Videos with Gaze- Moderated Embedded Visualizations. InProc. of CHI

  100. [110]

    Chen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer, Haijun Xia, Yingcai Wu, and Hanspeter Pfister. 2023. Sporthesia: Augmenting Sports Videos Using Natural Language.IEEE Trans. Vis. Comput. Graph.29, 1 (2023), 918–928. doi:10. 1109/TVCG.2022.3209497

  101. [112]

    Lijie Yao, Anastasia Bezerianos, Romain Vuillemot, and Petra Isenberg. 2022. Visualization in Motion: A Research Agenda and Two Evaluations.IEEE Trans. Vis. Comput. Graph.(2022). doi:10.1109/IEEETVCG.2022.3184993

  102. [2007]

    Wirevis: Visualization of categorical, time-varying data from financial transactions. InV AST. IEEE

  103. [2012]

    In Wayfinding and Spatial Configuration: evidencefrom street corners

    Wayfinding and Spatial Configuration: evidencefrom street corners. In Wayfinding and Spatial Configuration: evidencefrom street corners. https://api. semanticscholar.org/CorpusID:128846748

  104. [2018]

    Decision making with visualizations: a cognitive framework across disci- plines.Cognitive research: principles and implications(2018)

  105. [2021]

    Left, right, and gender: Exploring interaction traces to mitigate human biases.IEEE TVCG(2021)

  106. [2022]

    Threat assessment, sense making, and critical decision-making in police, military, ambulance, and fire services.Cognition, Technology & Work(2022)

  107. [2024]

    Accuracy-Time Tradeoffs in AI-Assisted Decision Making under Time Pressure. InIUI. ACM. doi:10.1145/3640543.3645206

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.