REVIEW 3 major objections 4 minor 34 references
Adaptive AR interfaces must be evaluated over trajectories—sequences of shifting contexts with longitudinal trust tracking—not single-snapshot sessions.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Adaptive AR interface evaluation should shift from one-shot snapshots to trajectory-based benchmarks covering task interference, context transitions, and longitudinal trust.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection A strong position paper whose benchmark does not yet measure what it claims: the distribution of contexts an adaptive interface actually encounters is endogenous, not scripted. the 3 major comments →
You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that an interface whose behavior over time is only partially specified at design time cannot be validly evaluated at a point; it must be evaluated along the trajectory of contexts it will actually face. Current practice remains snapshot-based—fixed conditions, single sessions, aggregate workload—so it cannot see trade-offs between concurrent tasks, adaptation latency at transitions, or trust recalibration across repeated use. The proposed remedy is a shared evaluation scaffold in which each scenario specifies a physical task, a digital task, at least one scripted context transition, and a mapping onto a criticality-by-coupling-by-contention coordinate system, and
What carries the argument
The trajectory is the carrying object: a sequence of contexts with at least one parameterized transition (crowd-density escalation, interruption, tracking loss, hazard injection), sampled from a standardized distribution of unpredictability rather than rigidly scripted events. Three outputs attach to it. The Performance Operating Characteristic (POC) frontier is the empirically measured set of achievable primary-task/secondary-task performance pairs, made visible by varying adaptation parameters or comparing interfaces. Average regret is the mean Euclidean distance, in a min-max normalized [0,1] performance space, between observed performance and the frontier across all N context points. The
Load-bearing premise
The framework assumes that the contexts an adaptive interface will actually encounter can be represented as a sampleable distribution of scripted, parameterized transitions (escalating crowd density, interruptions, tracking loss, injected hazards) without losing what makes real contexts distinct.
What would settle it
Run a three-session longitudinal study of two adaptive AR notification interfaces and collect both a single-snapshot evaluation (fixed context, SUS/TLX, task metrics) and a trajectory evaluation (POC frontier, average regret, override-rate recovery). If the snapshot metrics predict the trajectory outcomes—trust recovery and regret ranking—as well as the trajectory metrics do, the claim that snapshots are structurally insufficient would be falsified for that interface class.
If this is right
- Single-session, fixed-condition studies become insufficient evidence for any interface that adapts at runtime; evaluations must co-report primary and secondary task metrics.
- POC frontiers make trade-offs explicit: an interface that improves one task while degrading the other is no longer mistaken for a pure win.
- Average regret gives a trajectory-native score that is independent of session length, so studies can compare how well different adaptation logics navigate the same distribution of context shifts.
- Trust-sensitive evaluations require at least three sessions with injected errors and recovery measurement; failure of override rates to recover indicates broken trust that snapshot usability scores cannot see.
- Benchmark cells spanning criticality, coupling, and contention give the community a shared scaffold so results from different adaptive AR systems become comparable.
Where Pith is reading between the lines
- Beyond the paper, the trajectory logic likely applies to any adaptive user interface—voice assistants, recommender systems, alerting systems—wherever the designer specifies adaptation logic rather than the full set of runtime states.
- Beyond the paper, average regret's dependence on the choice of normalization (min-max vs. z-score vs. domain weighting) means cross-task comparisons are only meaningful after the community fixes conventions; until then regret is safest within matched task pairs.
- Beyond the paper, the per-task workload argument implies a testable bound: an adaptive system trained on aggregate workload labels cannot learn to throttle one task without throttling the other, so per-task ground truth is a precondition for task-sensitive adaptation.
- Beyond the paper, field trajectories collected through passive sensing would provide the true context distribution; comparing them with lab-sampled parameterized trajectories would directly test whether the benchmark's standardized unpredictability preserves ecological validity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that AI-mediated heads-up AR interfaces ('fluid interfaces') cannot be validly evaluated with the traditional fixed-interface paradigm of one context, one session, and aggregate metrics. It identifies four gaps in current OST-HMD multitasking evaluation practice: interference is not measured directly, contexts are frozen, time is absent, and ground-truth workload does not decompose across tasks. The paper proposes a trajectory-based evaluation framework with three components: Performance Operating Characteristic (POC) interference frontiers, an average regret metric over context points, and longitudinal trust measurement. A four-cell benchmark instantiation is presented, including a worked example for an adaptive notification manager (NOTIFADAPT).
Significance. If the framework is accepted, it could shift evaluation practice for adaptive AR interfaces and provide a concrete reporting standard for the community. The paper makes a genuine conceptual contribution by articulating the structural mismatch between interface dynamics and evaluation protocols, and it is commendably transparent about the limitations of its proposal (Section 3.2). The worked example and the regret formula are clearly specified, and the paper explicitly frames the benchmark as a first draft for community refinement. The main challenges are that the operationalization does not fully deliver the paper's central promise—evaluation under the distribution the interface will actually encounter—and the empirical claim that current practice is 'largely snapshot-based' rests on selected examples rather than a documented systematic review.
major comments (3)
- [Section 1 and Section 3, Table 2] The core definition of trajectory-based evaluation is 'sampled from the distribution the interface will actually encounter,' but the proposed benchmark feeds every interface the same pre-scripted context schedule (e.g., crowd-density escalation, tracking loss). A fluid interface changes the contexts it encounters: an aggressive notification-suppression policy can slow the user and extend exposure to crowded segments; an error-prone overlay can increase hazard exposure. Thus the 'distribution actually encountered' is endogenous, and average regret computed under the fixed schedule measures performance under an exogenous test trajectory, not under deployment. Section 3.2 acknowledges ecological-validity approximations but does not address this closed-loop issue. The authors should either justify the exogenous schedule as a conservative test, add closed-loop simulation/field trajectories, o
- [Section 2 and Abstract] The abstract and Section 2 assert that 'current evaluation practice remains largely snapshot-based' on the basis of 'a systematic reading' of the literature. No search protocol, inclusion criteria, or quantitative synthesis is reported; the support is a set of selected examples (GlassMessaging, ParaGlassMenu, etc.). Because this empirical generalization motivates the proposal, it should be either substantiated by a documented systematic review or explicitly downgraded to an observation about common patterns. As written, the statement overstates the evidence.
- [Table 2 and Section 3] The abstract defines a trajectory as 'a sequence of contexts with transitions,' but each benchmark cell contains exactly one scripted context transition (e.g., crowd-density increase, tracking loss). The worked example (Section 3.1) uses a single transition and computes regret over two segments. This does not instantiate the promised sequence of shifts, and it cannot capture adaptation latency over repeated transitions or within-session trust dynamics. The authors should extend cells to multiple transitions or clearly label the current version as a minimal template rather than the proposed trajectory benchmark.
minor comments (4)
- [Equation (1)] The regret formula is split across lines in the text; use display math. Also clarify whether 'distance to the frontier' means Euclidean distance to the nearest point on the Pareto frontier or distance to the ideal point (1,1).
- [Section 3.1] The phrase 'achieves a hypothetical1' uses a footnote marker that reads like a typo; consider rewording to avoid ambiguity.
- [Section 2.5] The rationale for the three task meta-properties (criticality, coupling, contention) is plausible but not validated. A short justification for why these three, not others, are the axes would help.
- [Section 2] References [9] and [22] are systematic literature reviews; citing them directly to support the claim of snapshot-based practice would strengthen the argument.
Circularity Check
No significant circularity: the trajectory-evaluation argument is a self-contained measurement critique; its illustrative numbers are explicitly hypothetical and its self-citations are examples, not premises.
full rationale
The paper argues for a shift from snapshot to trajectory evaluation of fluid AR interfaces. The central claim is a methodological argument (Section 1, Section 3) rather than a derivation from fitted quantities. The only numerical example (NOTIFADAPT, Section 3.1) is explicitly hypothetical ('Values are illustrative'), so no fitted input is renamed as a prediction. The POC frontier and average regret are defined as measurement constructs; the definition of regret as mean distance to an empirically measured frontier is not circular because the frontier is constructed from fixed policies and the adaptive system is scored against it independently. The paper explicitly flags the normalization choice as an open issue (Section 3.2), so it does not hide a convention as a result. Self-citations ([2], [3], [11], [31]) are used as illustrative examples of existing systems or as a general motivating claim about adaptation; they are not invoked as a uniqueness theorem, a fitted constraint, or the sole justification for the framework. The potential gap identified by a skeptical reader — that 'the distribution the interface will actually encounter' is endogenous while the benchmark standardizes scripted transitions — is a validity/correctness concern about ecological validity (acknowledged in Section 3.2: parameterized transitions 'remain approximations of naturalistic context shifts'), not a circular reduction of the paper's argument to its own inputs. No step in the paper exhibits the pattern X defined in terms of Y, or a fitted parameter reported as a prediction. Therefore no significant circularity is present.
Axiom & Free-Parameter Ledger
free parameters (2)
- Illustrative NOTIFADAPT performance values =
hypothetical (e.g., 0.97, 3.1; 0.96, 3.0; 0.90, 2.6)
- Normalization convention for average regret =
Min-max scaling to [0,1]
axioms (6)
- domain assumption Adaptive AI-mediated interfaces have behavior that is only partially specified at design time
- domain assumption The contexts a fluid interface will encounter can be modeled as a sampleable distribution of scripted transitions
- domain assumption User trust forms, evolves, and deteriorates on multi-session timescales and is measurable
- domain assumption Per-task workload decomposition is needed for adaptation decisions and cannot be obtained from aggregate measures
- standard math Euclidean distance in a min-max normalized [0,1] space is a valid measure of regret
- domain assumption The four identified patterns characterize current OST-HMD multitasking evaluation practice
invented entities (2)
-
Average regret metric (trajectory-level performance measure)
no independent evidence
-
Task meta-properties coordinate system (criticality x coupling x contention)
no independent evidence
Cite this review
Pith. "Pith review of You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction." pith.science (2026). https://pith.science/paper/VVPOGNWA
@misc{pith2026260801656,
author = {Pith},
title = {Pith review of: You Cannot Optimize What You Cannot Measure: Multitasking Evaluation as the Missing Foundation of AI-Mediated Heads-Up Interaction},
year = {2026},
howpublished = {\url{https://pith.science/paper/VVPOGNWA}},
note = {Machine review of arXiv:2608.01656}
}
read the original abstract
AI-mediated heads-up augmented reality (AR) replaces fixed interfaces with dynamically adapting ones that decide what information to present, in what form, and when, based on a continually changing context that cannot be fully anticipated beforehand. Although it remains an interface, its behavior over time is only partially specified at design time. We argue that this shift requires a corresponding change in evaluation: from snapshots to trajectories. A fixed interface is evaluated in a snapshot --- one context, one session, one set of task-performance metrics. A fluid interface must be evaluated over a trajectory --- a sequence of contexts with transitions, sampled from the distribution the interface will actually encounter, and tracked long enough for user trust to form, evolve, and potentially deteriorate. Drawing on the literature for heads-up AR multitasking enabled by optical see-through head-mounted displays (OST-HMDs), we find that current evaluation practice remains largely snapshot-based. Most studies use fixed-condition, single-session designs; interference between concurrent tasks is rarely quantified directly; and commonly used workload measures cannot disentangle cognitive load attributable to individual tasks. To address these limitations, we argue for three shifts: from isolated metrics to Performance Operating Characteristic (POC) interference frontiers, from fixed conditions to evaluation over context trajectories, and from single-session snapshots to longitudinal trust measurement.
Reference graph
Works this paper leans on
-
[1]
D. Ariansyah, J. A. Erkoyuncu, I. Eimontaite, T. Johnson, A.-M. Oost- veen, S. Fletcher, and S. Sharples. A head mounted augmented reality design practice for maintenance assembly: Toward meeting perceptual and cognitive needs of AR users.Applied Ergonomics, 98:103597, Jan. 2022. doi: 10.1016/j.apergo.2021.103597 2, 4
-
[3]
R. Cai, N. Janaka, S. Zhao, and M. Sun. ParaGlassMenu: Towards Social-Friendly Subtle Interactions in Conversations. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, pp. 1–21. Association for Computing Machinery, New York, NY , USA, Apr. 2023. doi: 10.1145/3544548.3581065 2
arXiv 2023
-
[4]
Y . Cheng, Y . Yan, X. Yi, Y . Shi, and D. Lindlbauer. SemanticAdapt: Optimization-based Adaptation of Mixed Reality Layouts Leverag- ing Virtual-Physical Semantic Connections. InThe 34th Annual ACM Symposium on User Interface Software and Technology, UIST ’21, pp. 282–297. Association for Computing Machinery, New York, NY , USA, 2021. doi: 10.1145/347274...
arXiv 2021
-
[5]
Y . Choi and Y . S. Kim. An Adaptive UI Based on User-Satisfaction Prediction in Mixed Reality.Applied Sciences (Switzerland), 12(9),
-
[6]
Y . B. Eisma, C. D. D. Cabrall, and J. C. F. de Winter. Visual Sampling Processes Revisited: Replicating and Extending Senders (1983) Us- ing Modern Eye-Tracking Equipment.IEEE Transactions on Human- Machine Systems, 48(5):526–540, Oct. 2018. doi: 10.1109/THMS. 2018.2806200 2
arXiv 1983
-
[7]
R. Guarese, J. Becker, H. Fensterseifer, M. Walter, C. Freitas, L. Nedel, and A. Maciel. Augmented Situated Visualization for Spa- tial and Context-Aware Decision-Making. InProceedings of the 2020 International Conference on Advanced Visual Interfaces, A VI ’20. As- sociation for Computing Machinery, New York, NY , USA, 2020. doi: 10.1145/3399715.3399838 2, 3
-
[8]
J. He, W. Choi, J. S. McCarley, B. S. Chaparro, and C. Wang. Texting while driving using Google Glass™: Promising but not distraction- free.Accident Analysis and Prevention, 81:218 – 229, 2015. doi: 10. 1016/j.aap.2015.03.033 2, 3
work page 2015
- [9]
-
[10]
R. Jain, J. Shi, R. Duan, Z. Zhu, X. Qian, and K. Ramani. Ubi- TOUCH: Ubiquitous Tangible Object Utilization through Consistent Hand-object interaction in Augmented Reality. InProceedings of the 36th Annual ACM Symposium on User Interface Software and Tech- nology, UIST ’23. Association for Computing Machinery, New York, NY , USA, 2023. doi: 10.1145/35861...
arXiv 2023
-
[11]
N. Janaka, J. Gao, L. Zhu, S. Zhao, L. Lyu, P. Xu, M. Nabokow, S. Wang, and Y . Ong. GlassMessaging: Towards Ubiquitous Mes- saging Using OHMDs.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, 7(3):100:1–100:32, Sept. 2023. doi: 10.1145/3610931 1, 2, 3, 4
-
[12]
M. Jannat, P. Dhaka, K. Katsuragawa, and K. Hasan. Exploring Aug- mented Reality User Interface Transitions Across Mid-Air, On-Body and Physical Surfaces. In2024 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 1177–1186, Oct. 2024. doi: 10.1109/ISMAR62088.2024.00134 2
arXiv 2024
-
[13]
J. Kuge, T. Grundgeiger, P. Schlosser, P. Sanderson, and O. Hap- pel. Design and Evaluation of a Head-Worn Display Application for Multi-Patient Monitoring. InProceedings of the 2021 ACM Design- ing Interactive Systems Conference, DIS ’21, pp. 879–890. Associa- tion for Computing Machinery, New York, NY , USA, 2021. doi: 10. 1145/3461778.3462011 2, 3
arXiv 2021
-
[14]
J. Lee, T. Lim, and W. Kim. Investigating the Usability of Collabo- rative Robot Control Through Hands-Free Operation Using Eye Gaze and Augmented Reality. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 4101–4106, Oct. 2023. doi: 10.1109/IROS55552.2023.10342045 2
arXiv 2023
-
[15]
J. Li, S. Wang, G. Wang, J. Zhang, S. Feng, Y . Xiao, and S. Wu. The effects of complex assembly task type and assembly experience on users’ demands for augmented reality instructions.International Jour- nal of Advanced Manufacturing Technology, 131(3-4):1479 – 1496,
- [16]
-
[17]
F. Lu, S. Davari, L. Lisle, Y . Li, and D. A. Bowman. Glanceable AR: Evaluating Information Access Methods for Head-Worn Augmented Reality. In2020 IEEE Conference on Virtual Reality and 3D User In- terfaces (VR), pp. 930–939, Mar. 2020. doi: 10.1109/VR46266.2020. 00113 2
arXiv 2020
-
[18]
F. Lu, L. Pavanatto, and D. A. Bowman. In-the-Wild Experiences with an Interactive Glanceable AR System for Everyday Use. InPro- ceedings of the 2023 ACM Symposium on Spatial User Interaction, SUI ’23. Association for Computing Machinery, New York, NY , USA,
work page 2023
-
[19]
F. Lu and Y . Xu. Exploring Spatial UI Transition Mechanisms with Head-Worn Augmented Reality. InProceedings of the 2022 CHI Con- ference on Human Factors in Computing Systems, CHI ’22. Associa- tion for Computing Machinery, New York, NY , USA, 2022. doi: 10. 1145/3491102.3517723 2
-
[20]
F. Lu, Y . Xu, X. Xu, B. Jones, and L. Malamed. Exploring the Im- pact of User and System Factors on Human-AI Interactions in Head- Worn Displays. In2023 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), pp. 109–118, Oct. 2023. doi: 10.1109/ ISMAR59233.2023.00025 2
-
[21]
P. Manakhov, L. Sidenmark, K. Pfeuffer, and H. Gellersen. Gaze on the Go: Effect of Spatial Reference Frame on Visual Target Acquisi- tion During Physical Locomotion in Extended Reality. InProceed- ings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI ’24. Association for Computing Machinery, New York, NY , USA, 2024. doi: 10.1145/361...
- [22]
-
[23]
F. Necker, M. L. Melcher, S. Busque, C. W. Leuze, P. Ghanouni, C. Le Castillo, E. Nguyen, and B. L. Daniel. Nested Semi-Transparent Isosurface Simulated V olume-Rendering (NESTIS-VR) – An efficient on-device rendering approach for Augmented Reality headsets in- creasing surgeon confidence of kidney donor arterial anatomy.Com- puters in Biology and Medicin...
- [24]
-
[25]
Z. Stefanidi, G. Margetis, S. Ntoa, and G. Papagiannakis. Real- Time Adaptation of Context-Aware Intelligent User Interfaces, for En- hanced Situational Awareness.IEEE Access, 10:23367–23393, 2022. doi: 10.1109/ACCESS.2022.3152743 2
- [26]
-
[27]
F. Tan, P. Xu, A. Ram, W. Z. Suen, S. Zhao, Y . Huang, and C. Hurter. AudioXtend: Assisted Reality Visual Accompaniments for Audio- book Storytelling During Everyday Routine Tasks. InProceedings of the 2024 CHI Conference on Human Factors in Computing Sys- tems, CHI ’24, pp. 1–22. Association for Computing Machinery, New York, NY , USA, May 2024. doi: 10....
arXiv 2024
-
[28]
C. D. Wickens, W. S. Helton, J. G. Hollands, and S. Banbury.Engi- neering Psychology and Human Performance. Routledge, New York, 5 ed., Sept. 2021. doi: 10.4324/9781003177616 2
-
[29]
C. D. Wickens, R. Martin-Emerson, and I. Larish. Attentional tunnel- ing and the head-up display. Apr. 1993. 2
work page 1993
-
[30]
K. Zhang, B. R. Cochran, R. Chen, L. Hartung, B. Sprecher, R. Tredin- nick, K. Ponto, S. Banerjee, and Y . Zhao. Exploring the Design Space of Optical See-through AR Head-Mounted Displays to Support First Responders in the Field. InProceedings of the 2024 CHI Confer- ence on Human Factors in Computing Systems, CHI ’24. Association for Computing Machinery,...
-
[31]
S. Zhao, F. Tan, and K. Fennedy. Heads-Up Computing Moving Be- yond the Device-Centered Paradigm.Communications of the ACM, 66(9):56–63, Aug. 2023. doi: 10.1145/3571722 1
doi:10.1145/3571722 2023
- [2020]
-
[2022]
doi: 10.3390/app12094559 2
- [2023]
-
[2024]
doi: 10.1007/s00170-024-13091-z 2, 3
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.