REVIEW 3 major objections 6 minor 35 references
Exploring AR Label Placements in Visually Cluttered Scenarios
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read In visually cluttered AR scenes, replacing per-item labels with one grouped label per cluster makes users faster.
desk verdict A first controlled study of grouped in-view AR labels that honestly rejects its own H2, but its headline recommendation overreaches because the experiment never varies data values within a group. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The design space is defined by two choices: whether labels are individual or grouped, and where a grouped label sits. Grouped layouts are built by DBSCAN clustering on item positions, using a user-set Euclidean distance to gather same-type items into spatial groups; a Situated Grouped label is anchored above the whole cluster, and a Centered Grouped label sits at the cluster's center, pushed away from items by a force-directed method. The study's three tasks, Identify, Compare, and Summarize, vary how many target types and subgroups a participant must track, and the outcome measures are completion time, accuracy, NASA-TLX workload, and rank preference.
What would settle it
A real-environment AR study with 50 or more objects, free head motion, and partial occlusion would falsify the central claim if Situated Grouped labels turn out no faster than Situated Individual labels, or produce lower accuracy.
Extended reading notes
Core claim
The central claim is that a Situated Grouped label placement, where one representative label sits above each spatial cluster of same-type items, outperforms the conventional Situated Individual approach of labeling every item separately. Participants completed the three tasks significantly faster with Situated Grouped (14.0s ± 1.51s) than with Situated Individual (16.5s ± 2.61s), at p < 0.05, while accuracy stayed statistically equivalent across all three layouts. Centered Grouped labels, placed at the cluster's center, were also faster than individual labels and statistically indistinguishable from Situated Grouped, so the paper rejects its second hypothesis that centering would help further. The paper concludes that using a label for spatially grouped items helps users identify, compare, and summarize data in terms of completion time.
Load-bearing premise
The load-bearing premise is that a static VR screen with simple same-size objects reproduces the visual clutter and search behavior of real augmented-reality scenes; if it does not, the grouped-label speed advantage may not carry over.
Editorial extensions
If this is right
- If the finding holds, AR designers should prefer one label per spatial group of same-type items when many items fill the user's field of view.
- Grouped labels deliver their speed benefit without lowering accuracy: the study found no statistically significant accuracy difference among the three placements.
- Centering the group label does not improve on placing it above the group, so leader-line length can be traded against label-item overlap.
- The benefit extends across the three investigated tasks: identifying a target type, comparing ratings within one type, and summarizing multiple types.
Reading between the lines
- The paper leaves the grouping threshold as a user-adjusted parameter; an automated threshold based on clutter level and object size is a natural next step and is already flagged by the authors.
- Because the study fixed the field of view and prohibited head movement, the results say nothing yet about out-of-view group labels or about how head motion and occlusion in real AR change the trade-off.
- A testable extension is to apply the grouped-label idea to dynamic scenes, where clusters move and the representative label must track them, building on the dynamic-scene line of prior label-placement work.
- The preference data are suggestive but not statistically significant: more participants chose Centered Grouped, while Situated Grouped won on completion time, so designers should weigh objective speed against subjective clarity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Park et al. present a user study comparing three AR label placement techniques—Situated Individual (SI), Situated Grouped (SG), and Centered Grouped (CG)—for in-view objects in visually cluttered scenes. The study uses a VR simulation with 50 items, 15 participants, three tasks (Identify, Compare, Summarize), and a Latin-square counterbalanced within-subject design. The main quantitative finding is that SG labels yield significantly faster completion times than SI labels (14.0s ± 1.51s vs 16.5s ± 2.61s, p < 0.05), with equivalent accuracy across conditions and no significant difference between SG and CG. The authors conclude that using a representative label for spatially grouped items helps users identify, compare, and summarize data.
Significance. If the result holds, the paper offers a simple and actionable guideline for AR label placement in cluttered scenes: prefer a single situated label per spatial group over per-item labels, with no measured accuracy cost. The study's strengths include a task set grounded in established visualization task taxonomies, Latin-square counterbalancing, mixed-effects modeling of accuracy, explicit rejection of H2, and a limitations section that candidly acknowledges the VR-to-AR simulation caveat. However, the experimental materials enforce value homogeneity within each spatial group, so the central claim's scope is narrower than the stated conclusion. The missing statistical details and the omitted CG-versus-SI comparison further weaken the current support for the headline recommendation.
major comments (3)
- [Section 4 and Section 5.2] The conclusion in Section 8 states that using a label for spatially grouped items helps users to identify, compare, and summarize data, but in the Compare and Summarize tasks all items within a subgroup share the same rating ('All items within a subgroup share the same rating'). Because the grouping rule in Section 4 groups items by 'same type and ratings, if available', the representative label always conveys the entire value distribution of its group; the study never tests the target scenario of spatially proximate same-type items with differing ratings, which is common in the motivating library/bookstore example. This is a construct-validity gap rather than only an external-validity limitation, and it should be addressed either by adding a condition with heterogeneous ratings within spatial clusters or by explicitly restricting the recommendation to groups with homogeneous displayed attributes.
- [Section 6] The statistical reporting is incomplete. The paper reports only 'significant differences among label placement methods (p < 0.05)' and the SG-SI pairwise comparison, without F statistics, degrees of freedom, exact p-values, or effect sizes (e.g., partial eta-squared) for the completion-time ANOVA, and without the test statistic or p-value for the GLMM accuracy model. In addition, the paper never reports the CG-SI pairwise comparison, so H1 (both grouped methods better than Situated Individual) is not actually evaluated; the discussion's claim of 'supporting the H1 hypothesis' is therefore too strong based on the reported results.
- [Section 5.2 and Figure 1] The information content of a representative label is unspecified. In the Compare task participants must report 'how many items exist at the [highest] rating', and in the Summarize task they must compute average ratings across subgroups; if the grouped label displays only type and rating but not item count, participants in the SG and CG conditions must obtain counts by inspecting leader lines, whereas SI participants can read counts directly from individual labels. The paper should state precisely what the representative label shows and, ideally, report participants' solution strategies or an additional analysis, so that the completion-time advantage of SG can be attributed to reduced clutter rather than to differences in the displayed information.
minor comments (6)
- [Abstract] The phrase 'the user' FOV' should be 'the user's FOV'.
- [Section 7] In the Limitations paragraph, 'within users' FOW' should be 'within users' FOV'.
- [References] References [29] and [30] are the same Vollick et al. paper; the duplicate entry should be removed or redirected to the original source.
- [Section 5.1] The sentence 'The ratings were displayed directly above their labels (e.g., )' contains an empty placeholder; include the rating glyph or a verbal description of how the rating was shown.
- [Section 6] The NASA-TLX results report means with +/- values but do not specify whether these are standard errors or confidence intervals; please clarify the error-bar convention.
- [Section 6] The preference counts (8, 4, 3) would benefit from a reported chi-squared statistic and p-value, although the text correctly notes the test was not significant.
Circularity Check
No significant circularity: the study's conclusions are empirical comparisons of label placements, not derived from fitted parameters or self-citation chains.
full rationale
The paper reports a controlled user study comparing three label placement methods (Situated Individual, Situated Grouped, Centered Grouped) across three tasks. The central claim, that grouped labels reduce completion time, is supported by direct measurements: 'Participants performed the tasks significantly faster (p < 0.05) with SG (14.0s ± 1.51s) than with SI (16.5s ± 2.61s)' (Section 6). There is no mathematical derivation chain, no fitted parameter that is later renamed as a prediction, and no quantity defined in terms of the outcome it is supposed to explain. The hypotheses are empirical and falsifiable; notably H2 (Centered Grouped faster than Situated Grouped) was explicitly rejected, which would be impossible if the results were forced by construction. The experimental conditions, including the Situated Individual baseline, are taken from independent prior work by Lin et al. [20] and Chen et al. [8], not from the present authors' own prior results. The paper's acknowledged limitations (VR simulation, simple objects, no head movement, manual distance thresholds; Section 7) are external-validity and construct-validity concerns, not circularity: they do not make the measured outcome equivalent to the input by definition. The skeptic's concern about within-group rating homogeneity is a valid threat to generalizing the recommendation to mixed-value groups, but it does not constitute circularity. No self-citation is load-bearing: the authors rely on established methods and external baselines, and conclusions stand or fall on the reported experimental data. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (2)
- DBSCAN Euclidean distance threshold (epsilon) =
not reported
- Force-directed layout parameters for Centered Grouped =
not reported
assumptions (3)
- domain assumption A VR simulation with a static FOV and simple objects adequately represents AR label placement contexts.
- domain assumption Identify, Compare, and Summarize are representative AR label tasks.
- domain assumption Spatial clustering by type and Euclidean distance reflects the groups users perceive.
Cite this review
Pith. "Pith review of Exploring AR Label Placements in Visually Cluttered Scenarios." pith.science (2026). https://pith.science/paper/4EGQ3QQN
@misc{pith2026250700198,
author = {Pith},
title = {Pith review of: Exploring AR Label Placements in Visually Cluttered Scenarios},
year = {2026},
howpublished = {\url{https://pith.science/paper/4EGQ3QQN}},
note = {Machine review of arXiv:2507.00198}
}
read the original abstract
We investigate methods for placing labels in AR environments that have visually cluttered scenes. As the number of items increases in a scene within the user' FOV, it is challenging to effectively place labels based on existing label placement guidelines. To address this issue, we implemented three label placement techniques for in-view objects for AR applications. We specifically target a scenario, where various items of different types are scattered within the user's field of view, and multiple items of the same type are situated close together. We evaluate three placement techniques for three target tasks. Our study shows that using a label to spatially group the same types of items is beneficial for identifying, comparing, and summarizing data.
Figures
Reference graph
Works this paper leans on
-
[1]
R. Azuma, Y . Baillot, R. Behringer, S. Feiner, S. Julier, and B. MacIn- tyre. Recent advances in augmented reality.IEEE Computer Graphics and Applications, 21(6):34–47, 2001. doi: 10.1109/38.963459 1
-
[2]
R. Azuma and C. Furmanski. Evaluating label placement for aug- mented reality view management. IEEE and ACM International Sym- posium on Mixed and Augmented Reality , pp. 66–75, 2003. doi: 10. 1109/ISMAR.2003.1240689 1, 2
arXiv 2003
-
[3]
M. R. Beck, M. C. Lohrenz, and J. G. Trafton. Measuring search efficiency in complex visual search tasks: global and local clutter. J Exp Psychol Appl, 16(3):238–50, 2010. doi: 10.1037/a0019633 4
-
[4]
M. A. Bekos, B. Niedermann, and M. N ¨ollenburg. External label- ing techniques: A taxonomy and survey. Computer Graphics Forum, 38(3):833–860, 2019. doi: 10.1111/cgf.13729 2, 3
-
[5]
B. Bell, S. Feiner, and T. H ¨ollerer. View management for virtual and augmented reality. ACM Symposium on User Interface Software and Technology, p. 101–110, 2001. doi: 10.1145/502348.502363 2
arXiv 2001
-
[6]
M. Brehmer and T. Munzner. A multi-level typology of abstract vi- sualization tasks. IEEE Transactions on Visualization and Computer Graphics, 19(12):2376–2385, 2013. doi: 10.1109/TVCG.2013.124 2, 3
-
[7]
N. Bressa, H. Korsgaard, A. Tabard, S. Houben, and J. Vermeulen. What’s the situation with situated visualization? a survey and per- spectives on situatedness. IEEE Transactions on Visualization and Computer Graphics, 28(1):107–117, 2021. doi: 10.1109/TVCG.2021 .3114835 1
-
[8]
Z. Chen, D. Chiappalupi, T. Lin, Y . Yang, J. Beyer, and H. Pfister. Rl- label: A deep reinforcement learning approach intended for ar label placement in dynamic scenarios. IEEE Transactions on Visualization and Computer Graphics, pp. 1–11, 2023. doi: 10.1109/TVCG.2023. 3326568 1, 2, 3, 4
Show all 35 references
-
[9]
ElSayed, B
N. ElSayed, B. Thomas, K. Marriott, J. Piantadosi, and R. Smith. Sit- uated analytics. IEEE Big Data Visual Analytics, pp. 1–8, 2015. doi: 10.1109/BDV A.2015.7314302 1
2015
-
[10]
Ester, H.-P
M. Ester, H.-P. Kriegel, J. Sander, and X. Xu. A density-based algo- rithm for discovering clusters in large spatial databases with noise. In International Conference on Knowledge Discovery and Data Mining, KDD’96, p. 226–231. AAAI Press, 1996. 3
1996
-
[11]
Fink, J.-H
M. Fink, J.-H. Haunert, A. Schulz, J. Spoerhase, and A. Wolff. Al- gorithms for labeling focus regions. IEEE Transactions on Visual- ization and Computer Graphics , 18(12):2583–2592, 2012. doi: 10. 1109/TVCG.2012.193 2
2012
-
[12]
Gebhardt, B
C. Gebhardt, B. Hecox, B. van Opheusden, D. Wigdor, J. Hillis, O. Hilliges, and H. Benko. Learning cooperative personalized poli- cies from gaze data. ACM Symposium on User Interface Software and Technology, p. 197–208, 2019. doi: 10.1145/3332165.3347933 2
2019
-
[13]
Grasset, T
R. Grasset, T. Langlotz, D. Kalkofen, M. Tatzgern, and D. Schmal- stieg. Image-driven view management for augmented reality browsers. IEEE Symposium on Mixed and Augmented Reality , pp. 177–186,
-
[14]
G ¨otzelmann, K
T. G ¨otzelmann, K. Hartmann, and T. Strothotte. Contextual grouping of labels. Simulation und Visualisierung, pp. 245–258, 2006. 2
2006
-
[15]
Hegde, J
S. Hegde, J. Maurya, R. Hebbalaguppe, and A. Kalkar. Smar- toverlays: A visual saliency driven label placement for intelligent human-computer interfaces. pp. 1110–1119, 2020. doi: 10.1109/ W ACV45572.2020.9093587 2
2020
-
[16]
Huynh, J
B. Huynh, J. Orlosky, and T. H ¨ollerer. In-situ labeling for augmented reality language learning. IEEE Conference on Virtual Reality and 3D User Interfaces , pp. 1606–1611, 2019. doi: 10.1109/VR.2019. 8798358 1
2019 doi
-
[17]
J. Jia, S. Elezovikj, H. Fan, et al. Semantic-aware label placement for augmented reality in street view. The Visual Computer, 37:1805– 1819, 2021. doi: 10.1007/s00371-020-01939-w 2
2021 doi
-
[18]
K ¨oppel, M
T. K ¨oppel, M. E. Gr ¨oller, and H. Y . Wu. Context-responsive labeling in augmented reality. IEEE Symposium on Pacific Visualization, pp. 91–100, 2021. doi: 10.1109/PacificVis52677.2021.00020 2
2021
-
[19]
B. Lee, M. Sedlmair, and D. Schmalstieg. Design patterns for situ- ated visualization in augmented reality. IEEE Transactions on Visu- alization and Computer Graphics , 30(1):1324–1335, 2024. doi: 10. 1109/TVCG.2023.3327398 2
2024
-
[20]
T. Lin, Y . Yang, J. Beyer, and H. Pfister. Labeling out-of-view ob- jects in immersive analytics to support situated visual searching.IEEE Transactions on Visualization and Computer Graphics , 29(3):1831– 1844, 2023. doi: 10.1109/TVCG.2021.3133511 1, 2, 3, 4
2023
-
[21]
J. B. Madsen, M. Tatzgern, C. B. Madsen, D. Schmalstieg, and D. Kalkofen. Temporal coherence strategies for augmented reality labeling. IEEE Transactions on Visualization and Computer Graph- ics, 22(4):1415–1423, 2016. doi: 10.1109/TVCG.2016.2518318 1, 2
2016
-
[22]
Marquardt, C
A. Marquardt, C. Trepkowski, T. D. Eibich, J. Maiero, E. Kruijff, and J. Sch¨oning. Comparing non-visual and visual guidance methods for narrow field of view augmented reality displays. IEEE Transactions on Visualization and Computer Graphics , 26(12):3389–3401, 2020. doi: 10....
2020
-
[23]
M ¨uhler and B
K. M ¨uhler and B. Preim. Automatic textual annotation for surgical planning. Proceedings of the Vision, Modeling, and Visualization Workshop, pp. 277–284, 2009. 2
2009
-
[24]
Rakholia, S
N. Rakholia, S. Hegde, and R. Hebbalaguppe. Where to place: A real-time visual saliency based label placement for augmented reality applications. IEEE International Conference on Image Processing , pp. 604–608, 2018. doi: 10.1109/ICIP.2018.8451052 2
2018
-
[25]
Tatzgern, D
M. Tatzgern, D. Kalkofen, R. Grasset, and D. Schmalstieg. Hedge- hog labeling: View management techniques for external labels in 3d space. IEEE Virtual Reality, pp. 27–32, 2014. doi: 10.1109/VR.2014 .6802046 2
2014 doi
-
[26]
Tatzgern, D
M. Tatzgern, D. Kalkofen, and D. Schmalstieg. Dynamic compact visualizations for augmented reality. IEEE Virtual Reality, pp. 3–6,
-
[27]
Tatzgern, V
M. Tatzgern, V . Orso, D. Kalkofen, G. Jacucci, L. Gamberini, and D. Schmalstieg. Adaptive information density for augmented reality displays. IEEE Virtual Reality, pp. 83–92, 2016. doi: 10.1109/VR. 2016.7504691 2
2016
-
[28]
Verghese and S
P. Verghese and S. P. McKee. Visual search in clutter. Vision Re- search, 44(12):1217–1225, 2004. Visual Attention. doi: 10.1016/j. visres.2003.12.006 2
2004 doi
-
[30]
V ollick, D
I. V ollick, D. V ogel, M. Agrawala, and A. Hertzmann. Specifying label layout style by example. ACM Symposium on User Interface Software and Technology, p. 221–230, 2007. doi: 10.1145/1294211. 1294252 4
2007 doi
-
[31]
Z. Wen, W. Zeng, L. Weng, Y . Liu, M. Xu, and W. Chen. Effects of view layout on situated analytics for multiple-view representations in immersive visualization. IEEE Transactions on Visualization and Computer Graphics, 29(1):440–450, 2023. doi: 10.1109/TVCG.2022 .3209475 1
2023 doi
-
[32]
Whitlock, S
M. Whitlock, S. Smart, and D. A. Szafir. Graphical perception for immersive analytics. IEEE Conference on Virtual Reality and 3D User Interfaces, pp. 616–625, 2020. doi: 10.1109/VR46266.2020.00084 2
2020
-
[33]
Z. Zhou, L. Wang, and V . Popescu. A partially-sorted concentric lay- out for efficient label localization in augmented reality. IEEE Trans- actions on Visualization and Computer Graphics, 27(11):4087–4096,
-
[2012]
doi: 10.1109/ISMAR.2012.6402555 2
2012
-
[2013]
doi: 10.1109/VR.2013.6549347 2
2013
-
[2021]
doi: 10.1109/TVCG.2021.3106492 1, 2 5
2021
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.