REVIEW 3 major objections 4 minor 135 references
What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision
T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper argues that visually distinguishing objects by importance level in a wearable AR display shifts people with low vision toward high-importance objects, at the cost of recalling fewer objects overall.
desk verdict The design exploration is genuinely useful, but the headline quantitative claim about importance-based attention shift doesn't survive contact with the paper's own thresholds or its AR baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
SceneGlance: a HoloLens-based wearable system that streams video to a server running fine-tuned RTMDet segmentation models (fine-tuned on kitchen and street datasets), raycasts detections onto the environmental mesh for 3D placement, and renders three base augmentations (static outline, solid overlay, icon label) combined into three distinction methods (by form, by color, by additional visual information). The load-bearing idea is the importance ranking itself, derived from a six-participant formative study that categorized objects as safety-related, visually challenging, or frequently used, and rated them by risk severity and visual difficulty into primary versus secondary importance; this
What would settle it
Re-run Study I with per-trial logging of which objects were actually augmented; restrict the analysis to trials where every primary-important object in the layout was detected and augmented. If the first-noticed rate no longer differs from the Reality baseline, the attention effect is an artifact of detection, not of importance-based distinction.
Extended reading notes
Core claim
On its own terms, the paper establishes that AR distinction—rendering objects of different importance with different visual treatment—is a workable strategy for guiding attention in complex scenes for people with low vision. SceneGlance detects and segments up to 11 kitchen and 21 street object classes, assigns them to primary- or secondary-importance, and augments them accordingly. In a controlled mock-kitchen study, participants first noticed primary-important objects in 79.2% of SceneGlance trials versus 20.8% with only their own vision (p < 0.001), and recall ratios for primary-important objects rose significantly, while overall object recall fell by about 8 points in both AR conditions.
Load-bearing premise
The claim that AR distinction shifted attention assumes the object detector's live performance was accurate enough that the effect comes from the importance-based rendering, not from which objects happened to get augmented; the reported offline false-negative rates (29.3% kitchen, 42.3% street) mean undetected objects—especially transparent glasses and curb cuts—could be driving both the attention shift and the recall drop.
Editorial extensions
If this is right
- Importance-differentiated AR augmentation reliably shifts first attention toward higher-importance objects for people with low vision in cluttered scenes.
- Multi-object augmentation imposes an attention-recall tradeoff: gains in high-importance recall come with reduced recall of un-augmented non-important objects.
- Adjacent augmentations of the same importance merge into misleading shapes, so distinction systems must also consider spatial relations, not only per-object importance.
- In outdoor scenes, continuous surfaces and dynamic objects need different treatments: outlines for boundaries, and importance based on collision likelihood rather than fixed object type.
- Color-based distinction degrades under variable outdoor lighting, while form-based distinction (overlay vs outline, outline thickness) remains recognizable.
Reading between the lines
- If the attention-recall tradeoff generalizes, importance-based AR acts like a spotlight that narrows the visual field; adaptive granularity (group outlines with counts) could recover breadth without losing the spotlight.
- The reported effect may partly reflect detection asymmetries: with false negatives highest for transparent objects and curb cuts, the attention shift could be an artifact of which objects were actually augmented, so per-trial augmentation logging is needed to separate design from detection.
- A testable extension: combining importance distinction with predicted trajectory (e.g., a car at a crosswalk rising to primary-importance as it approaches) should strengthen perceived safety benefits relative to static importance ranking.
- The merging of adjacent augmentations suggests a concrete design heuristic: vary color or texture across adjacent same-importance objects, or use outline thickness, to break uniform connectedness—something the paper implies but does not implement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents SceneGlance, a wearable AR system that detects important objects in complex scenes and visually distinguishes them by importance level (primary vs. secondary) using forms, colors, and/or additional information. The work is grounded in a formative study with six PLV that characterized important object categories and derived design guidelines. The system is evaluated in two studies: a controlled lab study with 12 PLV performing scene-perception tasks in a mock kitchen under three conditions (Reality baseline, AR baseline, SceneGlance), and a free-form outdoor think-aloud study with 13 PLV. The paper claims that AR distinction shifted attention toward primary-important objects, supported new perception strategies, and produced an attention-recall tradeoff, while also surfacing AR-specific challenges and design implications.
Significance. If the central claims were fully supported, this would be a valuable contribution to AR-based vision enhancement for low vision in complex scenes, an underexplored area. The paper's strengths include: a formative study that directly informs the system design; a fully implemented, near-real-time wearable system with a technical evaluation; a controlled three-condition lab study with PLV; and a complementary outdoor study that identifies realistic deployment challenges. The qualitative findings about augmentation interference, occlusion, and spatial misalignment are useful for future system design. However, the quantitative evidence does not currently support the abstract's strong claim that the importance-distinction design, rather than generic augmentation, drove the attention shift and recall tradeoff. The paper's own Bonferroni-corrected analyses show no significant difference between SceneGlance and the AR baseline on attention or recall measures, and two of the reported 'significant' pairwise comparisons exceed the stated threshold. These issues are fixable through revised analyses and reporting, but they are load-bearing for the paper's central message.
major comments (3)
- [§5.3.2, §5.3.3, Abstract] The abstract's claim that 'AR distinction on object importance shifted PLV's attention toward objects of higher importance' is not supported by the reported statistics. For primary-important recall ratio, SceneGlance does not differ from the AR baseline (p=0.927); for first-noticed objects, SceneGlance also does not differ from the AR baseline (p=0.149), and the AR baseline falls in between without significant separation from Reality (p=0.426). The only significant contrasts are versus the no-AR Reality baseline, which is consistent with a generic augmentation effect rather than a distinction-specific effect. The conclusion in Section 5.3.2 should be tempered, and the abstract should not attribute the effects to 'AR distinction' without significant SceneGlance-vs-AR-baseline evidence or additional analysis that isolates the distinction manipulation.
- [§5.3.3, §5.3.2] The reported pairwise p-values are inconsistent with the stated Bonferroni threshold (α=0.0056). In §5.3.3, overall recall is described as significantly lower under SceneGlance (diff=-0.079, p=0.013), but 0.013 > 0.0056, so this contrast is not significant by the paper's own criterion. Likewise, in §5.3.2 the primary-important recall ratio contrasts for SceneGlance (p=0.006) and the AR baseline (p=0.018) are labeled 'significantly higher,' yet both exceed 0.0056. This is an internal inconsistency that affects the abstract and the discussion of the attention-recall tradeoff. The authors should re-run or re-report the post-hoc comparisons with the stated correction and revise the claims accordingly.
- [§4.3.3, §5.3] The attention-shift and recall-reduction results relative to Reality could be confounded by differential detection coverage across importance levels. Offline false-negative rates are high and vary strongly by class (e.g., glasses 39.4%, jars 37.7%, curb cut 89.0%, crosswalk 79.1%), and no per-trial online augmentation coverage is reported. If undetected objects were systematically clustered in primary-important or secondary-important classes, the Reality-vs-SceneGlance differences could reflect which objects received augmentations rather than the importance-distinction design. The AR baseline partially controls for the detection pipeline, but because SceneGlance and the AR baseline are not significantly separated, the distinction-specific interpretation is not established. Please report per-condition augmentation coverage, and/or re-analyze attention and recall restricted to objects that
minor comments (4)
- [§4.3.2] The text says the models 'demonstrated strong robustness' and 'high accuracy,' but the reported mAP values (0.435 kitchen, 0.324 street) and false-negative rates (29.3% and 42.3%) suggest modest performance. Please soften this characterization to match the data.
- [§5.3.2] The phrase '(p=0.034 > 0.0056 with correction)' is awkward; the paper uses the threshold both as a significance level and as a post-hoc correction. Please clarify that no significance is claimed when p > 0.0056.
- [Table 2 and §5.3.2] The first-noticed analysis treats the two trials per condition as independent responses; this should be acknowledged or handled with a repeated-measures model. Also, the 22 vs 24 valid responses across conditions are not tested for the effect of missingness.
- [§5.1.3] In the AR baseline, participants chose their preferred single base augmentation, but the choice is not recorded as a factor. Differences in the chosen augmentation (outline vs overlay vs icon) could affect attention and recall; at least a sensitivity analysis or descriptive summary would help.
Circularity Check
No significant circularity: SceneGlance is an empirical design-and-evaluation study; importance labels are inputs, and attention/recall outcomes are measured independently against Reality and AR baselines.
full rationale
The paper makes no formal derivation or predictive claim that reduces to its inputs. The formative study (Section 3) elicited important objects and importance levels from six PLV; these labels were used as design inputs for SceneGlance (Section 4). The evaluation (Section 5) then measured attention allocation via recall ratios and first-noticed-object distribution under three conditions (Reality, AR baseline, SceneGlance). The outcome measures are behaviorally coded and statistically compared, not computed from the importance labels or from any fitted parameter. The AR baseline condition explicitly controls for generic augmentation, so the SceneGlance-vs-AR-baseline comparison is the right contrast for the distinction-specific claim; the fact that those comparisons were not significant (primary-important recall ratio p=0.927; first-noticed distribution p=0.149, Section 5.3.2) is a validity/interpretation concern, not a circularity. Self-citations (e.g., [123] for yellow color preference, [13,58,126] for augmentation designs) are used only to motivate design choices and describe related work; no load-bearing argument reduces to a self-cited theorem. The limitations stated in Section 7.2 (small sample, carryover, non-controlled Study II) are acknowledged threats to generalizability, not evidence of circular derivation. Accordingly, no circular step can be quoted with a specific reduction; the correct finding is no significant circularity.
Assumptions & free parameters
free parameters (4)
- Importance category assignment (primary vs secondary) =
Not disclosed
- Detection confidence threshold =
0.45
- IoU threshold for detection-level metrics =
0.3
- Scene complexity parameters (mock kitchen) =
30-33 objects; 12 categories; 150cm x 60cm counter
assumptions (5)
- domain assumption Participants are representative of the broader PLV population
- domain assumption The mock kitchen and outdoor route approximate real-world complexity
- ad hoc to paper Object importance can be treated as a stable two-level property
- domain assumption Recognition errors did not systematically bias attention measures
- standard math Standard statistical assumptions (normality, independence, chi-square approximation)
Cite this review
Pith. "Pith review of What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision." pith.science (2026). https://pith.science/paper/GVZV2IQN
@misc{pith2026260710902,
author = {Pith},
title = {Pith review of: What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision},
year = {2026},
howpublished = {\url{https://pith.science/paper/GVZV2IQN}},
note = {Machine review of arXiv:2607.10902}
}
read the original abstract
People with low vision (PLV) struggle to perceive complex scenes like busy kitchens and crowded streets, which contain many objects, visual clutter, and dynamic elements. Prior AR systems for low vision either enhance low-level visual features or augment task-relevant objects for single tasks in simple settings, leaving multi-object augmentation in complex scenes underexplored. Informed by a formative study characterizing important objects and their perceived importance for PLV, we built SceneGlance, a wearable AR system that recognizes important objects and visually distinguishes them by importance level. Through a controlled lab study with 12 PLV in a mock-up kitchen scene and a free-form think-aloud study with 13 PLV navigating an outdoor route, we found that AR distinction on object importance shifted PLV's attention toward objects of higher importance, and supported perception strategies such as building mental snapshots from the augmentation distribution and hierarchical scanning by importance. However, this attention shift came with a tradeoff of reduced overall scene recall. The studies also surfaced challenges posed by AR augmentations in complex scenes, such as adjacent augmentations blending or interfering with each other, yielding design implications for more practical AR vision enhancement systems in the complex real world.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Walid I Al-Atabany, Muhammad A Memon, Susan M Downes, and Patrick A Degenaar. 2010. Designing and testing scene enhancement algorithms for patients with retina degenerative disorders.Biomedical engineering online9, 1 (2010), 27. ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal Wu et al
2010
-
[2]
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. 2016. Social lstm: Human trajectory prediction in crowded spaces. InProceedings of the IEEE conference on computer vision and pattern recognition. 961–971
2016
-
[3]
George A Alvarez and Patrick Cavanagh. 2004. The capacity of visual short- term memory is set both by visual information load and by number of objects. Psychological science15, 2 (2004), 106–111
2004
-
[4]
American Optometric Association. 2023. Legal blindness in America — aoa.org. https://www.aoa.org/news/clinical-eye-care/diseases-and-conditions/ legal-blindness-in-america?sso=y. [Accessed 12-03-2025]
2023
-
[5]
Anastasios Nikolas Angelopoulos, Hossein Ameri, Debbie Mitra, and Mark Humayun. 2019. Enhanced depth navigation through augmented reality depth mapping in patients with low vision.Scientific reports9, 1 (2019), 11230
2019
-
[6]
Richard A Armstrong. 2014. When to use the Bonferroni correction.Ophthalmic and physiological optics34, 5 (2014), 502–508
2014
-
[7]
Marie Claire Bilyk, Jessica M Sontrop, Gwen E Chapman, Susan I Barr, and Linda Mamer. 2009. Food experiences and eating patterns of visually impaired and blind people.Canadian Journal of Dietetic practice and research70, 1 (2009), 13–18
2009
-
[8]
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology3, 2 (2006), 77–101
2006
Show all 135 references
-
[9]
Virginia Braun and Victoria Clarke. 2024. Thematic analysis. InEncyclopedia of quality of life and well-being research. Springer, 7187–7193
2024
-
[10]
Xiaoqian J Chai, Noa Ofen, Lucia F Jacobs, and John DE Gabrieli. 2010. Scene complexity: influence on perception, memory, and development in the medial temporal lobe.Frontiers in human neuroscience4 (2010), 1021
2010
-
[11]
Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. Worldscribe: Towards context-aware live visual descriptions. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–18
2024
-
[12]
Xiaojun Chang, Pengzhen Ren, Pengfei Xu, Zhihui Li, Xiaojiang Chen, and Alex Hauptmann. 2021. A comprehensive survey of scene graphs: Generation and application.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 1 (2021), 1–26
2021
-
[13]
Ruijia Chen, Junru Jiang, Pragati Maheshwary, Brianna R Cochran, and Yuhang Zhao. 2025. Visimark: Characterizing and augmenting landmarks for people with low vision in augmented reality to support indoor navigation. InProceedings of the 2025 CHI Conference on Human Factors in ...
2025
-
[14]
Ruijia Chen, Yuheng Wu, Charlie Houseago, Filipe Gaspar, Filippo Ale- otti, Dorian Gálvez-López, Oliver Johnston, Diego Mazala, Guillermo Garcia- Hernando, Maryam Bandukda, Gabriel Brostow, and Jessica Van Brummelen
-
[15]
2013.Statistical power analysis for the behavioral sciences
Jacob Cohen. 2013.Statistical power analysis for the behavioral sciences. Rout- ledge
2013
-
[16]
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling Egocentric Vision: The EPIC- KITCHENS Dataset. InEuropean Conference on Computer...
2018
-
[17]
Raghavendra Singh Dasila, Meet Trivedi, Shubham Soni, M Senthil, and M Narendran. 2017. Real time environment perception for visually impaired. In 2017 IEEE Technological Innovations in ICT for Agriculture and Rural Development (TIAR). IEEE, 168–172
2017
-
[18]
Johanne Desrosiers, Marie-Chantal Wanet-Defalque, Khatoune Témisjian, Jacques Gresset, Marie-France Dubois, Judith Renaud, Claude Vincent, Jacque- line Rousseau, Mathieu Carignan, and Olga Overbury. 2009. Participation in daily activities and social roles of older adults with ...
2009
-
[19]
Juan C Dibene and Enrique Dunn. 2022. HoloLens 2 Sensor Streaming.arXiv preprint arXiv:2211.02648(2022)
2022 arXiv
-
[20]
Olive Jean Dunn. 1961. Multiple comparisons among means.Journal of the American statistical association56, 293 (1961), 52–64
1961
-
[21]
Martin Eckert, Matthias Blex, Christoph M Friedrich, et al. 2018. Object detection featuring 3D audio localization for Microsoft HoloLens. InProc. 11th Int. Joint Conf. on Biomedical Engineering Systems and Technologies, Vol. 5. 555–561
2018
-
[22]
Elkin, Matthew Kay, James J
Lisa A. Elkin, Matthew Kay, James J. Higgins, and Jacob O. Wobbrock. 2021. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. InThe 34th Annual ACM Symposium on User Interface Software and Technology(Virtual Event, USA)(UIST ’21). Association for Computing Mac...
2021
-
[23]
Niklas Elmqvist and Philippas Tsigas. 2008. A taxonomy of 3d occlusion manage- ment for visualization.IEEE transactions on visualization and computer graphics 14, 5 (2008), 1095–1109
2008
-
[24]
MR Everingham, BT Thomas, T Troscianko, et al. 1999. Head-mounted mobility aid for low vision using scene classification techniques.The International Journal of Virtual Reality3, 4 (1999), 3
1999
-
[25]
Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision88 (2010), 303–338
2010
-
[26]
Fox, Ahmad Ahmadzada, Clara T
Dylan R. Fox, Ahmad Ahmadzada, Clara T. Friedman, Shiri Azenkot, Marlena A. Chu, Roberto Manduchi, and Emily A. Cooper. 2023. Using augmented reality to cue obstacles for people with low vision.Opt. Express31, 4 (Feb 2023), 6827–6848. doi:10.1364/OE.479258
2023 doi
-
[27]
Bhanuka Gamage, Nicola McDowell, Dijana Kovacic, Leona Holloway, Thanh- Toan Do, Arthur James Lowery, Nicholas Price, and Kim Marriott. 2025. Smart Glasses for CVI: Co-Designing Extended Reality Solutions to Support Environ- mental Perception by People with Cerebral Visual Imp...
2025
-
[28]
Yun Gao, Dan Wu, Jie Song, Xueyi Zhang, Bangbang Hou, Hengfa Liu, Junqi Liao, and Liang Zhou. 2025. A wearable obstacle avoidance device for visually impaired individuals with cross-modal learning.Nature Communications16, 1 (2025), 2857
2025
-
[29]
Google. 2025. Protocol Buffers. https://developers.google.com/protocol-buffers. Accessed: 2025-03-30
2025
-
[30]
Tenen- baum, Antonio Torralba, Florian Shkurti, and Liam Paull
Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallab- hula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty El- lis, Rama Chellappa, Chuang Gan, Celso Miguel de Melo, Joshua B. Tenen- baum, Antonio Torralba, Florian Shkurti, and Liam Paull. 2...
2023 arXiv
-
[31]
Fangli Guan, Zhixiang Fang, Lubin Wang, Xucai Zhang, Haoyu Zhong, and Haosheng Huang. 2022. Modelling people’s perceived scene complexity of real-world environments using street-view panoramas and open geodata.ISPRS Journal of Photogrammetry and Remote Sensing186 (2022), 315–331
2022
-
[32]
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi
-
[33]
Yu Hao, Alexey Magay, Hao Huang, Shuaihang Yuan, Congcong Wen, and Yi Fang. 2024. ChatMap: A Wearable Platform Based on the Multi-modal Foun- dation Model to Augment Spatial Cognition for People with Blindness and Low Vision. In2024 IEEE/RSJ International Conference on Intelli...
2024
-
[34]
Yu Hao, Fan Yang, Hao Huang, Shuaihang Yuan, Sundeep Rangan, John-Ross Rizzo, Yao Wang, and Yi Fang. 2024. A multi-modal foundation model to assist people with blindness and low vision in environmental interaction.Journal of Imaging10, 5 (2024), 103
2024
-
[35]
Sharon A Haymes, Alan W Johnston, and Anthony D Heyes. 2002. Relationship between vision impairment and ability to perform activities of daily living. Ophthalmic and Physiological Optics22, 2 (2002), 79–91
2002
-
[36]
John M Henderson, Myriam Chanceaux, and Tim J Smith. 2009. The influence of clutter on real-world scene search: Evidence from search efficiency and eye movements.Journal of vision9, 1 (2009), 32–32
2009
-
[37]
John M Henderson, Taylor R Hayes, Candace E Peacock, and Gwendolyn Rehrig
-
[38]
Marion Hersh. 2022. Wearable travel aids for blind and partially sighted people: A review with a focus on design issues.Sensors22, 14 (2022), 5454
2022
-
[39]
Stephen L Hicks, Iain Wilson, Louwai Muhammed, John Worsfold, Susan M Downes, and Christopher Kennard. 2013. A depth-based head-mounted visual display to aid navigation in partially sighted individuals.PLOS ONE8, 7 (2013), e67695
2013
-
[40]
Jonathan Huang, Max Kinateder, Matt J Dunn, Wojciech Jarosz, Xing-Dong Yang, and Emily A Cooper. 2019. An augmented reality sign-reading assistant for users with reduced vision.PloS one14, 1 (2019), e0210630
2019
-
[41]
Alex D Hwang and Eli Peli. 2014. An augmented-reality edge enhancement application for Google Glass.Optometry and vision science91, 8 (2014), 1021– 1030
2014
-
[42]
Md Touhidul Islam, Imran Kabir, Elena Ariel Pearce, Md Alimoor Reza, and Syed Masum Billah. 2024. Identifying Crucial Objects in Blind and Low-Vision Individuals’ Navigation. InProceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–8
2024
-
[43]
Watthanasak Jeamwatthanachai, Mike Wald, and Gary Wills. 2019. Indoor navigation by blind people: Behaviors and challenges in unfamiliar spaces and buildings.British Journal of Visual Impairment37, 2 (2019), 140–153. doi:10. 1177/0264619619833723
2019
-
[44]
Nabila Jones, Hannah Elizabeth Bartlett, and Richard Cooke. 2019. An analysis of the impact of visual impairment on activities of daily living and vision-related quality of life in a visually impaired adult population.British Journal of Visual Impairment37, 1 (2019), 50–63
2019
-
[45]
Amy A Kalia, Gordon E Legge, and Nicholas A Giudice. 2008. Learning building layouts with non-geometric visual information: The effects of visual impairment and age.Perception37, 11 (2008), 1677–1699. SceneGlance ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal
2008
-
[46]
Avyay Ravi Kashyap. 2020. Behaviors, Problems and Strategies of Visually Impaired Persons During Meal Preparation in the Indian Context: Challenges and Opportunities for Design. InProceedings of the 22nd International ACM SIGACCESS Conference on Computers and Accessibility (AS...
2020
-
[47]
Liam Kettle and Yi-Ching Lee. 2022. Augmented reality for vehicle-driver communication: a systematic review.Safety8, 4 (2022), 84
2022
-
[48]
Doaa Khattab, Julie Buelow, and Donna Saccuteli. 2015. Understanding the barriers: Grocery stores and visually impaired shoppers.Journal of accessibility and design for all: JACCES5, 2 (2015), 157–173
2015
-
[49]
Daniel Killough, Justin Feng, Zheng Xue Ching, Daniel Wang, Rithvik Dyava, Yapeng Tian, and Yuhang Zhao. 2025. VRSight: An AI-Driven Scene Description System to Improve Virtual Reality Accessibility for Blind People. InProceedings of the 38th Annual ACM Symposium on User Inter...
2025
-
[50]
Benjamin Kommey, Kumbong Herrman, and Ernest Ofosu Addo. 2019. A smart vision based navigation aid for the visually impaired.Asian Journal of Research in Computer Science4, 3 (2019), 1–8
2019
-
[51]
Masaki Kuribayashi, Kohei Uehara, Allan Wang, Shigeo Morishima, and Chieko Asakawa. 2025. Wanderguide: Indoor map-less robotic guide for exploration by blind people. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–21
2025
-
[52]
Alexandra Kuznetsova, Per B Brockhoff, and Rune HB Christensen. 2017. lmerTest package: tests in linear mixed effects models.Journal of statistical software82 (2017), 1–26
2017
-
[53]
MiYoung Kwon, Chaithanya Ramachandra, PremNandhini Satgunam, Bartlett W Mel, Eli Peli, and Bosco S Tjan. 2012. Contour enhancement benefits older adults with simulated central field loss.Optometry and vision science89, 9 (2012), 1374– 1384
2012
-
[54]
Cameron Kyle-Davidson, Elizabeth Yue Zhou, Dirk B Walther, Adrian G Bors, and Karla K Evans. 2023. Characterising and dissecting human perception of scene complexity.Cognition231 (2023), 105319
2023
-
[55]
Mikko Kytö, Barrett Ens, Thammathip Piumsomboon, Gun A Lee, and Mark Billinghurst. 2018. Pinpointing: Precise head-and eye-based target selection for augmented reality. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–14
2018
-
[56]
Florian Lang and Tonja Machulla. 2021. Pressing a button you cannot see: evaluating visual designs to assist persons with low vision through augmented reality. InProceedings of the 27th ACM Symposium on Virtual Reality Software and Technology. 1–10
2021
-
[57]
Jaewook Lee, Yang Li, Dylan Bunarto, Eujean Lee, Olivia H Wang, Adrian Rodriguez, Yuhang Zhao, Yapeng Tian, and Jon E Froehlich. 2024. Towards ai-powered ar for enhancing sports playability for people with low vision: An exploration of arsports. In2024 IEEE International Sympo...
2024
-
[58]
Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. InProceedings of the 37th Annual ...
2024
-
[59]
Meesung Lee, Heerim Lee, Sungjoo Hwang, and Minji Choi. 2021. Under- standing the impact of the walking environment on pedestrian perception and comprehension of the situation.Journal of Transport & Health23 (2021), 101267
2021
-
[60]
Franklin Mingzhe Li, Jamie Dorst, Peter Cederberg, and Patrick Carrington. 2021. Non-Visual Cooking: Exploring Practices and Challenges of Meal Preparation by People with Visual Impairments. InProceedings of the 23rd International ACM SIGACCESS Conference on Computers and Acce...
2021
-
[61]
Ke Li and Jitendra Malik. 2016. Amodal instance segmentation. InEuropean Conference on Computer Vision. Springer, 677–693
2016
-
[62]
Zhipeng Li, Christoph Gebhardt, Yves Inglin, Nicolas Steck, Paul Streli, and Christian Holz. 2024. Situationadapt: Contextual ui optimization in mixed reality with situation awareness via llm reasoning. InProceedings of the 37th Annual ACM Symposium on User Interface Software ...
2024
-
[63]
Haozhe Lin, Jiangtao Gong, Yu Wang, Jinsong Zhang, Bing Bai, Yan Zhang, Luyao Wang, Chenyu Wei, Yancheng Cao, Kun Li, et al . 2025. AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments.npj Artificial Intelligence1,...
2025
-
[64]
Lawrence Zitnick, and Piotr Dollár
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár
-
[65]
David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-aware online adaptation of mixed reality interfaces. InProceedings of the 32nd annual ACM symposium on user interface software and technology. 147–160
2019
-
[66]
Alice Lo Valvo, Daniele Croce, Domenico Garlisi, Fabrizio Giuliano, Laura Giarré, and Ilenia Tinnirello. 2021. A navigation and augmented reality system for visually impaired people.Sensors21, 9 (2021), 3061
2021
-
[67]
Gang Luo and Eli Peli. 2006. Use of an augmented-vision device for visual search by patients with tunnel vision.Investigative ophthalmology & visual science47, 9 (2006), 4152–4159
2006
-
[68]
Chengqi Lyu, Wenwei Zhang, Haian Huang, Yue Zhou, Yudong Wang, Yanyi Liu, Shilong Zhang, and Kai Chen. 2022. Rtmdet: An empirical study of designing real-time object detectors.arXiv preprint arXiv:2212.07784(2022)
2022 arXiv
-
[69]
Márcio CF Macedo and Antonio L Apolinario. 2021. Occlusion handling in augmented reality: Past, present and future.IEEE Transactions on Visualization and Computer Graphics29, 2 (2021), 1590–1609
2021
-
[70]
Sean P MacEvoy and Russell A Epstein. 2011. Constructing scenes from objects in human occipitotemporal cortex.Nature neuroscience14, 10 (2011), 1323–1329
2011
-
[71]
Alexey Magay, Dhurba Tripathi, Yu Hao, and Yi Fang. 2024. A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision. InEuropean Conference on Computer Vision. Springer, 323–339
2024
-
[72]
Florian Mathis and Johannes Schöning. 2025. LifeInsight: Design and Evaluation of an AI-Powered Assistive Wearable for Blind and Low Vision People Across Multiple Everyday Life Scenarios. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–25
2025
-
[73]
2017.Designing experiments and analyzing data: A model comparison perspective
Scott E Maxwell, Harold D Delaney, and Ken Kelley. 2017.Designing experiments and analyzing data: A model comparison perspective. Routledge
2017
-
[74]
Microsoft. 2017. Seeing AI - Talking Camera for the Blind — seeingai.com. https://www.seeingai.com/. [Accessed 09-09-2025]
2017
-
[75]
Margrain, Yu-Kun Lai, and Parisa Eslambolchilar
Hein Min Htike, Tom H. Margrain, Yu-Kun Lai, and Parisa Eslambolchilar. 2021. Augmented reality glasses as an orientation and mobility aid for people with low vision: a feasibility study of experiences and requirements. InProceedings of the 2021 CHI Conference on Human Factors...
2021
-
[76]
Aliaksei Miniukovich and Antonella De Angeli. 2014. Quantification of interface visual complexity. InProceedings of the 2014 international working conference on advanced visual interfaces. 153–160
2014
-
[77]
Pilar Montero. 2005. Nutritional assessment and diet quality of visually impaired Spanish children.Annals of Human Biology32, 4 (2005), 498–512
2005
-
[78]
Karin Müller, Christin Engel, Claudia Loitsch, Rainer Stiefelhagen, and Gerhard Weber. 2022. Traveling more independently: a study on the diverse needs and challenges of people with visual or mobility impairments in unfamiliar indoor environments.ACM Transactions on Accessible...
2022
-
[79]
National Eye Institute. 2025. Low Vision | National Eye Institute — nei.nih.gov. https://www.nei.nih.gov/learn-about-eye-health/eye-conditions- and-diseases/low-vision. [Accessed 19-06-2024]
2025
-
[80]
Mark B Neider and Gregory J Zelinsky. 2011. Cutting through the clutter: Searching for targets in evolving complex scenes.Journal of Vision11, 14 (2011), 7–7
2011
-
[81]
Gerhard Neuhold, Tobias Ollmann, Samuel Rota Bulo, and Peter Kontschieder
-
[82]
Aude Oliva, Michael L Mack, Mochan Shrestha, and Angela Peeper. 2004. Iden- tifying the perceptual dimensions of visual complexity of scenes. InProceedings of the annual meeting of the cognitive science society, Vol. 26
2004
-
[83]
Rafael Padilla, Wesley L Passos, Thadeu LB Dias, Sergio L Netto, and Eduardo AB Da Silva. 2021. A comparative analysis of object detection metrics with a companion open-source toolkit.Electronics10, 3 (2021), 279
2021
-
[84]
Stephen Palmer and Irvin Rock. 1994. Rethinking perceptual organization: The role of uniform connectedness.Psychonomic bulletin & review1, 1 (1994), 29–55
1994
-
[85]
Ji Hwan Park, Braden Roper, Amirhossein Arezoumand, and Tien Tran. 2025. Exploring AR Label Placements in Visually Cluttered Scenarios. In2025 IEEE Visualization and Visual Analytics (VIS). IEEE, 336–340
2025
-
[86]
Karl Pearson. 1900. X. On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling.The London, Edinburgh, and Dublin Philosophical Magazine a...
1900
-
[87]
Denis G Pelli and Katharine A Tillman. 2008. The uncrowded window of object recognition.Nature neuroscience11, 10 (2008), 1129–1135
2008
-
[88]
Gonzalez Penuela, Jazmin Collins, Cynthia L
Ricardo E. Gonzalez Penuela, Jazmin Collins, Cynthia L. Bennett, and Shiri Azenkot. 2024. Investigating Use Cases of AI-Powered Scene Description Appli- cations for Blind and Low Vision People.Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems(2024). ...
2024
-
[89]
Dwight J Peterson and Marian E Berryhill. 2013. The Gestalt principle of similarity benefits visual working memory.Psychonomic bulletin & review20, 6 (2013), 1282–1289
2013
-
[90]
Sanjeev U Rao, Swaroop Ranganath, TS Ashwin, Guddeti Ram Mohana Reddy, et al. 2021. A Google glass based real-time scene analysis for the visually impaired.IEEE Access9 (2021), 166351–166369. ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal Wu et al
2021
-
[91]
Elena T Remillard, Lyndsie M Koon, Tracy L Mitzner, and Wendy A Rogers. 2024. Everyday Challenges for Individuals Aging with Vision Impairment: Technology Implications.The Gerontologist64, 6 (2024), gnad169
2024
-
[92]
Ruth Rosenholtz, Yuanzhen Li, and Lisa Nakano. 2007. Measuring visual clutter. Journal of vision7, 2 (2007), 17–17
2007
-
[93]
Arezoo Sadeghzadeh, Md Baharul Islam, Md Nur Uddin, and Tarkan Aydin
-
[94]
Ahmed M Sayed, Mohamed Abou Shousha, MD Baharul Islam, Taher K Eleiwa, Rashed Kashem, Mostafa Abdel-Mottaleb, Eyup Ozcan, Mohamed Tolba, Jane C Cook, and Richard K Parrish. 2020. Mobility improvement of patients with peripheral visual field losses using novel see-through digit...
2020
-
[95]
Wilk, et al
S Shapiro, M.B. Wilk, et al . 1965. An analysis of variance test for normality. Biometrika52, 3 (1965), 591–611
1965
-
[96]
Sandra D Starke, Eugenie Golubova, Michael D Crossland, and James S Wolff- sohn. 2020. Everyday visual demands of people with low vision: A mixed methods real-life recording study.Journal of Vision20, 9 (2020), 3–3
2020
-
[97]
Froehlich
Lee Stearns, Leah Findlater, and Jon E. Froehlich. 2018. Design of an Augmented Reality Magnification Aid for Low Vision Users. InProceedings of the 20th Inter- national ACM SIGACCESS Conference on Computers and Accessibility(Galway, Ireland)(ASSETS ’18). Association for Compu...
2018
-
[98]
Sarit Szpiro, Yuhang Zhao, and Shiri Azenkot. 2016. Finding a store, searching for a product: a study of daily challenges of low vision people. InProceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing(Heidelberg, Germany)(UbiComp ’16)....
2016
-
[99]
Sarit Felicia Anais Szpiro, Shafeka Hashash, Yuhang Zhao, and Shiri Azenkot
-
[100]
Duje Tadin, Jeffrey B Nyquist, Kelly E Lusk, Anne L Corn, and Joseph S Lappin
-
[101]
Markus Tatzgern, Valeria Orso, Denis Kalkofen, Giulio Jacucci, Luciano Gam- berini, and Dieter Schmalstieg. 2016. Adaptive information density for aug- mented reality displays. In2016 IEEE Virtual Reality (VR). IEEE, 83–92
2016
-
[102]
Jodi Teitelman and Al Copolillo. 2005. Psychosocial issues in older adults’ adjustment to vision loss: findings from qualitative interviews and focus groups. The American journal of occupational therapy59, 4 (2005), 409–417
2005
-
[103]
Miguel Thibaut, Muriel Boucart, and Thi Ha Chau Tran. 2020. Object search in neovascular age-related macular degeneration: the crowding effect.Clinical and Experimental Optometry103, 5 (2020), 648–655
2020
-
[104]
John W Tukey. 1949. Comparing individual means in the analysis of variance. Biometrics(1949), 99–114
1949
-
[105]
Sandra Tullio-Pow, Hong Yu, and Megan Strickfaden. 2021. Do You See What I See? The shopping experiences of people with visual impairment.Interdisci- plinary Journal of Signage and Wayfinding5, 1 (2021), 42–61
2021
-
[106]
Kathleen A Turano, Duane R Geruschat, and Julie W Stahl. 1998. Mental effort required for walking: effects of retinitis pigmentosa.Optometry and Vision Science75, 12 (1998), 879–886
1998
-
[107]
Cashman, Bugra Tekin, Johannes L
Dorin Ungureanu, Federica Bogo, Silvano Galliani, Pooja Sama, Xin Duan, Casey Meekhof, Jan Stühmer, Thomas J. Cashman, Bugra Tekin, Johannes L. Schönberger, Pawel Olszta, and Marc Pollefeys. 2020. HoloLens 2 Research Mode as a Tool for Computer Vision Research. arXiv:2008.1123...
2020 arXiv
-
[108]
Unity Technologies. 2022. Unity - Manual: Raw Image — docs.unity3d.com. https: //docs.unity3d.com/2022.3/Documentation/Manual/script-RawImage.html. [Ac- cessed 07-04-2025]
2022
-
[109]
Joram J van Rheede, Iain R Wilson, Rose I Qian, Susan M Downes, Christopher Kennard, and Stephen L Hicks. 2015. Improving mobility performance in low vision with a distance-based representation of the visual scene.Investigative ophthalmology & visual science56, 8 (2015), 4802–4809
2015
-
[110]
Melissa Le-Hoa Võ. 2021. The meaning and structure of scenes.Vision Research 181 (2021), 10–20
2021
-
[111]
Melissa Le-Hoa Võ, Sage EP Boettcher, and Dejan Draschkow. 2019. Reading scenes: How scene grammar guides attention and aids perception in real-world environments.Current opinion in psychology29 (2019), 205–210
2019
-
[112]
Julian M Wallace, Susana TL Chung, and Bosco S Tjan. 2017. Object crowding in age-related macular degeneration.Journal of Vision17, 1 (2017), 33–33
2017
-
[113]
Ru Wang, Ruijia Chen, Anqiao Erica Cai, Zhiyuan Li, Sanbrita Mondal, and Yuhang Zhao. 2025. Characterizing Visual Intents for People with Low Vision through Eye Tracking. InProceedings of the 27th International ACM SIGACCESS Conference on Computers and Accessibility. 1–18
2025
-
[114]
Ru Wang, Zach Potter, Yun Ho, Daniel Killough, Linxiu Zeng, Sanbrita Mondal, and Yuhang Zhao. 2024. GazePrompt: Enhancing Low Vision People’s Reading Experience with Gaze-Aware Augmentations. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–17
2024
-
[115]
Ru Wang, Nihan Zhou, Tam Nguyen, Sanbrita Mondal, Bilge Mutlu, and Yuhang Zhao. 2023. Characterizing barriers and technology needs in the kitchen for blind and low vision people.arXiv preprint arXiv:2310.05396(2023)
2023 arXiv
-
[116]
David Whitney and Dennis M Levi. 2011. Visual crowding: A fundamental limit on conscious perception and object recognition.Trends in cognitive sciences15, 4 (2011), 160–168
2011
-
[117]
Sandro L Wiesmann and Melissa Le-Hoa Võ. 2023. Disentangling diagnostic object properties for human scene categorization.Scientific reports13, 1 (2023), 5912
2023
-
[118]
Pray before you step out
Michele A Williams, Amy Hurst, and Shaun K Kane. 2013. “Pray before you step out”: Describing Personal and Situational Blind Navigation Behaviors. In Proceedings of the 15th international ACM SIGACCESS conference on computers and accessibility. 1–8
2013
-
[119]
Wobbrock, Leah Findlater, Darren Gergle, and James J
Jacob O. Wobbrock, Leah Findlater, Darren Gergle, and James J. Higgins. 2011. The aligned rank transform for nonparametric factorial analyses using only anova procedures. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(Vancouver, BC, Canada)(CHI ’1...
2011
-
[120]
Jeremy M Wolfe. 2021. Guided Search 6.0: An updated model of visual search. Psychonomic bulletin & review28, 4 (2021), 1060–1092
2021
-
[121]
Limin Zeng. 2015. A survey: outdoor mobility experiences by the visually impaired. InMensch und Computer 2015–Workshopband. De Gruyter Oldenbourg, 391–397
2015
-
[122]
Xi Zhao, Kentaro Go, Kenji Kashiwagi, Masahiro Toyoura, Xiaoyang Mao, and Issei Fujishiro. 2019. Computational alleviation of homonymous visual field defect with ost-hmd: The effect of size and position of overlaid overview window. In2019 International Conference on Cyberworld...
2019
-
[123]
Yuhang Zhao, Michele Hu, Shafeka Hashash, and Shiri Azenkot. 2017. Under- standing Low Vision People’s Visual Perception on Commercial Augmented Reality Glasses. InProceedings of the 2017 CHI Conference on Human Factors in Computing Systems(Denver, Colorado, USA)(CHI ’17). Ass...
2017
-
[124]
Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR Visualizations to Facilitate Stair Navigation for People with Low Vision. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology(N...
2019
-
[125]
Yuhang Zhao, Sarit Szpiro, and Shiri Azenkot. 2015. Foresee: A customizable head-mounted vision enhancement system for people with low vision. InPro- ceedings of the 17th international ACM SIGACCESS conference on computers & accessibility. 239–249
2015
-
[126]
Yuhang Zhao, Sarit Szpiro, Jonathan Knighten, and Shiri Azenkot. 2016. CueSee: exploring visual cues for people with low vision to facilitate a visual search task. InProceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing. 73–84
2016
-
[127]
Yuhang Zhao, Sarit Szpiro, Lei Shi, and Shiri Azenkot. 2019. Designing and evaluating a customizable head-mounted vision enhancement system for people with low vision.ACM Transactions on Accessible Computing (TACCESS)12, 4 (2019), 1–46. SceneGlance ASSETS ’26, October 25–28, 2...
2019
-
[2012]
Peripheral vision of youths with low vision: motion perception, crowding, and visual search.Investigative ophthalmology & visual science53, 9 (2012), 5860–5868
2012
-
[2015]
arXiv:1405.0312 [cs.CV] https://arxiv.org/abs/1405.0312
Microsoft COCO: Common Objects in Context. arXiv:1405.0312 [cs.CV] https://arxiv.org/abs/1405.0312
-
[2016]
InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility
How people with low vision access computing devices: Understanding challenges and opportunities. InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility. 171–180
-
[2017]
InPro- ceedings of the IEEE international conference on computer vision
The mapillary vistas dataset for understanding of street scenes. InPro- ceedings of the IEEE international conference on computer vision. 4990–4999
-
[2018]
InProceedings of the IEEE conference on computer vision and pattern recognition
Social gan: Socially acceptable trajectories with generative adversarial networks. InProceedings of the IEEE conference on computer vision and pattern recognition. 2255–2264
-
[2019]
Meaning and attentional guidance in scenes: A review of the meaning map approach.Vision3, 2 (2019), 19
2019
-
[2024]
doi:10.1109/ACCESS.2024.3462628
ARVA: An Augmented Reality-Based Visual Aid for Mobility Enhancement Through Real-Time Video Stream Transformation.IEEE Access12 (2024), 137268– 137283. doi:10.1109/ACCESS.2024.3462628
2024
-
[2026]
InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26)
NaviNote: Enabling In-situ Spatial Annotation Authoring to Support Exploration and Navigation for Blind and Low Vision People. InProceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). As- sociation for Computing Machinery, New York, NY, USA, Ar...
2026
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.