REVIEW 3 major objections 5 minor 95 references
VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that VeasyGuide, a real-time motion-detection overlay that highlights and magnifies instructor pointing, marking, and sketching, lifts low-vision learners' action detection from 61% to 88% and cuts self-reported cognitive…
desk verdict A thoughtful cdesigned accessibility tool with real promise, but the headline detection gain rests on a between-video comparison that the paper's own analysis does not yet secure. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a two-stage activity recognition pipeline followed by a visualization module. The pipeline splits the video into shots with off-the-shelf shot detection, computes frame differences on one-third-second segments, keeps regions of change whose area exceeds $0.01\%$ of the frame, and builds a graph whose nodes are those regions, connected when they occur within 3 seconds and within 5% of the frame diagonal of each other. Edge weights come from Hu moments, seven shape-descriptor numbers that measure visual difference between regions, which lets the system merge or reject transient pointer trails. Connected components become activities, each rendered as a stationary, box-shaped, personalized highlight with a pointer icon and an optional auto-following zoom toggled with the Z key; a 1.5-second pre-activity trigger gives learners advance notice of where to look. The graph representation is what turns noisy per-frame motion into discrete, stable activities that can be highlighted and magnified.
What would settle it
Annotate a held-out set of presentation videos with the true locations of pointing, marking, and sketching actions, run VeasyGuide's pipeline on them, and compare the overlap of its activity boxes with the annotations; if low-vision users' success rate on those videos fails to beat the 61% baseline by a similar margin, the generalizability claim is not supported.
Extended reading notes
Core claim
The central claim is that the visual-search bottleneck for low-vision learners is not just the size of the content but knowing what to look for and where, and that an overlay can supply that knowledge in real time. On the paper's evidence, detection success rose from 61% to 88% (Welch's t-test $p<0.05$, Cohen's $d=1.30$), mean search time fell from 2.57s to 1.53s without reaching significance, and all NASA-TLX workload dimensions dropped significantly. In the guided condition no low-vision participant paused playback and fewer used screen magnification, suggesting the tool replaces part of the stop-and-search behavior of baseline viewing. The paper also reports that sighted viewers, already at ceiling on objective measures, subjectively preferred VeasyGuide for focus and comprehension, so the authors propose the mechanism as a broadly useful attention aid rather than only an accessibility tool.
Load-bearing premise
The load-bearing premise is that the fixed detection settings—how much motion counts as a region, how close regions must be in space and time to merge, and how short an activity can be—work on presentation videos beyond the six used in the study.
Editorial extensions
If this is right
- A learner who misses roughly four in ten instructor actions could miss roughly one in ten with the default settings, a per-participant mean improvement of 74.5%.
- The tool changes search strategy: in the guided condition no low-vision participant paused playback during the localization task, and screen-magnifier reliance dropped from six users to two.
- The design implications—consistent familiar visuals, predictable spatial context including a 1.5-second pre-trigger, real-time personalization with immediate feedback, and preserving user agency—can guide future accessibility tools for visual search in video.
- Because the pipeline uses lightweight motion detection and graph operations, the paper expects it to run on-device and to extend to live lectures and mainstream video platforms.
Reading between the lines
- The localization task fixed highlight style to system defaults, so the 88% figure is a default-settings result; testing with participants' own personalized styles might change detection rates in either direction and would separate the effect of personalization from the effect of highlighting per se.
- The pre-activity trigger likely explains part of the benefit; an ablation that removes the 1.5-second warning would isolate how much of the gain comes from anticipation versus visibility.
- The graph representation is blind to semantic content, so the same machinery could be pointed at other moving targets, such as cursors in coding screencasts or moving regions in sports and remote collaboration, by retuning only the thresholds.
- Since sighted viewers reported less clutter and better focus, the system may be a test bed for attention-as-accessibility design, where the same cue serves both perceptual and attentional functions.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents VeasyGuide, a web-based tool that uses motion detection to identify instructor actions (pointing, marking, sketching) in educational presentation videos and augments playback with personalized visual highlights and magnification. The system was developed through a co-design study with three low-vision (LV) participants. The evaluation involved 8 LV and 8 sighted participants in two tasks: a localization task (Task 1, videos V1–V4) measuring detection success and response time, and a viewing task (Task 2, videos V5–V6) measuring cognitive load and experience. The paper reports that VeasyGuide significantly improved LV users' detection of visual activities (61% to 88%), reduced NASA-TLX workload across all dimensions, and received positive feedback from sighted users, while detection time differences were not statistically significant. The main quantitative claim (RQ1) is based on an unpaired comparison of participants' success rates between conditions that used different video subsets, a design confound that the paper acknowledges but does not resolve.
Significance. If the reported effects hold, VeasyGuide addresses a real and under-explored accessibility barrier: LV learners missing instructor actions in slide-based videos. The co-design process and the derived design implications (DI1–DI4) are valuable and actionable, and the public web application makes the tool usable beyond the paper. The paper is honest about several limitations (Section 7), including the small sample and the lack of a systematic personalization evaluation in Task 1. The per-participant data in Table 7 are detailed and enable reanalysis, which is a strength. However, the headline success-rate claim currently rests on a between-video comparison, and the paper overstates the response-time finding, so the central contribution is not yet fully secured.
major comments (3)
- [§5.1.4, Table 7, Appendix E] The RQ1 success-rate analysis uses an unpaired Welch t-test on data that are paired at the participant level, but the baseline and VeasyGuide conditions used complementary subsets of videos V1–V4. Table 7 shows that the number of cued activities per condition differs within participants (e.g., L1: 11 vs 18; L7: 18 vs 11), and Appendix E reveals large per-video difficulty differences (V4 baseline success <25%). Consequently, the 61% vs 88% mean difference may be inflated or produced by which videos happened to appear in each condition rather than by the highlighting itself. Please reanalyze with a mixed-effects model that includes participant and video as random effects, report a stratified per-video comparison, or at minimum provide the paired per-participant difference and its confidence interval using the data in Table 7, and discuss how video difficulty is accounted for.
- [Abstract, §1, §5.2.2] The abstract and Section 1 claim that VeasyGuide yields 'faster response times' and 'improves the speed of detection,' but Section 5.2.2 states that the Mann-Whitney U test found no statistically significant difference (p > 0.05). Presenting the non-significant mean/median speed improvements as a demonstrated benefit overstates the evidence. Please either soften the speed claim to a non-significant trend or report a test that properly accounts for the paired design and video-level variation.
- [§4.1, §5.2.1] The activity recognition pipeline uses six hand-chosen thresholds (§4.1.1: RoC area 0.01% of frame, temporal closeness 3s, spatial closeness 5% of frame diagonal, Hu-moment merge threshold 0.5, minimum activity duration 5s, and 1.5s pre-trigger), but the paper reports no precision/recall evaluation of the detector on the six study videos. The user-success outcome in the VeasyGuide condition depends on the pipeline actually highlighting the cued activities; if recall is imperfect or the thresholds are implicitly tuned to the study videos, the 88% success rate is not a general property of the system. Please add a detection-accuracy evaluation (per-video recall/precision against the gold-standard activities) or a sensitivity analysis of the thresholds, and discuss the generalizability of these parameters to other presentation video styles.
minor comments (5)
- [§6.1.4] The statement that 'detection success gap reduced by 80.5%, and detection speed gap reduced by 27.3%' is not derived from the results in Section 5; please specify the formula (e.g., based on means or medians) or remove these numbers.
- [§5.1.3, Table 7] Please clarify how the total number of cued activities per condition in Table 7 was determined, including how V1–V4 were assigned to conditions for each participant and how activities shorter than one second were treated in the totals.
- [§5.1.1, Table 5] Two of the eight evaluation participants (L3 and L7) also participated in the co-design study; please acknowledge this overlap as a potential familiarity bias in the evaluation.
- [Figure 5] The horizontal axis of Figure 5 appears to be logarithmic; please label it as such or use a linear scale with clear units.
- [§4.2] The 1.5-second pre-activity highlight trigger is an important design element for predictability (DI3), but its isolated effect is not evaluated; please mention this explicitly in the limitations or future work.
Circularity Check
No circular derivation: VeasyGuide's central claims rest on an empirical user study with an external video corpus, and no fitted parameter is renamed as a prediction.
full rationale
I walked the paper's claimed derivation chain and found no circular step that reduces a prediction to its inputs. The activity recognition pipeline uses pre-specified, hand-chosen thresholds (0.01% of frame area, 3 s temporal closeness, 5% of frame diagonal spatial closeness, Hu-moment merge threshold 0.5, 5 s minimum activity duration, 1.5 s pre-trigger); these constants are not fit to the user-study outcome, and the paper does not claim they were derived from the measured success rates. The RQ1 result (61% vs. 88% mean success) is an empirical between-condition comparison on six external videos from YouTube, DeepLearning.AI, and Khan Academy, with success measured by participant keypresses rather than by any equation in the paper. Likewise, the NASA-TLX and reaction-time results are measured outcomes, not constructions. The self-citation to Sechayk et al. [66] for the initial 'highlight any notable visual change' prototype is a design-inspiration citation; the evaluation does not depend on the correctness of that prior work, and no uniqueness or forced-choice argument is imported from it. The participation of L3 and L7 in both co-design and evaluation, and the fact that Task 1 split V1-V4 between conditions, are genuine internal-validity concerns about attribution and generalizability, but they are not circularity: they do not make the measured effect true by definition or by fitted-parameter renaming. The reported statistics are consistent with a real, if possibly confounded, empirical comparison. Under the hard rule that circularity must be exhibited as a specific reduction in the paper's own reasoning, the burden is not met.
Assumptions & free parameters
free parameters (6)
- RoC area threshold =
0.01% of frame area
- Temporal closeness threshold =
3 seconds
- Spatial closeness threshold =
5% of frame diagonal
- Hu moments merge threshold =
0.5
- Minimum activity duration =
5 seconds
- Pre-activity highlight trigger =
1.5 seconds
assumptions (5)
- domain assumption Instructor actions such as pointing, marking, and sketching are frequent in educational presentation videos.
- domain assumption Low-vision users prefer to use residual vision rather than audio description.
- domain assumption Frame differencing with contour detection captures instructor actions in screen-shared presentation videos.
- domain assumption NASA-TLX self-reports are a valid measure of cognitive load in this setting.
- domain assumption The co-design findings from 3 participants generalize to the broader low-vision population.
Cite this review
Pith. "Pith review of VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos." pith.science (2026). https://pith.science/paper/FKZL247U
@misc{pith2026250721837,
author = {Pith},
title = {Pith review of: VeasyGuide: Personalized Visual Guidance for Low-vision Learners on Instructor Actions in Presentation Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/FKZL247U}},
note = {Machine review of arXiv:2507.21837}
}
read the original abstract
Instructors often rely on visual actions such as pointing, marking, and sketching to convey information in educational presentation videos. These subtle visual cues often lack verbal descriptions, forcing low-vision (LV) learners to search for visual indicators or rely solely on audio, which can lead to missed information and increased cognitive load. To address this challenge, we conducted a co-design study with three LV participants and developed VeasyGuide, a tool that uses motion detection to identify instructor actions and dynamically highlight and magnify them. VeasyGuide produces familiar visual highlights that convey spatial context and adapt to diverse learners and content through extensive personalization and real-time visual feedback. VeasyGuide reduces visual search effort by clarifying what to look for and where to look. In an evaluation with 8 LV participants, learners demonstrated a significant improvement in detecting instructor actions, with faster response times and significantly reduced cognitive load. A separate evaluation with 8 sighted participants showed that VeasyGuide also enhanced engagement and attentiveness, suggesting its potential as a universally beneficial tool.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Khan Academy. 2024. Elements and atoms | Atoms, compounds, and ions | Chemistry | Khan Academy. https://www.youtube.com/watch?v=IFKnq9QM6_A. Accessed: 2024-09-12
2024
-
[2]
Khan Academy. 2024. England in the Age of Exploration. https://www.youtube. com/watch?v=DZxG2miqgZI. Accessed: 2024-09-12
2024
-
[3]
Eler, Marouane Kessentini, and Ali Ouni
Wajdi Aljedaani, Mohammed Alkahtani, Stephanie Ludi, Mohamed Wiem Mkaouer, Marcelo M. Eler, Marouane Kessentini, and Ali Ouni. 2023. The State of Accessibility in Blackboard: Survey and User Reviews Case Study. In Proceedings of the 20th International Web for All Conference (<conf-loc>, <city>Austin</city>, <state>TX</state>, <country>USA</country>, </con...
-
[4]
Ali Selman Aydin, Shirin Feiz, Vikas Ashok, and IV Ramakrishnan. 2020. Towards making videos accessible for low vision screen magnifier users. In Proceedings of the 25th International Conference on Intelligent User Interfaces (Cagliari, Italy) (IUI ’20) . Association for Computing Machinery, New York, NY, USA, 10–21. doi:10.1145/3377325.3377494
arXiv 2020
-
[5]
Patrick Baudisch, Desney Tan, Maxime Collomb, Dan Robbins, Ken Hinckley, Maneesh Agrawala, Shengdong Zhao, and Gonzalo Ramos. 2006. Phosphor: explaining transitions in the user interface using afterglow effects. InProceedings of the 19th annual ACM symposium on User interface software and technology . ACM, New York, NY, USA, 169–178
2006
-
[6]
Porter, and I
Syed Masum Billah, Vikas Ashok, Donald E. Porter, and I. V. Ramakrishnan
-
[7]
Christian Bognar. 2022. Naughty Dog’s Obsession With Yellow Explained . https://gamerant.com/naughty-dog-yellow-color-coding-environments- progression-design/ Accessed: 2025-04-17
2022
-
[8]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597
2019
Show all 95 references
-
[9]
Sheryl E Burgstahler and Rebecca C Cory. 2010. Universal design in higher education: From principles to practice . Harvard Education Press
2010
-
[10]
Virgínia P Campos, Luiz MG Gonçalves, Wesnydy L Ribeiro, Tiago MU Araújo, Thaís G Do Rego, Pedro HV Figueiredo, Suanny FS Vieira, Thiago FS Costa, Caio C Moraes, Alexandre CS Cruz, et al. 2023. Machine generation of audio description for blind and visually impaired people.ACM ...
2023
-
[11]
Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. WorldScribe: Towards Context-Aware Live Visual Descriptions. In Proceedings of the 37th Annual ACM ASSETS ’25, October 26–29, 2025, Denver, CO, USA Sechayk et al. Symposium on User Interface Software and Technology . 1–18
2024
-
[12]
Coursera. [n. d.]. Coursera. https://www.coursera.org/. Accessed: 2023-10-05
2023
-
[13]
Cunningham, Joanne L
Sheila J. Cunningham, Joanne L. Brebner, Francis Quinn, and David J. Turk. 2014. The Self-Reference Effect on Memory in Early Child- hood. Child Development 85, 2 (2014), 808–823. doi:10.1111/cdev.12144 arXiv:https://srcd.onlinelibrary.wiley.com/doi/pdf/10.1111/cdev.12144
2014 doi
-
[14]
and Adam D
Robert D. and Adam D. 2024. PySceneDetect. https://github.com/Breakthrough/ PySceneDetect
2024
-
[15]
DeepLearning.AI. [n. d.]. DeepLearning.AI. https://www.deeplearning.ai/. Ac- cessed: 2024-08-10
2024
-
[16]
DeepLearningAI. 2024. #6 Machine Learning Specialization [Course 1, Week 1, Lesson 2]. https://www.youtube.com/watch?v=gG_wI_uGfIE. Accessed: 2024-09-12
2024
-
[17]
Laurent Denoue, Scott Carter, Matthew Cooper, and John Adcock. 2013. Real-time direct manipulation of screen-based videos. InProceedings of the companion publi- cation of the 2013 international conference on Intelligent user interfaces companion . 43–44
2013
-
[18]
NetworkX Developers. 2024. NetworkX. https://networkx.org/
2024
-
[19]
Alfred T D’Agostino. 2021. Accessible teaching and learning in the undergraduate chemistry course and laboratory for blind and low-vision students. Journal of Chemical Education 99, 1 (2021), 140–147
2021
-
[20]
doi:10.1145/3173574.3173594
-
[21]
Facebook
Inc. Facebook. 2024. React - A JavaScript library for building user interfaces. https://reactjs.org. Accessed: 2024-09-12
2024
-
[22]
Mirette Elias, Abi James, Edna Ruckhaus, Mari Carmen Suárez-Figueroa, Klaas An- dries De Graaf, Ali Khalili, Benjamin Wulff, Steffen Lohmann, and Sören Auer
-
[23]
In EC-TEL (Practitioner Proceedings)
SlideWiki-Towards a Collaborative and Accessible Platform for Slide Pre- sentations.. In EC-TEL (Practitioner Proceedings). 1–3
-
[24]
Fox, Ahmad Ahmadzada, Clara T
Dylan R. Fox, Ahmad Ahmadzada, Clara T. Friedman, Shiri Azenkot, Marlena A. Chu, Roberto Manduchi, and Emily A. Cooper. 2023. Using augmented reality to cue obstacles for people with low vision. Opt. Express 31, 4 (Feb 2023), 6827–6848. doi:10.1364/OE.479258
2023 doi
-
[25]
Tang, and Thomas Jaeger
Danyang Fan, Sasa Junuzovic, John C. Tang, and Thomas Jaeger. 2023. Improv- ing the Accessibility of Screen-Shared Presentations by Enabling Concurrent Exploration. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility, ASSETS 2023, N...
2023
-
[26]
Sue Grice and Janet Hughes. 2009. Can music and animation improve the flow and attainment in online learning?Journal of Educational Multimedia and Hypermedia 18, 4 (2009), 385–403
2009
-
[27]
Python Software Foundation. 2024. Python Programming Language. https: //www.python.org/
2024
-
[28]
Taralynn Hartsell and Steve Chi-Yin Yuen. 2006. Video streaming in online learning. AACE Review (Formerly AACE Journal) 14, 1 (2006), 31–43
2006
-
[29]
Sourojit Ghosh and Andrea Figueroa. 2023. Establishing TikTok as a Platform for Informal Learning: Evidence from Mixed-Methods Analysis of Creators and Viewers. In 56th Hawaii International Conference on System Sciences, HICSS 2023, Maui, Hawaii, USA, January 3-6, 2023, Tung X...
2023
-
[30]
Maija Hirvonen, Marika Hakola, and Michael Klade. 2023. Co-translation, consul- tancy and joint authorship: User-centred translation and editing in collaborative audio description. Journal of Specialised Translation 39 (2023), 26–51
2023
-
[31]
Tengda Han, Max Bain, Arsha Nagrani, Gül Varol, Weidi Xie, and Andrew Zisser- man. 2023. AutoAD: Movie description in context. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 18930–18940
2023
-
[32]
Gwo-Jen Hwang, Li-Hsueh Yang, and Sheng-Yuan Wang. 2013. A concept map- embedded educational computer game for improving students’ learning perfor- mance in natural science courses. Computers & Education 69 (2013), 121–130
2013
-
[33]
Alfonso Juan Hinojosa. 2015. Investigations on the impact of spatial ability and scientific reasoning of student comprehension in physics, state assessment tests, and STEM courses. The University of Texas at Arlington, Arlington, TX, USA
2015
-
[34]
Julie A Jacko and Andrew Sears. 1998. Designing interfaces for an overlooked user group: Considering the visual profiles of partially sighted users. In Proceedings of the third international ACM conference on Assistive technologies . 75–77
1998
-
[35]
Ming-Kuei Hu. 1962. Visual pattern recognition by moment invariants. IRE transactions on information theory 8, 2 (1962), 179–187
1962
-
[36]
Ji, Brianna R
Tiger F. Ji, Brianna R. Cochran, and Yuhang Zhao. 2022. VRBubble: Enhancing Peripheral Awareness of Avatars for People with Visual Impairments in Social Virtual Reality. In Proceedings of the 24th International ACM SIGACCESS Confer- ence on Computers and Accessibility, ASSETS ...
2022
-
[37]
Touhidul Islam and Syed Masum Billah
Md. Touhidul Islam and Syed Masum Billah. 2023. SpaceX Mag: An Automatic, Scalable, and Rapid Space Compactor for Optimizing Smartphone App Interfaces for Low-Vision Users. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7, 2 (2023), 59:1–59:36. doi:10.1145/3596253
2023 doi
-
[38]
Lucy Jiang and Richard Ladner. 2022. Co-Designing Systems to Support Blind and Low Vision Audio Description Writers. In Proceedings of the 24th International ACM SIGACCESS Conference on Computers and Accessibility (Athens, Greece) (ASSETS ’22). Association for Computing Machin...
2022
-
[39]
Gaurav Jain, Basel Hindi, Connor Courtien, Conrad Wyrick, Xin Yi Therese Xu, Michael C Malcolm, and Brian A. Smith. 2023. Towards Accessible Sports Broadcasts for Blind and Low-Vision Viewers. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Syste...
2023
-
[40]
Hyeungshik Jung, Hijung Valentina Shin, and Juho Kim. 2018. Dynamicslide: Exploring the design space of reference-based interaction techniques for slide- based lecture videos. In Proceedings of the 2018 Workshop on Multimedia for Accessible Human Computer Interface . ACM, New ...
2018
-
[41]
Lucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl, and Shiri Azenkot
-
[42]
Junhan Kong, Dena Sabha, Jeffrey P Bigham, Amy Pavel, and Anhong Guo. 2021. TutorialLens: authoring Interactive augmented reality tutorials through narration and demonstration. In Proceedings of the 2021 ACM Symposium on Spatial User Interaction. 1–11
2021
-
[43]
Richard E Ladner and Kyle Rector. 2017. Making your presentation accessible. Interactions 24, 4 (2017), 56–59
2017
-
[44]
Lucy Jiang, Mahika Phutane, and Shiri Azenkot. 2023. Beyond audio descrip- tion: Exploring 360 video accessibility with blind and low vision users through collaborative creation. In Proceedings of the 25th international ACM SIGACCESS conference on computers and accessibility . 1–17
2023
-
[45]
Open Source Computer Vision Library. 2024. OpenCV. https://opencv.org/
2024
-
[46]
Khan Academy. [n. d.]. Khan Academy. https://www.khanacademy.org/. Ac- cessed: 2024-08-10
2024
-
[47]
Chuan-Yu Mo, Chengliang Wang, Jian Dai, and Peiqi Jin. 2022. Video playback speed influence on learning effect from the perspective of personalized adaptive learning: A study based on cognitive load theory. Frontiers in Psychology 13 (2022), 839982
2022
-
[48]
Toni-Jan Keith Palma Monserrat, Shengdong Zhao, Kevin McGee, and An- shul Vikram Pandey. 2013. Notevideo: Facilitating navigation of blackboard- style lecture videos. In Proceedings of the SIGCHI conference on human factors in computing systems. 1139–1148
2013
-
[49]
Susan J Leat, Gordon E Legge, and Mark A Bullimore. 1999. What is low vision? A re-evaluation of definitions. Optometry and Vision Science 76, 4 (1999), 198–211
1999
-
[50]
Rosiana Natalie, Ruei-Che Chang, Smitha Sheshadri, Anhong Guo, and Kotaro Hara. 2024. Audio description customization. In Proceedings of the 26th Interna- tional ACM SIGACCESS Conference on Computers and Accessibility . 1–19
2024
-
[51]
Elke Mattheiss, Georg Regal, David Sellitsch, and Manfred Tscheligi. 2017. User- centred design with visually impaired pupils: A case study of a game editor for orientation and mobility training. International Journal of Child-Computer Interaction 11 (2017), 12–18. doi:10.1016...
2017 doi
- [52]
-
[53]
Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the CHI Conference on Hu...
2024
-
[54]
Zahra J Muhsin, Rami Qahwaji, Faruque Ghanchi, and Majid Al-Taee. 2024. Review of substitutive assistive tools and technologies for people with visual impairments: recent advancements and prospects. Journal on Multimodal User Interfaces 18, 1 (2024), 135–156
2024
-
[55]
Jaclyn Packer, Katie Vizenor, and Joshua A Miele. 2015. An overview of video description: history, benefits, and guidelines. Journal of Visual Impairment & Blindness 109, 2 (2015), 83–93
2015
-
[56]
Rosiana Natalie, Jolene Loh, Huei Suen Tan, Joshua Tseng, Ian Luke Yi-Ren Chan, Ebrima H Jarjue, Hernisa Kacorri, and Kotaro Hara. 2021. The efficacy of collaborative authoring of video scene descriptions. In Proceedings of the 23rd International ACM SIGACCESS Conference on Co...
2021
-
[57]
Amy Pavel, Gabriel Reyes, and Jeffrey P Bigham. 2020. Rescribe: Authoring and automatically editing audio descriptions. In Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology . 747–759
2020
-
[59]
American Council of the Blind. n.d.. Audio Description Project, Guidelines for Audio Describers. https://www.acb.org/adp/guidelines.html. Accessed: 2024-08- 10
2024
-
[60]
Yash Prakash, Akshay Kolgar Nayak, Sampath Jayarathna, Hae-Na Lee, and Vikas Ashok. 2024. Understanding Low Vision Graphical Perception of Bar Charts. In Proceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–10
2024
-
[61]
Suraiya Parveen and Javeria Shah. 2021. A motion detection system in python and opencv. In 2021 third international conference on intelligent communication technologies and virtual mobile networks (ICICV) . IEEE, IEEE, Virtual Conference, VeasyGuide ASSETS ’25, October 26–29, ...
2021
-
[62]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023. Robust speech recognition via large-scale weak supervision. In International conference on machine learning . PMLR, 28492–28518
2023
-
[63]
Ashwin Ram, Han Xiao, Shengdong Zhao, and Chi-Wing Fu. 2023. VidAdapter: Adapting Blackboard-Style Videos for Ubiquitous Viewing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 7, 3, Article 119 (sep 2023), 19 pages. doi:10. 1145/3610928
2023
-
[64]
Yi-Hao Peng, JiWoong Jang, Jeffrey P Bigham, and Amy Pavel. 2021. Say it all: Feedback for improving non-visual presentation accessibility. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12
2021
-
[65]
Anastasia Schaadhardt, Alexis Hiniker, and Jacob O Wobbrock. 2021. Understand- ing blind screen-reader users’ experiences of digital artboards. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–19
2021
-
[66]
Pallets Projects. 2024. Flask - A micro web framework for Python. https://flask. palletsprojects.com. Accessed: 2024-09-12
2024
-
[67]
Joel Snyder. 2005. Audio description: The visual made verbal. In International congress series, Vol. 1282. Elsevier, 935–939
2005
-
[68]
Abigale Stangl, Shasta Ihorn, Yue-Ting Siu, Aditya Bodi, Mar Castanon, Lothar D Narins, and Ilmi Yoon. 2023. The Potential of a Visual Dialogue Agent In a Tandem Automated Audio Description System for Videos. In Proceedings of the 25th International ACM SIGACCESS Conference on...
2023
-
[69]
Andreas Sackl, Franziska Graf, Raimund Schatz, and Manfred Tscheligi. 2020. Ensuring accessibility: Individual video playback enhancements for low vision users. In Proceedings of the 22nd International ACM SIGACCESS Conference on Computers and Accessibility. 1–4
2020
-
[70]
Lee Stearns, Leah Findlater, and Jon E Froehlich. 2018. Design of an augmented re- ality magnification aid for low vision users. InProceedings of the 20th international ACM SIGACCESS conference on computers and accessibility . 28–39
2018
-
[71]
Yotam Sechayk, Ariel Shamir, and Takeo Igarashi. 2024. SmartLearn: Visual- Temporal Accessibility for Slide-based e-learning Videos. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA 2024, Honolulu, HI, USA, May 11-16, 2024, Florian ’Flo...
2024 doi
-
[72]
Sarit Felicia Anais Szpiro, Shafeka Hashash, Yuhang Zhao, and Shiri Azenkot
-
[73]
Meini Tang, Roberto Manduchi, Susana Chung, and Raquel Prado. 2023. Screen Magnification for Readers with Low Vision: A Study on Usability and Perfor- mance. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–15
2023
-
[74]
Fleischmann, Meredith Ringel Morris, and Danna Gurari
Abigale Stangl, Nitin Verma, Kenneth R. Fleischmann, Meredith Ringel Morris, and Danna Gurari. 2021. Going Beyond One-Size-Fits-All Image Descriptions to Satisfy the Information Wants of People Who are Blind or Have Low Vision. In Proceedings of the 23rd International ACM SIGA...
2021
-
[75]
Chris Thompson
Dr. Chris Thompson. 2024. Introduction to Neuroscience 2: Lecture 26: Neu- roethology. https://www.youtube.com/watch?v=7pQyw1rHEEg. Accessed: 2024-09-12
2024
-
[76]
Satoshi Suzuki et al. 1985. Topological structural analysis of digitized binary images by border following. Computer vision, graphics, and image processing 30, 1 (1985), 32–46
1985
-
[77]
The Organic Chemistry Tutor. 2024. Biology - Intro to Cell Structure - Quick Review! https://www.youtube.com/watch?v=vwAJ8ByQH2U. Accessed: 2024- 09-12
2024
-
[78]
Ru Wang, Zach Potter, Yun Ho, Daniel Killough, Linxiu Zeng, Sanbrita Mondal, and Yuhang Zhao. 2024. GazePrompt: Enhancing Low Vision People’s Reading Experience with Gaze-Aware Augmentations. InProceedings of the CHI Conference on Human Factors in Computing Systems, CHI 2024, ...
2024
-
[79]
Ru Wang, Linxiu Zeng, Xinyong Zhang, Sanbrita Mondal, and Yuhang Zhao. 2023. Understanding how low vision people read using eye tracking. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM, New York, NY, USA, 1–17
2023
-
[80]
ThatsEngineering. 2024. Frame Assignment For Robotic Manipulators - Direct Kinematics I. https://www.youtube.com/watch?v=fNIyNF87q9I. Accessed: 2024- 09-12
2024
-
[81]
Yanan Wang, Yuhang Zhao, and Yea-Seul Kim. 2024. How Do Low-Vision Individ- uals Experience Information Visualization?. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–15
2024
-
[82]
Ayaka Tsutsui, Kenta Yamamoto, Yinan Zhao, Ippei Suzuki, Kengo Tanaka, and Yoichi Ochiai. 2024. Low Vision Boxing: Participatory Design of Adap- tive Kickboxing Experiences with Low Vision Person. In The 26th Interna- tional ACM SIGACCESS Conference on Computers and Accessibil...
2024
-
[83]
YouTube. [n. d.]. YouTube. https://www.youtube.com/. Accessed: 2024-08-10
2024
-
[84]
Beste F Yuksel, Pooyan Fazli, Umang Mathur, Vaishali Bisht, Soo Jung Kim, Joshua Junhee Lee, Seung Jung Jin, Yue-Ting Siu, Joshua A Miele, and Ilmi Yoon
-
[85]
Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR Visualizations to Facilitate Stair Navigation for People with Low Vision. InProceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology, ...
2019
-
[86]
Yujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang, Dingzeyu Li, and Lap-Fai Yu. 2021. Toward automatic audio description generation for accessible videos. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12
2021
-
[87]
Yuhang Zhao, Sarit Szpiro, Jonathan Knighten, and Shiri Azenkot. 2016. CueSee: exploring visual cues for people with low vision to facilitate a visual search task. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing (Heidelberg, ...
2016
-
[88]
Carmen Yip, Jie Mi Chong, Sin Yee Kwek, Yong Wang, and Kotaro Hara. 2021. Visionary Caption: Improving the Accessibility of Presentation Slides through Highlighting Visualization. In Proceedings of the 23rd International ACM SIGAC- CESS Conference on Computers and Accessibility . 1–4
2021
-
[89]
Zoom. [n. d.]. Zoom. https://zoom.us/. Accessed: 2024-08-10. A Empirical Evaluation of Visual Activities in Presentation Videos To better understand how visual activities appear in educational presentation videos in the wild, we conducted an analysis of 300 YouTube [83] videos...
2024
-
[93]
Yuhang Zhao, Elizabeth Kupferstein, Hathaitorn Rojnirun, Leah Findlater, and Shiri Azenkot. 2020. The Effectiveness of Visual and Audio Wayfinding Guid- ance on Smartglasses for People with Low Vision. In CHI ’20: CHI Confer- ence on Human Factors in Computing Systems, Honolul...
2020
-
[95]
Yuhang Zhao, Sarit Szpiro, Lei Shi, and Shiri Azenkot. 2020. Designing and Evaluating a Customizable Head-mounted Vision Enhancement System for Peo- ple with Low Vision. ACM Trans. Access. Comput. 12, 4 (2020), 15:1–15:46. doi:10.1145/3361866
2020 doi
-
[2016]
InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility
How people with low vision access computing devices: Understanding chal- lenges and opportunities. InProceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility . 171–180
-
[2018]
In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI 2018, Montreal, QC, Canada, April 21-26, 2018 , Regan L
SteeringWheel: A Locality-Preserving Magnification Interface for Low Vision Web Browsing. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems, CHI 2018, Montreal, QC, Canada, April 21-26, 2018 , Regan L. Mandryk, Mark Hancock, Mark Perry, and Anna L...
2018
-
[2020]
In Proceedings of the 2020 ACM Designing Interactive Systems Conference
Human-in-the-loop machine learning to increase video accessibility for visually impaired and blind users. In Proceedings of the 2020 ACM Designing Interactive Systems Conference. 47–60
2020
-
[2023]
doi:10.1145/3597638.3608411
ACM, 44:1–44:16. doi:10.1145/3597638.3608411
-
[2024]
It’s Kind of Context Dependent
"It’s Kind of Context Dependent": Understanding Blind and Low Vision People’s Video Accessibility Preferences Across Viewing Scenarios. In Proceed- ings of the CHI Conference on Human Factors in Computing Systems, CHI 2024, Honolulu, HI, USA, May 11-16, 2024, Florian ’Floyd’ M...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.